Touch detection for an input interface on a physical surface using detection of a surface tap
The head-mounted display device employs sensor fusion to overcome hand occlusion issues, enhancing touch detection accuracy and reducing false positives for virtual interfaces by confirming surface taps.
Patent Information
- Application Number
- PCT/US2024/023740
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-09
- Publication Date
- 2025-10-16
AI Technical Summary
Conventional display devices face challenges in accurately detecting surface touches due to hand occlusion, leading to false positives in virtual keyboard interactions.
A head-mounted display device uses a combination of sensor data from multiple sources, including computer vision, inertial measurement units, microphones, and LiDAR, to enhance touch detection accuracy by confirming surface taps through sensor fusion.
The system provides a more robust and accurate input interface by reducing false positives, enabling precise interactions with virtual keyboards and other interfaces on various surfaces.
Smart Images

Figure US2024023740_16102025_PF_FP_ABST
Abstract
Description
TOUCH DETECTION FOR AN INPUT INTERFACE ON A PHYSICAL SURFACE USING DETECTION OF A SURFACE TAPBACKGROUND
[0001] Some conventional display devices may place a virtual keyboard on a physical surface. However, it may be difficult for an outward-facing camera to detect a surface touch because the hand may occlude the finger.SUMMARY
[0002] This disclosure relates to a head-mounted display device that includes an input interface manager that detects a user’s interaction with an input interface on a physical surface using first sensor data captured by a head-mounted display device and / or second sensor data captured by one or more computing devices (a smartphone, a tablet, a laptop, or a wearable device such as a smartwatch, etc.) connected (e.g., wirelessly connected) to the head-mounted display device. To overcome one or more technical challenges of false positives driven by computer vision, the input interface manager may combine (e.g., fuse) sensor data from multiple sources, e.g., first sensor data for computer vision (e.g., coarse computer vision) for body part location (e.g., hand location), and second sensor data from a nearby microphone, inertial measurement unit (IMU), accelerator, light detection and ranging (e.g., LiDAR), and / or other sensing for tap detection. In some examples, a computing device with sparse sensing is in contact (e.g., direct contact) with the physical surface in which the input interface is positioned (e.g., anchored to) or the computing device is in contact (e.g., direct contact) with a user (e.g., worn by the user). By combining sensor data from multiple sensors, the head-mounted display device may provide a more robust and accurate input interface (e.g., a virtual keyboard, virtual computer mouse, etc.) that may be used on a plurality of different surfaces, including a body portion of the user such as a leg, a hand or an arm (e.g., any surface).
[0003] In some aspects, the techniques described herein relate to a method including: rendering an input interface on a physical surface; receiving image data from a sensor system on a head-mounted display device; receiving inertial measurement unit data from a computing device connected to the head-mounted display device; and detecting a touch interaction to aselectable element on the input interface using the image data and the inertial measurement unit data.
[0004] In some aspects, the techniques described herein relate to a computer program product storing executable instructions that when executed by at least one processor causes the at least one processor to execute operations, the operations including: rendering an input interface on a physical surface; receiving image data from a sensor system on a head-mounted display device; receiving inertial measurement unit data from a computing device connected to the headmounted display device; and detecting a touch interaction to a selectable element on the input interface using the image data and the inertial measurement unit data.
[0005] In some aspects, the techniques described herein relate to a head-mounted display device including: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to execute operations, the operations including: rendering an input interface on a physical surface; receiving first sensor data from a sensor system of the head-mounted display device; receiving second sensor data from a computing device connected to the head-mounted display device, the second sensor data including motion data; and detecting a touch interaction on the input interface using the first sensor data and the second sensor data, including determining whether a surface tap occurred on the physical surface using the motion data.
[0006] The details of one or more implementations are set forth in the accompanying drawings and the description below. Other features will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. 1 A illustrates a system for detecting a touch interaction on an input interface on a physical surface using multi-sensor sensor data according to an aspect.
[0008] FIG. IB illustrates a system for detecting a touch interaction on an input interface on a physical surface using multi-sensor sensor data according to another aspect.
[0009] FIG. 1C illustrates an example of the system with a head-mounted display device and a computing device according to an aspect.
[0010] FIG. ID illustrates an input interface manager with a tap detector that detects a surface tap on a physical surface according to an aspect.
[0011] FIG. 2 illustrates an input interface manager according to another aspect.
[0012] FIG. 3 illustrates an input interface manager according to another aspect.
[0013] FIG. 4 illustrates a flowchart depicting example operations of detecting a touch interaction on an input interface on a physical surface using multisensory sensor data according to an aspect.DETAILED DESCRIPTION
[0014] This disclosure relates to a head-mounted display device that includes an input interface manager configured to detect a touch interaction on an input interface rendered on a physical surface (e.g., a table, a body part, etc.) using first sensor data captured by the headmounted display device and / or second sensor data captured by a computing device (e.g., a smartphone, a tablet, a laptop, a wearable device such as a smartwatch, etc.) connected (e.g., wirelessly connected) to the head-mounted display device. The input interface manager includes a tap detector that determines whether a surface tap occurred on the physical surface using the first sensor data and / or the second sensor data, and the occurrence (or lack of occurrence) of the surface tap may be used to determine whether a touch interaction has occurred. In other words, the input interface manager may determine a selection to a selectable element of the input interface based on computer vision processing of image data, and that selection may be confirmed by the presence of a surface tap. If the surface tap is not detected, the input interface manager may determine that the image-based selection is a false positive, and thereby not detect a touch interaction. The input interface may be any type of interface that is selectable by human touch. In some examples, the input interface includes a virtual keyboard. However, the virtual input interface may be other types of input interfaces such as a virtual computer mouse, a computer interface, or any virtual user interface with one or more selectable elements. The input interface manager may overcome one or more technical problems for detecting a touch interaction to an input interface on a physical surface using multi-sensor sensor data, where the use of multi-sensor sensor data for tap detection on a physical surface may provide one or more technical benefits of increased tap accuracy (e.g., reducing or eliminating false positives).
[0015] The input interface manager may render an input interface on a physical surface. For example, the input interface manager may render a virtual keyboard, a virtual computer mouse, or a virtual user interface with selectable elements on a physical surface. The inputinterface manager may detect a selection to a selectable element (e.g., a particular key, right or left click of a virtual computer mouse, etc.) using the first image sensor data captured by the head-mounted display device. The first image sensor data may include image data from one or more cameras on the head-mounted display device. In some examples, using image data alone for detecting selection of a particular selectable element may result in selection inaccuracies because the user’s body part that makes the selection (e.g., a finger) may be occluded by another body part (e.g., other finger) or other object. However, the input interface manager may detect a surface tap on the physical surface using the first sensor data and / or the second sensor data, and the detection (or lack of detection) may be used to confirm or discard a selection of the selectable element. The input interface manager may provide one or more technical benefits of higher selection accuracy by including additional sensor data (e.g., besides image data from the headmounted display device) and using that additional sensor data to detect the occurrence of a surface tap on the physical surface.
[0016] In some examples, the second sensor data may include an inertial measurement unit (IMU) data from one or more IMUs on the computing device. In some examples, the computing device is positioned on the same surface that includes the input interface. In some examples, the computing device is a device that is not on the same surface but worn by the user or in the same room as the user. The IMU data may include information about the movement and / or position of the computing device. In some examples, a tap (e.g., finger tap) on the physical surface (e.g., a table or other flat surface) may cause vibrations which cause slight movements to the computing device, which may be detected by the IMU(s) of the computing device that is positioned on the same surface as the input interface. In some examples, the computing device is worn by the user, and the IMU data may include movement of the computing device when the user has moved their finger (or other body part) to select a selectable element on the input interface. Using the IMU data, the input interface manager may detect whether a surface tap has occurred. In some examples, the second sensor data includes acoustic data captured by one or more microphones on the computing device or the head-mounted display device. The acoustic data may include sound waves that are indicative of a surface tap. In some examples, the IMU data and the acoustic data are used to determine whether a surface tap occurred on the physical surface. In some examples, the second sensor data may include accelerator data. In some examples, the second sensor data may include light detection andranging (e.g., LiDAR) data. In some examples, the second sensor data may include other sensing data that may indicate a surface tap.
[0017] The input interface manager may use a sensor fusion approach from one or more nearby computing devices (e.g., a smartphone, a laptop, a wearable device such as a smartwatch, smart ring, etc.) to detect a surface tap on a physical surface, and the presence (or lack of presence) of the surface tap may assist with detecting whether a touch interaction occurred with respect to the input interface on the physical surface. In some examples, the input interface manager may use the second sensor data (e.g., the acoustic data and / or the IMU data or other sensing data from a connected computing device) for modeling touch (or no touch) with a hand finger (e.g., any hand finger). In some examples, the input interface manager combines the second sensor data with one or more headset based vision modules (e.g., computer vision engine, physical engine, etc.) to determine touch interactions (and, in some examples, finger localization). In some examples, the input interface manager may detect a touch interaction when computer vision detects a selection (e.g., when a hand is in contact with a selectable element of the input interface) and a surface tap detector detects an occurrence of a surface tap on the physical surface, which may provide one or more benefits of input accuracy including minimizing (or eliminating) false positives (e.g., when computer vision determines an imagebased selection but the user has not made a selection).
[0018] In some examples, the input interface manager includes a computer vision engine configured to receive image data and generate positional data about the user’s tracked body part relative to the physical surface. In some examples, the positional data includes the approximate location of the tracked body part (e.g., locations of the user’s finger). Computer vision may provide a general understanding of the user's hand position relative to the surface, which may be used to estimate the approximate location (e.g., 2D or 3D location) of the user's fingers, even when they are not directly visible to the camera.
[0019] In some examples, the input interface manager includes a tap detector that uses IMU data and / or acoustic data or other sensing data to determine whether or not a surface tap occurred on the physical surface. For example, one or more microphones on the computing device and / or the head-mounted display device may obtain sound waves generated by finger taps on the surface. By analyzing these sound signatures, the tap detector may determine the location and timing of taps, which may increase the accuracy of detecting touch interactions on the inputinterface. The IMU data may include information about the movements and accelerations of the computing device. In some examples, the IMU data may be used in conjunction with the acoustic data, which can provide additional confirmation of tap events and may enhance the overall accuracy of the system. In some examples, the input interface manager includes a physics engine, where the physics engine combined with the finger location may provide a spatiotemporal scale to enable multi- finger discrimination. For example, each finger’s tip may be a collider, and the physics engine may detect which was the last key to be accessed by one of the moving fingers. Other sensing data may include accelerator data, LiDAR data, or sensing for tap detection.
[0020] By combining the data from the sensors on the head-mounted display device and / or the computing device, the input interface manager may render an input interface (e.g., a virtual input interface) that is more reliable, responsive, and / or intuitive to use compared to some conventional approaches. The techniques discussed herein may reduce the over reliance on outward-facing cameras and their susceptibility to hand occlusion, by reducing false positives with event synchronization and sensor fusion. In some examples, the input interface manager may enable users to interact with input interfaces on a physical surface with greater precision and comfort. In some examples, the input interface manager may avoid the use of physical input devices such as keyboards and / or computing mouses.
[0021] FIGS. 1A to ID illustrate a system 100 with a head-mounted display device 102 having an input interface manager 124 that detects a touch interaction 125 on an input interface 152 positioned on a physical surface 115 using multi-sensor sensor data 122. The multi-sensor sensor data 122 includes at least one of sensor data 122a (e.g., first sensor data) captured by the head-mounted display device 102 or sensor data 122b (e.g., second sensor data) captured by a computing device 110 connected (e.g., wirelessly connected) to the head-mounted display device 102. The multi-sensor sensor data 122 can include sensor data 122a (e.g., first sensor data) captured by the head-mounted display device 102 and sensor data 122b (e.g., second sensor data) captured by a computing device 110 connected (e.g., wirelessly connected) to the head-mounted display device 102. The multi-sensor sensor data 122 may include data from two or more sensors on the head-mounted display device 102 and / or the computing device 110. Also, although one computing device 110 is illustrated in FIGS. 1A to ID, the head-mounted display device 102 may communicate with two or more computing devices 110 to obtain multi-sensorsensor data 122 across two or more computing devices 1 10. For example, the head-mounted display device 102 may receive sensor data from a first computing device (e.g., a smartphone), sensor data from a second computing device (e.g., a laptop), and / or sensor data from a third computing device (e.g., a smartwatch), and / or sensor data from a third computing device (e.g., a smart ring, etc.), and so forth.
[0022] In some examples, to overcome one or more technical challenges of false positives driven by computer vision, the input interface manager 124 may combine (e.g., fuse) sensor data from multiple sources. In some examples, the input interface manager 124 may receive first sensor data (e.g., image data 130a) for computer vision (e.g., coarse computer vision) for body part estimation (e.g., hand location), and second sensor data (e.g., IMU data 132b, acoustic data 134b, LiDAR data, etc.) from a nearby microphone (e.g., microphone 112a, microphone 112b) and / or IMU (e.g., IMU 108b) for tap detection. In some examples, a computing device 110 with an IMU 108b is in contact (e g., direct contact) with the physical surface 115 in which the input interface 152 is anchored to or the computing device 110 is in contact (e.g., direct contact) with a user (e.g., worn by the user). By combining the information from the sensors on the head-mounted display device 102 and / or the computing device 110, the input interface manager 124 may provide a more robust and accurate input interface 152 that may be used on a plurality of different surfaces (e.g., any surface).
[0023] In some examples, the head-mounted display device 102 is an XR device. In some examples, the head-mounted display device 102 is an AR device. In some examples, the head-mounted display device 102 is a VR device. The head-mounted display device 102 may include an optical head-mounted display (OHMD) device, a transparent heads-up display (HUD) device, an augmented reality (AR) device, or other devices such as goggles or headsets having sensors, display, and computing capabilities.
[0024] The input interface 152 may be any type of user interface that is selectable by human touch. The input interface 152 may be a virtual input device. In some examples, the input interface 152 is a virtual keyboard. In some examples, the input interface 152 is a virtual computer mouse. In some examples, the input interface 152 is a computer interface. Generally, the input interface 152 may be any type of interface such as a UI object, menu, virtual input device, and / or control. The input interface 152 includes one or more selectable elements 154. The selectable elements 154 may be icons, controls, or other elements, which, when selected,causes an action to be performed. In some examples, the selectable elements 154 are different keys of a virtual keyboard. In some examples, the selectable elements 154 are selectable controls on a virtual computer mouse such as a right “click” control or a left “click” control. In some examples, the selectable elements 154 are selectable user interface (UI) elements on a computer interface.
[0025] The head-mounted display device 102 includes a sensor system 104a. The sensor system 104a includes one or more cameras 106a configured to detect image data 130a (e.g., image frames) in the camera’s field of view. The image data 130a includes color data, and, in some examples, grayscale data for each of a plurality of pixels. In some examples, the camera (s) 106a includes one or more imaging sensors. In some examples, the camera(s) 106a includes a stereo pair of image sensors. In some examples, the camera(s) 106a includes a visual see through (VST) camera system that allows the user to see the real world through the camera’s lens while also seeing digital information (e.g., input interface 152) overlaid on the real world (e.g., a physical surface 115). In some examples, the camera(s) 106a may be referred to as an AR camera, a mixed reality camera, a head-mounted display camera, a transparent display camera, or a combiner camera. The camera’s field of view may be the angular extent of the scene that is captured by the camera(s) 106a. The field of view may be measured in degrees and may be specified as a horizontal field of view and / or a vertical field of view.
[0026] The sensor system 104a includes an inertial measurement unit (IMU) 108a configured to generate IMU data 132a about an acceleration and / or velocity of the head-mounted display device 102. In some examples, the IMU 108a may be referred to as a motion sensor, and the IMU data 132a may be referred to as motion data, wherein the motion data includes acceleration and / or rotation along two or three axes. In some examples, the IMU data 132a includes head movement data. The IMU 108a includes an accelerometer configured to measure an acceleration of the head-mounted display device 102 and generate accelerometer data. The accelerometer data includes information about the acceleration of the head-mounted display device 102, e.g., acceleration in an x-axis, a y-axis, and a z-axis. The IMU 108a includes a gyroscope configured to measure a velocity of the head-mounted display device 102 and generate gyroscope data. The gyroscope data includes information about the velocity of the head-mounted display device 102, e.g., information about the velocity in the x-axis, the y-axis, and the z-axis.
[0027] The sensor system 104a includes one or more microphones 112a configured to generate acoustic data 134a about the sound waves around the head-mounted display device 102. The acoustic data 134a includes an audio signal, and, in some examples, metadata associated with the audio signal. The audio signal may be a digital representation of the sound waves captured by the microphone(s) 112a. The audio signal may include digital data that represent the amplitude and frequency of the sound at specific moments in time. The metadata may include information about the sample rate (e.g., how many times per second the sound wave is being sampled and converted to a digital value), a bit depth (e.g., the number of bits used to represent the amplitude of each sample) and / or channel information about whether the recording is mono or a stereo channel.
[0028] The sensor system 104a may include one or more depth sensors 114a configured to generate depth data 136a about objects in the scene. The depth data 136a includes information about the relative distances of objects in the camera scene. In some examples, the depth data 136a includes a depth map. In some examples, the depth sensor(s) 114a includes time-of-flight (ToF) sensors. In some examples, the sensor system 104a includes a light detection and ranging (LiDAR) sensor. The LiDAR sensor may emit a laser, which may be used to compute the distances of objects. The sensor system 104a may include other sensors such as a magnetometer, an ambient light sensor, infrared sensor, an ultrasonic sensor, a microwave sensor, a gravity sensor, a curve vector sensor, and / or a tomographic sensor. In some examples, the sensor system 104a may include other sensors such as electrooculography (e.g., measures the electrical potential changes around the eyes) and / or electromagnetic coils (e.g., creating a magnetic field that interacts with the eye’s conductivity, recording movements).
[0029] The input interface manager 124 may obtain sensor data 122a from the sensor system 104a. In some examples, the sensor data 122a includes the image data 130a. In some examples, the sensor data 122a includes the IMU data 132a. In some examples, the sensor data includes the acoustic data 134a. In some examples, the sensor data 122a includes the depth data 136a. In some examples, the sensor data 122a includes a combination of two or more of the image data 130a, the IMU data 132a, the acoustic data 134a, or the depth data 136a.
[0030] The computing device 110 may be any type of user device with one or more sensors discussed herein. The computing device 110 may be a smartphone, a laptop, a tablet, or a wearable device (e.g., a smartwatch, fitness tracker, smart clothing, smart glasses, a hearabledevicejewelry wearable, etc.). The computing device 1 10 may include a sensor system 104b. The sensor system 104b may include one or more of the sensors included in the sensor system 104a on the head-mounted display device 102. For example, the sensor system 104b may include one or more IMUs 108b, one or more microphones 112b, one or more cameras 106b, and / or one or more depth sensors 114b.
[0031] The IMU 108b generates IMU data 132b about an acceleration and / or velocity of the computing device 110. In some examples, the IMU data 132b includes acceleration data. In some examples, the IMU data 132b includes velocity data. In some examples, the IMU data 132b includes device movement and / or orientation data. The IMU 108b includes an accelerometer configured to measure an acceleration of the computing device 110 and generate accelerometer data. The accelerometer data includes information about the acceleration of the computing device 110, e.g., acceleration in an x-axis, a y-axis, and a z-axis. The IMU 108b includes a gyroscope configured to measure a velocity of the computing device 110 and generate gyroscope data. The gyroscope data includes information about the velocity of the computing device 110, e.g., information about the velocity in the x-axis, the y-axis, and the z-axis. In some examples, the IMU 108b may be referred to as a motion sensor, and the IMU data 132b may be referred to as motion data, wherein the motion data includes acceleration and / or rotation along two or three axes.
[0032] The microphone(s) 112b may generate acoustic data 134b about the sound waves around the computing device 1 10. The acoustic data 134a includes an audio signal, and, in some examples, metadata associated with the audio signal. The audio signal may be a digital representation of the sound waves captured by the microphone(s) 112b. The audio signal may include digital data that represent the amplitude and frequency of the sound at specific moments in time. The metadata may include information about the sample rate (e.g., how many times per second the sound wave is being sampled and converted to a digital value), a bit depth (e.g., the number of bits used to represent the amplitude of each sample) and / or channel information about whether the recording is mono or a stereo channel.
[0033] The depth sensor(s) 114b may generate depth data 136b about objects in the scene. The depth data 136b includes information about the relative distances of objects in the camera scene. In some examples, the depth data 136b includes a depth map. In some examples, the depth sensor(s) 114b includes time-of-flight (ToF) sensors. In some examples, the sensorsystem 104b includes a light detection and ranging (LiDAR) sensor. The LiDAR sensor may emit a laser, which may be used to compute the distances of objects. The sensor system 104b may include other sensors such as a magnetometer, an ambient light sensor, infrared sensor, an ultrasonic sensor, a microwave sensor, a gravity sensor, a curve vector sensor, and / or a tomographic sensor. In some examples, the sensor system 104b may include other sensors such as electrooculography (e.g., measures the electrical potential changes around the eyes) and / or electromagnetic coils (e.g., creating a magnetic field that interacts with the eye’s conductivity, recording movements).
[0034] The input interface manager 124 may receive the sensor data 122b from one or more computing devices 110 connected to the head-mounted display device 102. In some examples, the computing device 110 is wirelessly connected to the head-mounted display device 102 over a network. In some examples, the network includes a short range communication network such as a Bluetooth network or a near-field communication (NFC) network. In some examples, the network includes an Internet network such as a Wi-Fi network. In some examples, the computing device 110 is connected to the head-mounted display device 102 via one or more wired protocols (e.g., one or more physical cords). In some examples, the sensor data 122b includes the IMU data 132b. In some examples, the sensor data 122b includes motion data. In some examples, the sensor data 122b includes acceleration data. In some examples, the sensor data 122b includes the acoustic data 134b. In some examples, the sensor data 122b includes the IMU data 132b and the acoustic data 134b. In some examples, the sensor data 122b includes the image data 130b. In some examples, the sensor data 122b includes the depth data 136b. In some examples, the sensor data 122b includes the LiDAR sensor data.
[0035] The input interface manager 124 may render the input interface 152 on a physical surface 115 (e.g., a real world surface) on a display 190 of the head-mounted display device 102. In some examples, the input interface manager 124 may render a virtual representation 117b of the user’s hands 117a in a position that is above the input interface 152. In some examples, the input interface manager 124 may detect a flat surface from the image data 130a and position the input interface 152 at an orientation in which the input interface 152 appears to be located on top of the physical surface 115. In some examples, the input interface manager 124 uses information (e.g., positional data about the position and orientation of the user’s hands in 3D space) from a hand tracker (e.g., hand tracker 241 of FIG. 2), and positions the input interface 152 at a location(e g., a fixed location) on the physical surface 1 15 that is proximate (e.g., below) the user’s hands. In some examples, if the user moves their hands to a different location, the input interface manager 124 may move the input interface 152 to the hands’ new location such that the input interface 152 is aligned with the user’s hands.
[0036] In some examples, the input interface 152 may include one or more 3D objects (e.g., a 3D object in the form of an input device such as a virtual keyboard or a virtual mouse). In some examples, the input interface 152 includes one or more 2D objects (e.g., a flat keyboard or virtual computer mouse, or a computer interface). In some examples, the input interface manager 124 may animate the input interface 152 (or a portion thereof) in response to user interaction with one or more selectable elements 154 of the input interface 152. For example, in response to the selection of a selectable element 154 (e.g., a user taps the physical surface on a virtual key), the input interface manager 124 may animate a portion of the input interface 152 such as changing the color of a pressed virtual key or other animation effects.
[0037] The input interface manager 124 may detect a touch interaction 125 to a selectable element 154 using the image data 130a captured by the head-mounted display device 102, the sensor data 122a from one or more other sensors on the head-mounted display device 102 and / or sensor data 122b from one or more sensors on the computing device 110. Using the image data 130a alone for detecting a selection of a selectable element 154 may result in selection inaccuracies because the user’s body part (e.g., a finger) that makes the selection may be occluded by another body part (e.g., other finger) or other object. In some examples, the input interface manager 124 may detect the touch interaction 125 using the image data 130a and at least one of the IMU data 132a, the acoustic data 134a, the depth data 136a, the IMU data 132b, the acoustic data 134b, the image data 130b, or the depth data 136b.
[0038] In some examples, the input interface manager 124 includes a tap detector 126 configured to detect a surface tap 128 on the physical surface 115 using the sensor data 122a and / or the sensor data 122b. In some examples, the detection (or lack of detection) of the surface tap 128 may be used to confirm or discard a selection to a selectable element 154 of the input interface 152. In some examples, the input interface manager 124 may use computer vision (e.g., image data 130a, and / or, in some examples, image data 130b) to detect a selection to a particular selectable element 154 (e.g., a user’s finger has interacted with the selectable element 154). The input interface manager 124 may use at least one of the IMU data 132a, the acoustic data 134a,the IMU data 132b, or the acoustic data 134b to confirm or discard the selection identified by computer vision. The input interface manager 124 may provide one or more technical benefits of higher selection accuracy on the input interface 152 by using detection of a surface tap 128 using the IMU data 132b and the acoustic data 134b (and, in some examples, the acoustic data 134a).
[0039] In some examples, the tap detector 126 determines whether a surface tap 128 occurred on the physical surface 115 using at least one of the IMU data 132b or the acoustic data 134b. In some examples, the tap detector 126 determines whether a surface tap 128 occurred on the physical surface 115 using the IMU data 132b. In some examples, the tap detector 126 determines whether a surface tap 128 occurred on the physical surface 115 using the acoustic data 134b. In some examples, the tap detector 126 determines whether a surface tap 128 occurred on the physical surface 115 using the acoustic data 134b and the IMU data 132b. In some examples, the tap detector 126 determines whether a surface tap 128 occurred on the physical surface 115 using the acoustic data 134a. In some examples, the tap detector 126 determines whether a surface tap 128 occurred on the physical surface using at least one of the acoustic data 134a, the acoustic data 134b, or the IMU data 132b.
[0040] A surface tap 128 may be determined when a user’s body part (e.g., user’s finger) taps on the physical surface 115. A surface tap 128 may refer to a quick touch on the physical surface 115. In some examples, a surface tap 128 may be referred to as a press or a touch. In some examples, the determination that the surface tap 128 occurred on the surface tap 128 indicates that the user has tapped the physical surface 115. In some examples, a surface tap 128 may be referred to as a press or a touch by a user’s finger with sufficient force that it produces a sound. In some examples, the computing device 110 is positioned on the same surface (e.g., the physical surface 115) that includes the input interface 152. In some examples, the computing device 110 is a device that is not on the same surface (e.g., the physical surface 115), but worn by the user or in the same physical space as the user.
[0041] The IMU data 132b may include information about the movement and / or position of the computing device 110. In some examples, a surface tap 128 on the physical surface 115 may cause vibrations which cause movements (e.g., slight movements) to the computing device 110, which may be detected by the IMU(s) 108b of the computing device 110 that is positioned on the same surface as the input interface 152. In some examples, the computing device 110 is worn by the user, and the IMU data 132b may include movement of the computing device 110when the user has moved their finger (or other body part) to select a selectable element 154. The acoustic data 134b or acoustic data 134a may include an audio signal with a portion that represents the sound of a tap on the physical surface 115.
[0042] In some examples, the tap detector 126 includes a machine-learning (ML) model configured to determine (e.g., estimate, predict, etc.) whether a surface tap 128 occurred on the physical surface 115 using at least one of the acoustic data 134a, the acoustic data 134b, or the IMU data 132b as input(s) to the ML model. In some examples, the ML model is a neural network. The ML model may be trained with training data to predict the occurrence of a surface tap 128 on the physical surface 115 using at least one of the acoustic data 134a, the acoustic data 134b, or the IMU data 132b as input(s) to the ML model.
[0043] In some examples, the input interface manager 124 detects a selection to a selectable element 154 on the input interface 152 using the image data 130a. The selection to a selectable element 154 on the input interface 152 can be detected from the image data 130a, such as for example by analyzing the image data 130a. In some examples, the input interface manager 124 detects a selection to the selectable element 154 using at least one of the image data 130a or the image data 130b. In some examples, the input interface manager 124 may estimate a location of a user’s finger using the image data 130a and determine that the location of the user’s finger corresponds to a location of the selectable element 154. In some examples, the input interface manager 124 may not detect a touch interaction 125 (e.g., disregard or ignore the selection to the selectable element 154) in response to the surface tap 128 being determined as not occurred on the physical surface 115. In some examples, the input interface manager 124 may select the selectable element 154 in response to the surface tap 128 being determined as occurred on the physical surface 115.
[0044] In some examples, the input interface manager 124 may determine that a selection to a selectable element 154 on the input interface 152 occurred at a first time and / or a first location using the sensor data 122a and determine that the surface tap 128 occurred on the physical surface 115 at a second time and / or a second location using the sensor data 122b. In some examples, the input interface manager 124 does not select the selectable element 154 in response to a temporal difference between the first time and the second time being greater than a first threshold level and / or to a spatial difference between the first location and the second location being greater than a second threshold level. In some examples, the input interfacemanager 124 selects the selectable element 154 in response to the temporal difference between the first time and the second time being equal to or less than the first threshold level and / or to the spatial difference between the first location and the second location being equal to or less than the second threshold level.
[0045] The input interface manager 124 may use a sensor fusion approach from one or more nearby computing devices 110 (e.g., a smartphone, a laptop, a wearable device such as a smartwatch, etc.) to determine whether a surface tap 128 occurred on the physical surface 115, and the presence (or lack of presence) of the surface tap 128 on the physical surface 115 may assist with detecting whether a touch interaction 125 has occurred with respect to the input interface 152. In some examples, the input interface manager 124 may use the sensor data 122b (e.g., the acoustic data 134b and / or the IMU data 132b) for modeling touch (or no touch) with a hand finger (e g., any hand finger). In some examples, the input interface manager 124 combines the sensor data 122b with one or more headset based vision modules (e.g., a computer vision engine 240 of FIG. 2, a physical engine of FIG. 2, etc.) to determine touch interactions 125 (and, in some examples, finger localization). In some examples, the input interface manager 124 may detect a touch interaction 125 when a hand is in contact with a selectable element 154 of the input interface 152 and a surface tap 128 on the physical surface 115 has occurred, which may provide one or more benefits of input accuracy including minimizing (or eliminating) false positives.
[0046] The head-mounted display device 102 may include one or more processors 101, one or more memory devices 103, and an operating system 105 configured to execute one or more applications 107. In some examples, the operating system 105 includes the input interface manager 124. In some examples, an application 107, executable by the operating system 105, includes the input interface manager 124. The processor(s) 101 may be formed in a substrate configured to execute one or more machine executable instructions or pieces of software, firmware, or a combination thereof. The processor(s) 101 can be semiconductor-based - that is, the processors can include semiconductor material that can perform digital logic. The memory device(s) 103 may include any type of storage device that stores information in a format that can be read and / or executed by the processor(s) 101. In some examples, the memory device(s) 103 is / are a non-transitory computer-readable medium. In some examples, the memory device(s) 103 includes a non-transitory computer-readable medium that includes executable instructions thatcause at least one processor (e.g., the processor(s) 101) to execute operations discussed with reference to the head-mounted display device 102. The applications 107 may be any type of computer program that can be executed by the head-mounted display device 102, including native applications that are installed on the operating system 105 by the user and / or system applications that are pre-installed on the operating system 105. The input interface 152 may be any interface associated with the operating system 105, an application 107, or generally any component of the head-mounted display device 102.
[0047] FIG. 2 illustrates an input interface manager 224 according to an aspect. The input interface manager 224 may be an example of the input interface manager 124 of FIGS. 1A to ID and may include any of the details discussed with reference to those figures. The input interface manager 224 may be included in a head-mounted display device 102 of FIGS. 1A to ID.
[0048] The input interface manager 224 may render an input interface 252 on a display of the head-mounted display device in which the input interface 252 is positioned on a physical surface. As shown in FIG. 2, the input interface 252 may be a virtual keyboard with a plurality of selectable elements 254 (e.g., virtual keys). For example, the selectable elements 254 include a selectable element 254-1, a selectable element 254-2, and a selectable element 254-3.However, the selectable elements 254 may include any number of selectable elements 254, including a number that corresponds to the number of virtual keys on a virtual keyboard. Further, although FIG. 2 illustrates a virtual keyboard as the input interface 252, the input interface 252 may be any type of input interface such as a computer screen or a virtual computer mouse.
[0049] The input interface manager 224 may detect a touch interaction on the input interface 252 using multi-sensor sensor data 222, and the detection of the touch interaction is based on whether or not a surface tap 228 occurred on the physical surface. The multi-sensor sensor data 222 may include image data 230a obtained from one or more cameras on the headmounted display device. The multi-sensor sensor data 222 may include IMU data 232b and / or acoustic data 234b received from one or more computing devices (e.g., the computing device 110 of FIGS. 1A to ID). However, it is noted that the multi-sensor sensor data 222 may include two or more types of sensor data as discussed with reference to FIGS. 1A to ID, which may be from the same device or different devices. In some examples, the input interface manager 224may detect a touch interaction on the input interface 252 using the image data 230a and at least one of the IMU data 232b or the acoustic data 234b.
[0050] The input interface manager 224 includes a computer vision engine 240 configured to receive image data 230a and generate positional data 235 about the user’ s tracked body part 244 (e.g., fingers) relative to the physical surface (e.g., the physical surface 115 of FIGS. 1 A to ID). In some examples, the positional data 235 includes the location(s) 242 of the tracked body part 244 (e.g., locations 242 of the user’s fingers). In some examples, the computer vision engine 240 includes a hand tracker 241 configured to generate positional data 235 about the user’s fingers. In some examples, the computer vision engine 240 may estimate the locations 242 of the user's fingers, even when they are not directly visible to the camera.
[0051] The input interface manager 224 includes a physics engine 260 that uses the positional data 235 to determine a selection to a particular selectable element 254 on the input interface 152. The physics engine 260 may determine whether a location 242 of a finger corresponds to a location of a selectable element 254 on the input interface 252. For example, if the location 242 of the index finger corresponds to the location of a selectable element 254-2, the physics engine 260 may determine that the selectable element 254-2 is a selected element 254a.
[0052] In some examples, the physics engine 260 may determine a selection time 264 associated with the selected element 254a, where the selection time 264 is a time in which the user’s finger has interacted with a boundary of the selected element 254a. In some examples, a finger (e g., each finger) has a collider 292 on the finger’s tip. The physics engine 260 may detect which was the last selectable element 254 to be accessed by one of the moving fingers. For example, the physics engine 260 may create a virtual representation 217b of a user’s hands. The virtual representation 117b may include separate colliders 292 for each finger (e.g., collider 292-1, collider 292-2, collider 292-3, collider 292-4, collider 292-5). The colliders 292 can be shaped like spheres, capsules, or other shapes depending on the desired level of detail. The physics engine 260 is configured to simulate how objects in the virtual world interact with each other. When a virtual finger collider (e.g., collider 292-3) comes into contact with another object (e.g., selectable element 254-2) in the virtual world that also has a collider, the physics engine 260 calculates how they should interact.
[0053] The input interface manager 224 includes a tap detector 226 configured to receive the acoustic data 234b and / or the IMU data 232b and use the acoustic data 234b and / or the IMUdata 232b to determine whether a surface tap 228 occurred on the physical surface. Detection of a surface tap 228 indicates that the user has tapped the physical surface. In some examples, the computing device (e.g., computing device 110 of FIGS. 1A to ID), which generated the IMU data 232b and / or the acoustic data 234b, is positioned on the same surface (e.g., the physical surface 115 of FIGS. 1A to ID) that includes the input interface 252. In some examples, the computing device is a device that is not on the same surface (e.g., the physical surface 115 of FIGS. 1 A to ID) but worn by the user or in the same physical space as the user.
[0054] The IMU data 232b may include information about the movement and / or position of the computing device. In some examples, a surface tap 228 on the physical surface may cause vibrations which cause movements (e.g., slight movements) to the computing device, which may be detected by the IMU(s) (e.g., the IMU(s) 108b of FIGS. 1A to ID) of the computing device. In some examples, the computing device is worn by the user, and the IMU data 232b may include movement of the computing device when the user has moved their finger (or other body part) to select a portion of the input interface 252. The acoustic data 234b may include an audio signal with a portion that represents the sound of a tap on the physical surface.
[0055] The tap detector 226 includes a machine-learning (ML) model 248 configured to determine (e.g., estimate, predict, etc.) whether a surface tap 228 occurred on the physical surface using the acoustic data 234b and / or the IMU data 232b as input(s) to the ML model 248. In some examples, the ML model 248 is a neural network 255. The ML model 248 may be trained with training data to predict the occurrence of a surface tap 228 using the acoustic data 234b and / or the IMU data 232b as input(s) to the ML model 248. In some examples, the tap detector 226 may determine where and when the surface tap 228 occurred on the physical surface. For example, the tap detector 226 may determine a source location 256 of the surface tap 228 and a time 258 in which the surface tap 228 occurred on the physical surface. The source location 256 may be the location of where the surface tap 228 has occurred.
[0056] The input interface manager 224 detects a touch interaction with a selected element 254a when the tap detector 226 determines that a surface tap 228 occurred on the physical surface using the acoustic data 234b and / or the IMU data 232b. In other words, the input interface manager 224 detects a touch interaction with a selectable element 254 when the physics engine 260 identifies a selected element 254a from the location 242 of a last moving finger and the tap detector 226 determines that a surface tap 228 occurred on the physicalsurface. In some examples, the input interface manager 224 does not detect a touch interaction with a selected element 254a when a surface tap 228 has not occurred on the physical surface.
[0057] In some examples, the input interface manager 224 may determine whether there is temporal alignment between a selection of a selectable element 254 using the image data 230a and a determination of a surface tap 228 on the physical surface using the acoustic data 234b and / or the IMU data 232b. For example, the input interface manager 224 determines whether the temporal difference between the selection time 264 and the time 258 of the surface tap 228 is within a threshold level (e.g., a first threshold level), and, if the temporal difference is equal to or below the threshold level, selects the selected element 254a. In some examples, the input interface manager 224 does not select the selected element 254a when the temporal difference between the selection time 264 and the time 258 of the surface tap 228 is greater than the threshold level (e g., the touch interaction is a false positive).
[0058] In some examples, the input interface manager 224 may determine whether there is spatial alignment between a selection of a selectable element 254 using the image data 230a and a determination of a surface tap 228 on the physical surface using the acoustic data 234b and / or the IMU data 232b. For example, the input interface manager 224 determines whether the spatial difference between the source location 256 of the surface tap 228 and the location 242 of the user’s finger is within a threshold level (e.g., a second threshold level), and, if so, detects the touch interaction with the selected element 254a (e.g., selects the selected element 254a). In some examples, the input interface manager 224 determines whether the spatial difference between the source location 256 of the surface tap 228 and the location 242 of the user’s finger is within a threshold level, and, if not, the input interface manager 224 does not detect the touch interaction with the selected element 254a (e.g., the selected element 254a is not selected (e.g., the touch interaction is a false positive).
[0059] FIG. 3 illustrates an input interface manager 324 according to an aspect. The input interface manager 324 may be an example of the input interface manager 124 of FIGS. 1A to ID and / or the input interface manager 224 of FIG. 3 and may include any of the details discussed with reference to those figures. The input interface manager 324 may be included in a head-mounted display device 102 of FIGS. 1A to ID.
[0060] In operation 370, the input interface manager 324 detects a selected element (e.g., selected element 254a of FIG. 2) on an input interface using a location 342 of the user’s finger(e g., last moving finger) and the location of the selectable element. When a location 342 of the user’s finger corresponds to a location of a selectable element, the input interface manager 324 may detect that selectable element as the selected element. In some examples, the input interface manager 324 may detect that a collider on a virtual representation of the user’s finger has interacted with a selectable element’s collider.
[0061] In operation 372, the input interface manager 324 may determine that the last finger has moved within a threshold period of time from the detection of the selected element. In operation 374, the input interface manager 324 may determine whether a surface tap occurred on the physical surface using the acoustic data 334b and / or the IMU data 332b from a computing device connected to the head-mounted display device. In response to a determination that a surface tap has not occurred on the physical surface, in operation 380, the input interface manager 324 may not select the selected element (e.g., the touch interaction is a false positive). In other words, although the input interface manager 324 may detect that the user has selected a selectable element from the image data, the input interface manager 324 may detect a false position if a surface tap has not occurred.
[0062] In response to a surface tap and a selection to a selectable element using the location 342 of the last moving figure (and, in some examples, the last finger is detected as moved within a threshold period of time), in operation 376, the input interface manager 324 may determine temporal and / or spatial alignment between the surface tap and the selection to the selectable image (e g., by the physics engine). For example, the input interface manager 324 determines whether the temporal difference between a selection time (e.g., the selection time 264 of FIG. 2) of the selected element and the time (e.g., the time 258 of FIG. 2) of the surface tap is within a threshold level (e.g., a first threshold level), and, if the difference is equal to or less than the threshold level, the input interface manager 324 determines the temporal alignment to be true. The input interface manager 324 may determine whether the spatial difference between a source location (e.g., the source location 256 of FIG. 2) of the surface tap and the location (e.g., the location 242 of FIG. 2) of the user’s finger is within a threshold level (e.g., the second threshold level), and, if so, determines the spatial alignment to be true. In operation 378, in response to the temporal alignment and / or the spatial alignment being true, the input interface manager 324 selected the selected element (e.g., presses the virtual key). In response to thetemporal alignment and / or the spatial alignment being false, in operation 380, the input interface manager 324 does not select the selected element (e.g., the touch interaction is a false positive).
[0063] FIG. 4 is a flowchart 400 depicting example operations of a head-mounted display device for detecting a touch interaction on an input interface positioned on a physical surface according to an aspect. The flowchart 400 may depict operations of a computer-implemented method. Although the flowchart 400 is explained with respect to the system 100 of FIGS. 1 A to ID, the operations may be executed by any of the examples discussed herein. Although the flowchart 400 of FIG. 4 illustrates the operations in sequential order, it will be appreciated that this is merely an example, and that additional or alternative operations may be included. Further, operations of FIG. 4 and related operations may be executed in a different order than that shown, or in a parallel or overlapping fashion.
[0064] Operation 402 includes rendering an input interface on a physical surface. Operation 404 includes obtaining first sensor data from a sensor system on a head-mounted display device. The first sensor data can be received from the sensor system and can comprise image data. Operation 406 includes receiving second sensor data from a computing device connected to the head-mounted display device, the second sensor data including motion data, which may include inertial measurement unit (IMU) data. Operation 408 includes detecting a touch interaction on the input interface using the first sensor data and the second sensor data, including determining whether a surface tap occurred on the physical surface using at least the IMU data. The touch interaction to a selectable element on the input interface can be detected using the image data and the motion data (e.g., the inertial measurement unit data).
[0065] Clause 1. A method comprising: rendering an input interface on a physical surface; receiving image data from a sensor system on a head-mounted display device; receiving inertial measurement unit data from a computing device connected to the head-mounted display device; and detecting a touch interaction to a selectable element on the input interface using the image data and the inertial measurement unit data.
[0066] Clause 2. The method of clause 1, wherein detecting the touch interaction includes: detecting a selection of the selectable element based on the image data; determining whether a surface tap occurred on the physical surface based on the inertial measurement unit data; and detecting the touch interaction in response to the surface tap being determined as occurred on the physical surface.
[0067] Clause 3. The method of clause 2, wherein detecting the selection of the selectable element based on the image data includes: estimating a location of a user's finger based on the image data; and determining that the location of the user's finger corresponds to a location of the selectable element.
[0068] Clause 4. The method of clause 2, wherein determining whether the surface tap occurred on the physical surface using the inertial measurement unit data includes: determining, by a model, a presence of the surface tap on the physical surface using the inertial measurement unit data as an input to the model.
[0069] Clause 5. The method of clause 4, further comprising: receiving acoustic data captured by the head-mounted display device or the computing device; and determining, by the model, the presence of the surface tap on the physical surface using the inertial measurement unit data and the acoustic data as inputs to the model.
[0070] Clause 6. The method of any one of clauses 1 to 5, wherein detecting the touch interaction includes: determining that a selection to the selectable element occurred at a first time using the image data; determining that a surface tap occurred on the physical surface at a second time using the inertial measurement unit data; and detecting the touch interaction in response to a temporal difference between the first time and the second time not achieving a threshold level.
[0071] Clause 7. The method of any one of clauses 1 to 5, wherein detecting the touch interaction includes: determining that a selection of the selectable element occurred at a first location using the image data; determining that a surface tap occurred on the physical surface at a second location using the inertial measurement unit data; and detecting the touch interaction in response to a spatial difference between the first location and the second location not achieving a threshold level.
[0072] Clause 8. The method of any one of clauses 1 to 5, further comprising: receiving acoustic data from one or more microphones on the computing device; and detecting the touch interaction based on the image data, the inertial measurement unit data, and the acoustic data.
[0073] Clause 9. A computer program product storing executable instructions that when executed by at least one processor causes the at least one processor to execute operations, the operations comprising: rendering an input interface on a physical surface; receiving image data from a sensor system on a head-mounted display device; receiving inertial measurement unit data from a computing device connected to the head-mounted display device; and detecting atouch interaction to a selectable element on the input interface using the image data and the inertial measurement unit data.
[0074] Clause 10. The computer program product of clause 9, wherein the operations further comprise: detecting a selection of the selectable element based on the image data; determining whether a surface tap occurred on the physical surface based on the inertial measurement unit data; and detecting the touch interaction in response to the surface tap being determined as occurred on the physical surface.
[0075] Clause 11. The computer program product of clause 10, wherein the operations further comprise: estimating a location of a user's finger based on the image data; and determining that the location of the user's finger corresponds to a location of the selectable element.
[0076] Clause 12. The computer program product of clause 10, wherein the operations further comprise: determining, by a model, a presence of the surface tap on the physical surface using the inertial measurement unit data as an input to the model.
[0077] Clause 13. A head-mounted display device comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to execute operations, the operations comprising: rendering an input interface on a physical surface; receiving first sensor data from a sensor system of the head-mounted display device; receiving second sensor data from a computing device connected to the head-mounted display device, the second sensor data including motion data; and detecting a touch interaction on the input interface using the first sensor data and the second sensor data, including determining whether a surface tap occurred on the physical surface using the motion data.
[0078] Clause 14. The head-mounted display device of clause 13, wherein the first sensor data includes image data from one or more cameras on the head-mounted display device, wherein the operations further comprise: detecting a selection to a selectable element on the input interface using the image data; ignoring the selection to the selectable element in response to the surface tap being determined as not occurred on the physical surface; and selecting the selectable element in response to the surface tap being determined as occurred on the physical surface.
[0079] Clause 15. The head-mounted display device of clause 14, wherein the operations further comprise: estimating a location of a user's finger based on the image data; and detectingthe selection to the selectable element by determining that the location of the user's finger corresponds to a location of the selectable element.
[0080] Clause 16. The head-mounted display device of any one of clauses 13 to 15, wherein the first sensor data includes image data from one or more cameras on the head-mounted display device, wherein the operations further comprise: determining that a selection to a selectable element on the input interface occurred at a first time and a first location using the first sensor data; determining that the surface tap occurred on the physical surface at a second time and a second location using the second sensor data; and selecting the selectable element in response to at least one of a temporal difference between the first time and the second time not achieving a first threshold level or a spatial difference between the first location and the second location not achieving a second threshold level.
[0081] Clause 17. The head-mounted display device of any one of clauses 13 to 16, wherein the second sensor data includes acoustic data from one or more microphones on the computing device, wherein the operations further comprise: determining whether the surface tap occurred on the physical surface using the motion data and the acoustic data.
[0082] Clause 18. The head-mounted display device of clause 17, wherein the operations further comprise: determining, by a model, a presence of the surface tap on the physical surface using the motion data and the acoustic data as inputs to the model.
[0083] Clause 19. The head-mounted display device of any one of clauses 13 to 18, wherein the first sensor data includes acoustic data from one or more microphones on the headmounted display device, wherein the operations further comprise: determining whether the surface tap occurred on the physical surface using the motion data and the acoustic data.
[0084] Clause 20. The head-mounted display device of any one of clauses 13 to 19, wherein the motion data is first motion data and the computing device is a first computing device, wherein the operations further comprise: receiving third sensor data from a second computing device connected to the head-mounted display device, the third sensor data including second motion data; and determining whether the surface tap occurred on the physical surface using the second sensor data and the third sensor data. .
[0085] Various implementations of the systems and techniques described here can be realized in digital electronic circuitry, integrated circuitry, specially designed ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinationsthereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0086] These computer programs (also known as programs, software, software applications or code) include machine instructions for a programmable processor and can be implemented in a high-level procedural and / or object-oriented programming language, and / or in assembly / machine language. As used herein, the terms “machine-readable medium” “computer- readable medium” refers to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine- readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0087] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0088] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communicationnetwork). Examples of communication networks include a local area network (“LAN”), a wide area network (“WAN”), and the Internet.
[0089] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0090] In this specification and the appended claims, the singular forms "a," "an" and "the" do not exclude the plural reference unless the context clearly dictates otherwise. Further, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context clearly dictates otherwise. For example, “A and / or B” includes A alone, B alone, and A with B. Further, connecting lines or connectors shown in the various figures presented are intended to represent example functional relationships and / or physical or logical couplings between the various elements. Many alternative or additional functional relationships, physical connections or logical connections may be present in a practical device. Moreover, no item or component is essential to the practice of the implementations disclosed herein unless the element is specifically described as “essential” or “critical”.
[0091] Terms such as, but not limited to, approximately, substantially, generally, etc. are used herein to indicate that a precise value or range thereof is not required and need not be specified. As used herein, the terms discussed above will have ready and instant meaning to one of ordinary skill in the art.
[0092] Moreover, use of terms such as up, down, top, bottom, side, end, front, back, etc. herein are used with reference to a currently considered or illustrated orientation. If they are considered with respect to another orientation, it should be understood that such terms must be correspondingly modified.
[0093] Further, in this specification and the appended claims, the singular forms "a," "an" and "the" do not exclude the plural reference unless the context clearly dictates otherwise. Moreover, conjunctions such as “and,” “or,” and “and / or” are inclusive unless the context clearly dictates otherwise. For example, “A and / or B” includes A alone, B alone, and A with B.
[0094] Although certain example methods, apparatuses and articles of manufacture have been described herein, the scope of coverage of this patent is not limited thereto. It is to be understood that terminology employed herein is for the purpose of describing particular aspectsand is not intended to be limiting. On the contrary, this patent covers all methods, apparatus and articles of manufacture fairly falling within the scope of the claims of this patent.
Claims
WHAT IS CLAIMED IS:
1. A method compri sing : rendering an input interface on a physical surface; receiving image data from a sensor system on a head-mounted display device; receiving inertial measurement unit data from a computing device connected to the headmounted display device; and detecting a touch interaction to a selectable element on the input interface using the image data and the inertial measurement unit data.
2. The method of claim 1, wherein detecting the touch interaction includes: detecting a selection of the selectable element based on the image data; determining whether a surface tap occurred on the physical surface based on the inertial measurement unit data; and detecting the touch interaction in response to the surface tap being determined as occurred on the physical surface.
3. The method of claim 2, wherein detecting the selection of the selectable element based on the image data includes: estimating a location of a user’s finger based on the image data; and determining that the location of the user’s finger corresponds to a location of the selectable element.
4. The method of claim 2, wherein determining whether the surface tap occurred on the physical surface using the inertial measurement unit data includes: determining, by a model, a presence of the surface tap on the physical surface using the inertial measurement unit data as an input to the model.
5. The method of claim 4, further comprising: receiving acoustic data captured by the head-mounted display device or the computing device; anddetermining, by the model, the presence of the surface tap on the physical surface using the inertial measurement unit data and the acoustic data as inputs to the model.
6. The method of any one of claims 1 to 5, wherein detecting the touch interaction includes: determining that a selection to the selectable element occurred at a first time using the image data; determining that a surface tap occurred on the physical surface at a second time using the inertial measurement unit data; and detecting the touch interaction in response to a temporal difference between the first time and the second time not achieving a threshold level.
7. The method of any one of claims 1 to 5, wherein detecting the touch interaction includes: determining that a selection of the selectable element occurred at a first location using the image data; determining that a surface tap occurred on the physical surface at a second location using the inertial measurement unit data; and detecting the touch interaction in response to a spatial difference between the first location and the second location not achieving a threshold level.
8. The method of any one of claims 1 to 5, further comprising: receiving acoustic data from one or more microphones on the computing device; and detecting the touch interaction based on the image data, the inertial measurement unit data, and the acoustic data.
9. A computer program product storing executable instructions that when executed by at least one processor causes the at least one processor to execute operations, the operations comprising: rendering an input interface on a physical surface; receiving image data from a sensor system on a head-mounted display device; receiving inertial measurement unit data from a computing device connected to the headmounted display device; anddetecting a touch interaction to a selectable element on the input interface using the image data and the inertial measurement unit data.
10. The computer program product of claim 9, wherein the operations further comprise: detecting a selection of the selectable element based on the image data; determining whether a surface tap occurred on the physical surface based on the inertial measurement unit data; and detecting the touch interaction in response to the surface tap being determined as occurred on the physical surface.
11. The computer program product of claim 10, wherein the operations further comprise: estimating a location of a user’s finger based on the image data; and determining that the location of the user’s finger corresponds to a location of the selectable element.
12. The computer program product of claim 10, wherein the operations further comprise: determining, by a model, a presence of the surface tap on the physical surface using the inertial measurement unit data as an input to the model.
13. A head-mounted display device comprising: at least one processor; and a non-transitory computer-readable medium storing executable instructions that cause the at least one processor to execute operations, the operations comprising: rendering an input interface on a physical surface; receiving first sensor data from a sensor system of the head-mounted display device; receiving second sensor data from a computing device connected to the headmounted display device, the second sensor data including motion data; and detecting a touch interaction on the input interface using the first sensor data and the second sensor data, including determining whether a surface tap occurred on the physical surface using the motion data.
14. The head-mounted display device of claim 13, wherein the first sensor data includes image data from one or more cameras on the head-mounted display device, wherein the operations further comprise: detecting a selection to a selectable element on the input interface using the image data; ignoring the selection to the selectable element in response to the surface tap being determined as not occurred on the physical surface; and selecting the selectable element in response to the surface tap being determined as occurred on the physical surface.
15. The head-mounted display device of claim 14, wherein the operations further comprise: estimating a location of a user’s finger based on the image data; and detecting the selection to the selectable element by determining that the location of the user’s finger corresponds to a location of the selectable element.
16. The head-mounted display device of any one of claims 13 to 15, wherein the first sensor data includes image data from one or more cameras on the head-mounted display device, wherein the operations further comprise: determining that a selection to a selectable element on the input interface occurred at a first time and a first location using the first sensor data; determining that the surface tap occurred on the physical surface at a second time and a second location using the second sensor data; and selecting the selectable element in response to at least one of a temporal difference between the first time and the second time not achieving a first threshold level or a spatial difference between the first location and the second location not achieving a second threshold level.
17. The head-mounted display device of any one of claims 13 to 16, wherein the second sensor data includes acoustic data from one or more microphones on the computing device, wherein the operations further comprise:determining whether the surface tap occurred on the physical surface using the motion data and the acoustic data.
18. The head-mounted display device of claim 17, wherein the operations further comprise: determining, by a model, a presence of the surface tap on the physical surface using the motion data and the acoustic data as inputs to the model.
19. The head-mounted display device of any one of claims 13 to 18, wherein the first sensor data includes acoustic data from one or more microphones on the head-mounted display device, wherein the operations further comprise: determining whether the surface tap occurred on the physical surface using the motion data and the acoustic data.
20. The head-mounted display device of any one of claims 13 to 19, wherein the motion data is first motion data and the computing device is a first computing device, wherein the operations further comprise: receiving third sensor data from a second computing device connected to the headmounted display device, the third sensor data including second motion data; and determining whether the surface tap occurred on the physical surface using the second sensor data and the third sensor data.
Citation Information
Patent Citations
Virtual keyboard
WO2024049463A1