Systems and methods for human interaction with virtual objects

Through a system combining touch-sensitive surfaces with reference layers, using image sensors or ultrasonic technology, accurate virtual object interactions in augmented reality or virtual reality systems are achieved, solving the problems of user fatigue and immersion reduction, and providing reliable input motion tracking and hand rest.

CN119356519BActive Publication Date: 2025-07-08FIREFLY DIMENSION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411218675.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-08-02
Filing Date
2019-08-02
Publication Date
2025-07-08
Estimated Expiration
2039-08-02

AI Technical Summary

Technical Problem

In the augmented reality and virtual reality systems, direct interaction methods cause user fatigue, indirect interaction methods reduce immersion, making it difficult to achieve accurate, responsive, and reliable input motion tracking.

Method used

A system that combines touch-sensitive surfaces and reference layers is used to detect and identify user interactions with virtual objects through image sensors or ultrasonic technology, and combine display devices and processors to achieve accurate virtual object tracking and interaction.

Benefits of technology

保持了增强现实或虚拟现实系统的沉浸感,提供了精确、响应的、可靠的交互体验,并为用户提供了手部休息的物理表面。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119356519B_ABST
    Figure CN119356519B_ABST
Patent Text Reader

Abstract

A system for human interaction with a virtual object, comprising: a touch-sensitive surface configured to detect the position of a contact made on the touch-sensitive surface; a reference layer rigidly attached to the touch-sensitive surface and including one or more patterns; a display device configured to display a virtual object recorded in reference coordinates fixed relative to the touch-sensitive surface; one or more image sensors rigidly attached to the display device, configured to capture an image of at least a portion of the one or more patterns; and at least one processor configured to determine the position and orientation of the display device relative to the touch-sensitive surface based on the captured image, and to identify an interaction with the virtual object based on the detected position of the contact made on the touch-sensitive surface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application is a divisional application of the application named "Systems and Methods for Human-Virtual Object Interaction" with the application number 201980064378.9 (PCT / US2019 / 044939), filed on August 2, 2019.

[0002] Related Applications

[0003] This application claims the priority of U.S. Provisional Application No. 62 / 713,887, filed on August 2, 2018, the entire content of which is incorporated herein by reference. Technical Field

[0004] This disclosure generally relates to the fields of augmented reality and virtual reality. Specifically, this disclosure relates to the interaction between humans and virtual objects. Background Art

[0005] The interaction between users and virtual objects is a hot topic widely studied in the current fields of augmented reality and virtual reality. Common interaction methods can generally be divided into direct interaction methods and indirect interaction methods. The direct interaction method tracks the user's gestures and postures or the spatial position of a handheld device, so that the handheld device can be accurately recorded in the virtual space, and the virtual space itself can be accurately overlaid or displayed on top of the real world. As a result, the user can directly interact with the virtual object at its display position. The indirect interaction method typically includes an indicator for interacting with virtual elements, and the user controls the movement and actions of the indicator via an input device. In addition, the movement of the indicator can be scaled or offset relative to the recorded input movement, so that for a small interaction area, the user can interact with virtual objects that are far away or distributed in a large space.

[0006] The direct interaction method can employ technologies such as hand tracking, body tracking, and six-degree-of-freedom handheld controllers. The hand tracking method typically has a gesture capture and recognition system, where gestures can be captured by an image acquisition sensor or a glove equipped with orientation and bending sensors to reconstruct the gesture. Then, similar to the situation where a hand interacts with a real object in the physical world, the reconstructed gesture or hand and its spatial position and orientation can be used to simulate the interaction with a virtual object. Similarly, the six-degree-of-freedom handheld controller has its translation and orientation tracking, so that it can be recorded in the virtual space and used to create an interaction with a virtual object at its display position.

[0007] Indirect interaction methods can use variants of traditional input devices (such as touchpads, mice, or keyboards). Users can interact with these devices to control the movement of an indicator in virtual space and interact with virtual elements. For example, an interaction system using a mouse or touchpad can record the movement of a user's hand or fingers on a table or touchpad surface, such that a virtual pointer moves according to the recorded movement and is used to apply an action to a virtual element. The movement of the virtual pointer can be scaled, and its position can be offset relative to the position of the hand or fingers, such that the user can navigate the pointer throughout the virtual space without running out of the space within the interaction area of the input device. Additionally, sometimes the input device only records relative movement, such that the user can further extend the movement range of the virtual pointer by making multiple relative movements on the input device.

[0008] While direct interaction using, for example, hand tracking or a handheld controller provides a realistic experience and a high degree of freedom in controlling virtual objects, the need to constantly hold the hand in the air and move the hand around in the physical world can quickly introduce fatigue and prevent long-term use. Additionally, it is challenging to achieve reliable and high-fidelity spatial tracking of the hand or body with relatively low-cost sensors and a streamlined user setup experience. Indirect interaction using, for example, a keyboard, mouse, or touchpad system provides a familiar experience from the personal computer / smartphone era and allows users to easily interact with virtual content, but it reduces the immersion created by an augmented reality or virtual reality system. Summary of the Invention

[0009] Methods and systems are disclosed that allow a user to view and directly interact with three-dimensional (3D) and two-dimensional (2D) virtual content that accurately overlays the interaction area of a touch input device. Such systems maintain the immersion of an augmented reality or virtual reality system by allowing the user to directly interact with virtual objects at their display location and interact with the recorded precise input movements. At the same time, the integration of a touch input system provides a physical surface for interaction and hand rest. Thus, precise, responsive, reliable tracking of input movements and an effortless interaction experience can be achieved.

[0010] In one aspect of the present disclosure, a system for a human to interact with a virtual object includes: a touch-sensitive surface configured to detect a position of a contact made on the touch-sensitive surface; a reference layer rigidly attached to the touch-sensitive surface and including one or more patterns; a display device configured to display a virtual object recorded in reference coordinates fixed relative to the touch-sensitive surface; one or more image sensors rigidly attached to the display device and configured to capture an image of at least a portion of the one or more patterns; and at least one processor configured to determine a position and orientation of the display device relative to the touch-sensitive surface based on the captured image and to identify an interaction with the virtual object based on the detected position of the contact made on the touch-sensitive surface.

[0011] In one aspect of the present disclosure, the virtual object can be a three-dimensional virtual object. In another aspect of the present disclosure, the virtual object can be a two-dimensional virtual element. In one aspect of the present disclosure, the display device is a see-through display device.

[0012] In one aspect of the present disclosure, the one or more patterns include one or more fiducial markers. In one aspect of the present disclosure, the one or more fiducial markers are configured to absorb infrared light, and the one or more image sensors are configured to sense infrared light. In one aspect of the present disclosure, each of the one or more fiducial markers includes a rectangle that includes an internal grid representation of a binary code. In one aspect of the present disclosure, each of the one or more fiducial markers includes a plurality of image features having known positions, where each image feature corresponds to a unique feature descriptor.

[0013] In one aspect of the present disclosure, the one or more patterns include a plurality of light sources having known positions. In one aspect of the present disclosure, the plurality of light sources are infrared light sources, and the one or more image sensors are configured to sense infrared light. In one aspect of the present disclosure, the plurality of light sources are configured to be turned on in a predetermined order.

[0014] In one aspect of the present disclosure, the one or more patterns include a mask and one or more light sources, where at least a portion of the light emitted from the one or more light sources and passing through the mask is captured by the one or more image sensors. In one aspect of the present disclosure, the one or more patterns further include a diffuser configured to diffuse the light emitted from the one or more light sources. In one aspect of the present disclosure, the one or more patterns further include a light guide plate configured to receive light emitted from the one or more light sources from at least one side of the light guide plate and to guide at least a portion of the light to the mask above the light guide plate.

[0015] In one aspect of the present disclosure, the touch-sensitive surface is at least partially transparent, and the reference layer is disposed beneath the touch-sensitive surface. In one aspect of the present disclosure, the reference layer is disposed adjacent to at least one side of the touch-sensitive surface. In one aspect of the present disclosure, the reference layer is disposed above the touch-sensitive surface.

[0016] In one aspect of the present disclosure, at least one processor is configured to: identify an interaction with a virtual object when a position of the detected contact matches a position of the virtual object. In one aspect of the present disclosure, the virtual object is lifted from the touch-sensitive surface, and at least one processor is configured to: identify an interaction with the virtual object when a position of the detected contact matches a position of a virtual footprint projected from the virtual object onto the touch-sensitive surface. In one aspect of the present disclosure, when the interaction is identified, the display device displays a virtual two-dimensional menu approximating the virtual footprint.

[0017] In one aspect of the present disclosure, a system for human interaction with a virtual object includes: a touch-sensitive surface configured to detect a position of a contact made on the touch-sensitive surface; a display device configured to display the virtual object; one or more ultrasonic transmitters rigidly attached to one of the touch-sensitive surface and the display device and configured to transmit ultrasonic signals; one or more ultrasonic receivers rigidly attached to the other of the touch-sensitive surface and the display device and configured to receive the ultrasonic signals transmitted by the ultrasonic transmitters; and at least one processor configured to determine a position and orientation of the display device relative to the touch-sensitive surface based at least on a time of flight of the received ultrasonic signals and to identify an interaction with the virtual object based on the detected position of the contact.

[0018] In one aspect of the present disclosure, the system further includes an inertial measurement unit rigidly attached to the touch-sensitive surface and / or an inertial measurement unit rigidly attached to the display device.

[0019] In one aspect of the present disclosure, a method for human interaction with a virtual object includes: detecting a position of a contact made on a touch-sensitive surface; displaying, using a display device, a virtual object recorded in reference coordinates fixed relative to the touch-sensitive surface; capturing, using one or more image sensors rigidly attached to the display device, an image of at least a portion of one or more patterns on a reference layer rigidly attached to the touch-sensitive surface; determining a position and orientation of the display device relative to the touch-sensitive surface based on the captured image; and identifying an interaction with the virtual object based on the detected position of the contact.

[0020] In one aspect of the present disclosure, one or more patterns include one or more fiducial marks. In one aspect of the present disclosure, one or more patterns include a plurality of light sources with known positions. In one aspect of the present disclosure, the one or more patterns include a mask and one or more light sources, wherein at least a portion of the light emitted from the one or more light sources and passing through the mask is captured by the one or more image sensors. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 FIG. 6 shows a structural diagram of an exemplary user interaction system according to one aspect of the present disclosure.

[0022] Figure 2 FIG. 10 shows a structural diagram of a first exemplary touch input device according to one aspect of the present disclosure.

[0023] Figure 3 FIG. 14 shows a structural diagram of a second exemplary touch input device according to one aspect of the present disclosure.

[0024] Figure 4 FIG. 18 shows a structural diagram of a third exemplary touch input device according to one aspect of the present disclosure.

[0025] Figure 5A FIG. 22 shows a first exemplary fiducial mark according to one aspect of the present disclosure.

[0026] Figure 5B FIG. 26 shows a second exemplary fiducial mark according to one aspect of the present disclosure.

[0027] Figure 5C FIG. 30 shows a third exemplary fiducial mark according to one aspect of the present disclosure.

[0028] Figure 6 FIG. 34 shows a cross-sectional view of an exemplary reference layer.

[0029] Figure 7 FIG. 38 shows a cross-sectional view of another exemplary reference layer.

[0030] Figure 8 FIG. 42 shows a cross-sectional view of a third exemplary reference layer. DETAILED DESCRIPTION

[0031] Figure 1FIG. 0 shows a structural diagram of an exemplary system 100 that allows a user to interact with a virtual object according to an aspect of the present disclosure. The system 100 may include: a touch input device 101, which includes a reference layer 102 and an interaction surface 103; a display device 105; a pose tracking device 106; and a computing unit 107, which is operably connected to the touch input device 101, the display device 105, and the pose tracking device 106 wirelessly or via one or more cables. In some embodiments, the touch input device 101 may include a touchpad that includes a haptic sensor, wherein the interaction surface 103 may be an active surface of the haptic sensor that detects contact of a finger or an object (e.g., a stylus) with the surface. In some embodiments, the reference layer 102 may include a predetermined set of fiducial patterns. In some embodiments, the reference layer 102 may include a plurality of light sources, such as light-emitting diodes (LEDs), arranged in a predetermined pattern. In some embodiments, an exemplary virtual object 108 is recorded in a reference coordinate 104 fixed relative to the touch input device 101, and the relative position and orientation between the reference coordinate 104 and the interaction surface 103 are known, such that the position and movement of the detected touch input can be directly mapped to a virtual space. In some embodiments, the pose tracking device 106 may include a single or multiple image sensors rigidly attached to the display device 105. Multiple image sensors may be advantageous in cases where a wider field of view is required or where an object such as the user's hand obscures a portion of the field of view. In some embodiments, each image sensor may further include a filter that allows light having a predetermined wavelength range to pass through while attenuating the intensity of light having other wavelengths.

[0032] In some embodiments, the display device 105 may be a see-through display device through which a viewer can perceive computer-generated virtual content as well as the real world. As a result, the system 100 can be used in augmented reality applications. In some embodiments, the display device 105 may be opaque such that it can block light from the real world and display only computer-generated virtual content. As a result, the system 100 can be used in virtual reality applications. In some embodiments, the display device 105 may be a head-mounted device that is placed in front of a viewer's eyes 109 during use.

[0033] In some embodiments, the interaction surface 103 is capable of detecting and reporting the exact position of a touch event on the surface, where the touch event may be generated by contact between an object and the surface, and the object may be the user's finger or hand, or a stylus, etc. In some embodiments, the touch input device 101 may also detect and report the shape of the contact area of the touch event. In some embodiments, the touch input device 101 may further detect and report the force distribution on the contact area of the touch event.

[0034] In some embodiments, the reference layer 102 can be sensed by the pose tracking device 106 to determine the position and orientation of the touch input device 101 relative to the pose tracking device 106. In some embodiments, the reference layer 102 can be a fiducial pattern layer, which can include a predetermined set of points, lines, or shapes. In some embodiments, the reference layer 102 can include a layer of light-emitting diodes arranged in a predetermined pattern.

[0035] In some embodiments, the interaction surface 103 can include a haptic sensor that precisely measures the position of contact between the sensor and a finger or object. In some embodiments, the interaction surface 103 can be fully transparent or translucent. As a result, the reference layer 102 can be disposed beneath the interaction surface 103 while still being able to be sensed by the pose tracking device 106. In some embodiments, the interaction surface 103 can be opaque. As a result, the reference layer 102 can be disposed on top of the interaction surface 103 or attached to one or more sides of the interaction surface 103. In some embodiments, the haptic sensor can further measure the area and / or force distribution of the contact. In some embodiments, the haptic sensors described above can be resistive sensing or capacitive sensing. The haptic sensor can be any type that those skilled in the art consider suitable for performing the functions described herein.

[0036] As described above, the touch input device 101 can include the reference layer 102. In some embodiments, the reference layer 102 can include a predetermined set of fiducial patterns, where the fiducial patterns include a predetermined combination of features, the features including shapes, lines, and points, where the dimensions, positions, or orientations of these features are known. As a result, when a portion or all of the fiducial pattern is captured by one or more image sensors, the position and orientation of the pattern can be determined. In some embodiments, for example, the fiducial pattern can be printed or etched on a support substrate layer with a material that absorbs visible light and / or infrared light. In some embodiments, the fiducial pattern can be created by applying an opaque mask above a diffused illumination source and partially removing it. The fiducial pattern can be fabricated in many forms that those skilled in the art consider suitable for performing the functions described herein.

[0037] Figure 5, Figure 6 and Figure 7 illustrate non-limiting examples of fiducial markers that can be used with embodiments of the present disclosure.

[0038] In some embodiments, the reference layer may include a plurality of light sources, such as light-emitting diodes (LEDs), arranged in a predetermined pattern, where the position of each LED is known. As a result, when a part or all of the LEDs are captured by one or more image sensors, the position and orientation of the pattern can be determined. In some embodiments, the LEDs are sequentially lit such that: in each frame, only one or a few LEDs are captured by the image sensors. Because if all the LEDs are lit simultaneously, multiple LEDs may share similar characteristics in the captured image frame, and this may cause uncertainty problems when establishing a correspondence between each observed LED and each known position. Therefore, the uncertainty problem can be solved by turning on the LEDs in a predetermined order. At the same time, since not all the LEDs need to be on all the time for the pose tracking device to determine the position and orientation, sequentially lighting the LEDs can save valuable battery power.

[0039] In some embodiments, the pose tracking device 106 may include a single or multiple image sensors. In some embodiments, each image sensor may further include a filter that allows only light with a predetermined wavelength range (e.g., infrared light) to pass through while attenuating the intensity of light with other wavelengths (e.g., visible light), where the predetermined range may depend on the wavelength of the light reflected or emitted by the above-mentioned pattern. As a result, the pattern can be clearly captured by the image sensor, while other features in the field of view of the sensor may be partially or completely invisible to the sensor. In some embodiments, the pose tracking device 106 may further include an illumination device, where the illumination device may include a single or multiple light-emitting diodes.

[0040] The computing unit 107 may include one or more processors. Although in Figure 1In the example shown, the computing unit 107 is incorporated into a head-mounted device that includes a pose tracking device 106 and a display device 105. However, the computing unit 107 can alternatively be arranged in a separate computer or incorporated into a touch input device 101. The computing unit 107 can generate a virtual space that includes virtual content, where the virtual content can include three-dimensional virtual objects and / or two-dimensional virtual elements. The computing unit 107 can also generate information about the virtual content, where the information can include spatial position, orientation, shape, kinematics, dynamics, and graphical rendering characteristics, etc. When a touch event is detected, the touch position reported by the touch input device 101 and the position of the virtual object overlying the interaction surface 103 will allow the computing unit 107 to determine which virtual object is being interacted with. Additionally, the type of touch event and the characteristics of the virtual object can define different interactions and responses. Meanwhile, the pose tracking device 106 provides the necessary measurement data to the computing unit 107, enabling the determination of the relative spatial relationship between the display device 105 and the touch input device 101. The relative spatial relationship generally refers to the relative translation and rotation between two entities. Thus, the virtual objects and the responses triggered by user interactions can be appropriately rendered and displayed from the viewer's perspective.

[0041] For example, the virtual object can be magnified while interacting with it, and optionally, text information associated with the specific object can be displayed. As another example, the virtual object can be dragged from a first position to a second position through touch interaction.

[0042] In one embodiment, the fiducial pattern can include multiple square-based fiducial markers, each fiducial marker containing an outer boundary and an internal grid representation of a binary code. An example of such a fiducial marker is shown in Figure 5A When captured by an image sensor, the fiducial markers in the captured frame are extracted by discarding any shapes that are not 4-vertex polygons, do not have a valid binary code, or do not meet the size threshold. The corners of each valid fiducial marker are detected by a corner detection algorithm. Using the known positions of the corners of the fiducial marker, the relative position and orientation of the marker with respect to the image sensor can be estimated, for example, by solving the perspective-n-point problem. Thus, fiducial pattern-based pose estimation is achieved.

[0043] In another embodiment, the fiducial pattern may include a predefined image target that contains a plurality of features having known positions and known descriptors in the image target. Features in computer vision or image processing are distinct local structures found in an image, such as "edges" (a set of points in the image with strong gradient magnitudes), "corners / interest points" (a set of points where the direction of the gradient changes rapidly within a local region), or local image patches. The descriptor encodes the characteristics of the feature, e.g., the magnitude and orientation of the local gradient of pixel intensities, a vector of intensity comparisons between pairs of pixels around the feature. The descriptor can be in various forms, including numerical values, numerical vectors, or vectors of boolean variables. The descriptor can be used to uniquely identify the corresponding feature in the image. For example, in the case of the BRIEF (Binary Robust Independent Elementary Features) descriptor, the Hamming distance between the known descriptor and the descriptor of the candidate feature is calculated, and if the distance is less than a threshold, a match is confirmed.

[0044] When the image target is captured by the image sensor, all candidate feature points within the captured frame are extracted using a feature detection algorithm, and corresponding descriptors are calculated for the candidate feature points. By comparing the descriptors, some of the candidate feature points are matched with the known feature points in the image target. For example, by solving the Perspective-n-Point problem, the matched pairs are used to estimate the relative position and orientation of the image target with respect to the image sensor. Thus, pose estimation based on the fiducial pattern is achieved.

[0045] Figure 2 is a structural diagram of a first exemplary touch input device. In Figure 2 the illustrated embodiment, the interaction layer 202 is disposed on top of the reference layer 201, where the interaction layer 202 is fully or partially transparent. The interaction layer 202 may include the tactile sensors as described above. The reference layer 201 may include the predefined set of fiducial patterns as described above. Alternatively, the reference layer 201 may include a plurality of light-emitting diodes (LEDs) arranged in a predefined pattern as described above.

[0046] Figure 3 is a structural diagram of a second exemplary touch input device. In Figure 3 the illustrated embodiment, the reference layer 301 is rigidly attached to one side of the interaction layer 302, where the interaction layer 302 may be opaque. The interaction layer 302 may include the tactile sensors as described above. The reference layer 301 may include the predefined set of fiducial patterns as described above. Alternatively, the reference layer may include a plurality of light-emitting diodes (LEDs) arranged in a predefined pattern as described above.

[0047] As an optional example, as Figure 1 shown, the reference layer 102 may be arranged around the four sides of the rectangular interaction surface 103.

[0048] In some embodiments, the touch input device may include a plurality of ultrasonic transmitters placed in a predetermined pattern, and the pose tracking device may include a plurality of ultrasonic receivers. In some embodiments, the pose tracking device may include a plurality of ultrasonic transmitters placed in a predetermined pattern, and the touch input device may include a plurality of ultrasonic receivers. As a result, the distance between the transmitter and the receiver can be determined by measuring the propagation time of the ultrasonic signal. Therefore, the position and orientation of the touch input device can be determined.

[0049] Figure 4 is a structural diagram of a third exemplary touch input device employing ultrasonic transmitters or receivers. In Figure 4 the illustrated embodiment, the support frame 402 is attached to four sides of the interaction layer 401, and four ultrasonic transmitters 403, 404, 405, 406 are disposed on top of the support frame 402 and distributed at four corners. The interaction layer 402 may include the tactile sensors as described above. Alternatively, four ultrasonic receivers 403, 404, 405, 406 may be disposed on top of the support frame 402.

[0050] Various methods can be used to track the position and orientation of a virtual object using ultrasonic receivers and transmitters. For example, in one implementation, three ultrasonic receivers are rigidly attached to one of the display device and the touch input device in a non - collinear arrangement, three transmitters are rigidly attached to the other of the display device and the touch input device, and a computing unit is coupled to the transmitters and receivers. The three transmitters respectively generate three ultrasonic pulses of different frequencies. Each of the three receivers separates the received ultrasonic waves with three different frequencies into three signals, resulting in a total of nine signals. Based on the time - of - flight principle, the nine signals are processed into nine distances between each of the three transmitters and each of the three receivers. Therefore, the relative orientation and position between the transmitter assembly and the receiver assembly can be estimated.

[0051] In another embodiment, an ultrasonic transmitter and a 9-axis inertial measurement unit (IMU) are rigidly attached to one of a display device and a touch input device, and three ultrasonic receivers and a 9-axis IMU are rigidly attached to the other of the display device and the touch input device, wherein the receivers are arranged in a non-collinear arrangement, and a computing unit is coupled to the transmitter, the receivers, and the IMU. Alternatively, three transmitters and one receiver may be used. The transmitter generates ultrasonic sound pulses of a known frequency, while the receivers convert the received ultrasonic pulses into three signals. Based on the time-of-flight principle, the signals respectively produce three distances between the transmitter and the three receivers. As a result, the relative position between the transmitter and the receiver assembly can be calculated. These IMUs measure the absolute orientation of the display device and the touch input device, such that the relative orientation between them can be determined.

[0052] In some embodiments, footprints may be displayed for the lifted virtual objects. Since user interaction is sensed through the touch input device, the interaction is restricted to near the 2D plane. However, the virtual content may be displayed at a non-negligible vertical distance above the touch input device. To overcome this limitation, in one embodiment, virtual footprints projected from the lifted virtual objects onto the interaction layer are displayed through the display device. As a result, the user can interact with the lifted virtual objects via their virtual footprints using various touch gestures. For example, the user can perform a pinch gesture on the touch input device over the area of the virtual footprint to scale the corresponding virtual object, or the user can perform a press-and-drag gesture on the virtual footprint to move the corresponding virtual object. Additionally, when the user touches the area of the virtual footprint, virtual two-dimensional menu elements may be displayed on the interaction layer near the virtual footprint to provide additional operations for the corresponding virtual object. The user can click on different areas of the interaction layer displaying the menu items to activate the relevant functions. For example, the additional operations may include, but are not limited to: starting an animation associated with the corresponding virtual object, deleting the corresponding virtual object, or changing the properties of the corresponding virtual object.

[0053] Figure 6 is a cross-sectional view of an exemplary reference layer. In Figure 6 the illustrated embodiment, a mask layer 603 is disposed above a light source layer 601, and an optional light diffuser 602 may be located between the light source layer 601 and the mask layer 603. The mask layer 603 includes light-transmitting portions illustrated as bright color blocks and light-blocking portions illustrated as dark color blocks. The light emitted by the light sources in the light source layer 601 may be diffused by the light diffuser 602 and pass through the light-transmitting portions of the mask layer 603 to form one or more of the above patterns when captured by the image sensor. In some embodiments, the light from the light source may be infrared light, and the light-transmitting portions and the light-blocking portions respond to the infrared wavelength.

[0054] In some embodiments, the light-shielding portion may comprise a polymer having a light-shielding additive. In some embodiments, the light-shielding portion may comprise a light-shielding coating deposited on a substrate. In some embodiments, the light-transmitting portion is simply formed by voids or lacks any material. The material of the mask layer 603 may be any type that those skilled in the art deem suitable for performing the functions described herein.

[0055] In some embodiments, the interaction surface is located above the mask layer 603. In some embodiments, the interaction surface is located below the mask layer 603. In some embodiments, the light-shielding material is directly deposited on the interaction surface to form the mask layer 603.

[0056] Figure 7 is a cross-sectional view of another exemplary reference layer. In Figure 7 the illustrated embodiment, the mask layer 703 is disposed above the light guide plate 701 and may have the same configuration as the mask layer 603 described above. The light source 702 is arranged to laterally emit light from at least one side of the light guide plate to the light guide plate 701. The light emitted from the light source 702 is redirected by the light guide plate 701 to the mask layer 703 and passes through the mask layer 703 to form one or more of the above-described patterns when captured by an image sensor. In some embodiments, the light guide plate 701 may comprise a microlens array on the bottom surface, and the light from the light source 702 is internally reflected along the length of the light guide plate 701 (e.g., from right to left in Figure 7 ), and redirected by the microlens array to pass through the top surface of the plate. The light guide plate 701 may be any type that those skilled in the art deem suitable for performing the functions described herein. In some embodiments, the light from the light source may be infrared light, and the light-transmitting and light-shielding portions of the mask layer 703 respond to the infrared wavelength.

[0057] In some embodiments, the interaction surface is located above the mask layer 703. In some embodiments, the interaction surface is located below the mask layer 703. In some embodiments, the light-shielding material is directly deposited on the touch-sensitive surface to form the mask layer 703.

[0058] Figure 8 is a cross-sectional view of a third exemplary reference layer. In Figure 8 the illustrated embodiment, the light guide plate 801 is located above a reference layer 803 that includes one or more reference patterns. The light source 802 is arranged to laterally emit light from at least one side of the light guide plate 801 to the light guide plate 801. The light emitted by the light source is redirected by the leading light guide plate to illuminate the reference layer 803 and form one or more of the above-described reference patterns in the image frame when captured by an image sensor. In some embodiments, the light guide plate 801 may comprise a microlens array on the top surface, and, the light from the light source is internally reflected along the length of the light guide plate 801 (e.g., in Figure 8from right to left), and redirected by the microlens array to pass through the bottom surface of the plate. The light guide plate 801 can be of any type that those skilled in the art deem suitable for performing the functions described herein. In some embodiments, the light from the light source can be infrared light, and the fiducial pattern responds to the infrared wavelength.

[0059] In some embodiments, the interaction surface is located above the light guide plate 801. In some embodiments, the interaction surface is located below the fiducial layer 803. In some embodiments, the interaction surface is located between the light guide plate 801 and the fiducial layer 803.

[0060] For purposes of illustration, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention and various embodiments with various modifications suited to the particular use contemplated.

Claims

1. A system for human - virtual object interaction, comprising: A touch - sensitive surface configured to detect the position of a contact made on the touch - sensitive surface; A reference layer rigidly attached to at least one side of the touch - sensitive surface, wherein, The reference layer includes: A light - guide plate; A fiducial layer including one or more patterns; and A light source positioned such that light emitted from the light source enters the light - guide plate through a lateral side of the light - guide plate, is reflected internally along the length of the light - guide plate, and is redirected to illuminate the fiducial layer and form the one or more patterns in an image frame; A wearable display device configured to display a virtual object recorded in reference coordinates fixed relative to the touch - sensitive surface; One or more image sensors rigidly attached to the display device and configured to capture an image of at least a portion of the one or more patterns in the reference layer; and At least one processor configured to determine the position and orientation of the display device relative to the touch - sensitive surface at least in part based on image data associated with one or more patterns found in the captured image, Obtain touch data from the touch - sensitive surface, the touch data including the position of the detected contact made on the touch - sensitive surface; Cause the virtual object to be displayed by the wearable display device; Generate an operation on the virtual object based on the position of the detected contact made on the touch - sensitive surface; Cause the operation on the virtual object to be displayed by the wearable display device.

2. The system according to claim 1, wherein The one or more patterns include transmissive portions and light - blocking portions.

3. The system according to claim 1, wherein The one or more patterns include light - absorbing portions.

4. The system according to claim 1, wherein The virtual object is a two - dimensional virtual object.

5. The system according to claim 1, wherein The virtual object is a three - dimensional virtual object.

6. The system according to claim 1, wherein The display device is a see - through display device.

7. The system according to claim 1, wherein, The one or more patterns include one or more fiducial marks.

8. The system according to claim 7, wherein The one or more fiducial marks are configured to absorb infrared light, and the one or more image sensors are configured to sense infrared light.

9. The system according to claim 7, wherein Each of the one or more fiducial marks includes a rectangle containing an internal grid representation of a binary code.

10. The system according to claim 7, wherein, Each of the one or more fiducial marks includes a plurality of image features with known positions, wherein each of the image features corresponds to a unique feature descriptor.

11. The system according to claim 1, wherein, A planar light - source layer is arranged below the fiducial layer, and the transmissive portions and light - blocking portions of the fiducial layer are configured as masks.

12. The system according to claim 1, wherein, A planar light - source layer is arranged between the touch - sensitive surface and the fiducial layer.

13. A method for human - virtual object interaction, comprising: Detect the position of a contact made on a touch - sensitive surface; Use a display device to display a virtual object recorded in reference coordinates fixed relative to the touch - sensitive surface; Use a light source to emit light into a lateral side of a light - guide plate to illuminate one or more patterns of a fiducial layer of a reference layer rigidly attached to at least one side of the touch - sensitive surface; Capturing an image of at least a portion of the one or more patterns using one or more image sensors rigidly attached to the display device; Determining a position and an orientation of the display device relative to the touch-sensitive surface based on the captured image; Obtaining touch data from the touch-sensitive surface, the touch data including positions of detected contacts made on the touch-sensitive surface; Causing a virtual object to be displayed by the wearable display device; Generating an operation on the virtual object based on the positions of the detected contacts made on the touch-sensitive surface; Causing the operation on the virtual object to be displayed by the wearable display device.

14. The method according to claim 13, wherein The one or more patterns of the reference layer include light-transmissive portions and light-blocking portions, and the method further includes: Causing light to pass through the light-transmissive portions using the one or more patterns of the reference layer; and Causing light to be blocked by the light-blocking portions using the one or more patterns of the reference layer.

15. The method according to claim 13, wherein, The one or more patterns of the reference layer include light-transmissive portions, and the method further includes: causing light to be absorbed by light-blocking portions using the one or more patterns of the reference layer.

16. The method according to claim 13, wherein, The one or more patterns include one or more fiducial marks.

17. The method according to claim 13, wherein The one or more patterns include a plurality of light sources having known positions.

18. The method according to claim 13, wherein, The one or more patterns include a mask and a light source, wherein at least a portion of the light emitted from the light source and passing through the mask is captured by the one or more image sensors.

Citation Information

Patent Citations

  • System and method for facilitating interaction with virtual space via touch sensitive surface

    CN103677620A

  • Systems and methods for viewport-based augmented reality haptic effects

    CN105094311A