Methods of tracking the location of a device
By installing image sensors and displaying markers on head-mounted devices, the problem that existing systems cannot accurately track the location of independent mobile devices is solved, and virtual representations of these devices are realized in a computer-generated real environment, improving the user interaction experience.
Patent Information
- Application Number
- CN202210304288.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-11
- Filing Date
- 2019-09-12
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2039-09-12
AI Technical Summary
Existing systems fail to accurately, consistently and effectively track the location of independent mobile devices relative to content-providing devices, resulting in the inability to display representations of these devices in a computer-generated real-world environment, affecting user interaction with electronic content.
By installing an image sensor and a display on a first device (such as a head mounted device) and displaying markers on a second device (such as a mobile phone), determining the relative position and orientation of the second device using the image sensor, generating a control signal for use as a 3D controller or pointer, and displaying a virtual representation of the second device on the display of the first device.
Accurate location tracking and virtual representation of independent mobile devices are realized in computer-generated real-life environments, improving the user's interactive experience with electronic content.
Smart Images

Figure CN114647318B_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese invention patent application with the application date of September 12, 2019, application number 201910866381.4, and invention name “Method for tracking the location of a device”.
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims the benefit of U.S. Provisional Application Serial No. 62 / 731,285, filed September 14, 2018, which is incorporated herein by reference in its entirety. Technical Field
[0004] The present disclosure generally relates to electronic devices for providing and interacting with content, and particularly to systems, methods, and devices that track relative positions of electronic devices and use such positions to provide interactivity, for example, with a computer generated reality (CGR) environment. Background Art
[0005] In order to enable users to interact with electronic content, it may be advantageous to enable users to provide input via independent real-world devices, such as the touch screen of independent mobile devices. However, existing systems are unable to adequately track the position of such independent devices relative to the content providing device, and therefore are unable to display such separate devices or their representations to guide the user's interaction. For example, a user wearing a head mounted device (HMD) presenting a CGR environment will need to see a representation of his mobile phone in the CGR environment in order to use the touch screen of the mobile phone as an input device. However, without being able to accurately, consistently and effectively track the relative position of the mobile phone with respect to the HMD, the representation of the mobile phone cannot be displayed in the CGR environment at a position corresponding to the real-world position. Summary of the invention
[0006] Various embodiments disclosed herein include devices, systems, and methods for providing an improved user interface for interacting with electronic content using multiple electronic devices. Some embodiments relate to a first device (e.g., a head mounted device (HMD)) having an image sensor (e.g., a camera) and one or more displays, and a second device (e.g., a mobile phone) having a display. A marker is displayed on the display of the second device, and the first device determines the relative position and orientation of the second device relative to the first device based on the marker. In some embodiments, the marker is an image containing texture / information that allows the image to be detected and makes it possible to determine the posture of the image relative to the camera. In some embodiments, a control signal is generated based on the relative position and orientation of the second device (e.g., the first device uses the position and orientation of the second device to enable the second device to be used as a three-dimensional (3D) controller, a 3D pointer, a user interface input device, etc.). In some embodiments, based on the determined relative position of the second device, a representation of the second device including virtual content replacing the marker is displayed on the display of the first device.
[0007] In some implementations, the second device has a touch screen, and in some implementations, the virtual content positioned in place of the marker includes controls (e.g., buttons) corresponding to interactions with the user experience / content provided by the first device. For example, the first device may display a CGR environment that includes a virtual remote control having virtual buttons that are representations of a mobile phone. The virtual remote control is displayed at a location corresponding to the real-world location of the mobile phone. When a user virtually touches a virtual button on the virtual remote control, the user actually touches a corresponding portion of the touch screen of the second device, which is recognized as an input to control or otherwise initiate interaction with the virtual environment.
[0008] In some implementations, the relative position and orientation of the second device relative to the first device is adjusted over time based on motion tracking of the first device and the second device over time, for example, inertial measurement unit (IMU) data from an inertial measurement unit (IMU) sensor of the first device or the second device. In addition, in some implementations, the relative position and orientation of the second device relative to the first device is adjusted over time based on additional images depicting markers.
[0009] In some implementations, an estimated error associated with the relative position and orientation of the second device relative to the first device is detected to be greater than a threshold (e.g., drift). Based on detecting that the estimated error is greater than the threshold, an additional image including a marker is obtained. The relative position and orientation of the second device relative to the first device is adjusted over time based on the additional image. In some implementations, the marker is determined based on attributes of the physical environment (e.g., lighting conditions). In some implementations, the marker in the additional image is adaptive (e.g., based on changes in conditions over time). In some implementations, the marker is positioned on only a portion of the second display based on detecting an obstruction between the image sensor and the display. In addition, the marker can be positioned on a portion of the second display based on detecting a touch event on the second display (e.g., a user's finger blocking another portion of the display).
[0010] In some implementations, a light source on a second device (e.g., an array of LEDs, a pixel-based display, a visible light source, an infrared light source (IR) that produces light that is generally invisible to humans, etc.) generates light at a given moment that is encoded in light and can be used to synchronize motion data generated via the second device (e.g., accelerometer data, IMU data, etc.) with processing performed by a first device. In some implementations, a method involves using an image sensor of a first device to obtain an image of a physical environment. The image includes a description of a second device. The description of the second device includes a description of a light-based indicator provided via a light source on the second device.
[0011] The method synchronizes motion data generated via a second device and processing (e.g., interpretation of an image) performed by a first device based on a description of a light-based indicator. For example, the light-based indicator of the second device can be a plurality of LEDs that produce a binary pattern of light that encodes current motion data generated at the second device. As another example, such LEDs can produce a binary pattern of light that encodes time data associated with the generation of motion data via the second device, for example, the time at which a motion sensor on the device captures the data relative to the time at which the binary pattern is provided. In other specific implementations, the second device includes a pixel-based display that displays a pattern that encodes motion or time data of the second device. In other specific implementations, the device includes an IR light source that produces a pattern of IR light that encodes information, such as motion data generated at the second device. The first device can synchronize the motion data of the second device with positioning data determined by computer vision processing of the image, for example, associating the current motion of the second device provided in the light-based indicator with the current relative position of the second device determined via computer vision.
[0012] The method may generate a control signal based on synchronization of the motion data with the image. For example, if the motion data of the second device is associated with movement of the second device intended to move an associated cursor displayed on the first device, the method may generate an appropriate signal to cause such movement of the cursor.
[0013] According to some specific implementations, a device includes one or more processors, non-volatile memory, and one or more programs; the one or more programs are stored in the non-volatile memory and are configured to be executed by one or more processors, and the one or more programs include instructions for performing or causing the execution of any of the methods described herein. According to some specific implementations, a non-volatile computer-readable storage medium stores instructions that, when executed by one or more processors of the device, cause the device to perform or cause the execution of any of the methods described herein. According to some specific implementations, a device includes: one or more processors, non-volatile memory, and a device for performing or causing the execution of any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] So that the present disclosure can be understood by those of ordinary skill in the art, a more detailed description may be obtained with reference to aspects of some exemplary embodiments, some of which are illustrated in the accompanying drawings.
[0015] Figure 1 is a block diagram of an exemplary physical environment including a user, a first device, and a second device according to some implementations.
[0016] Figure 2 According to some specific implementations, a display including a first device is shown. Figure 1 An exemplary physical environment.
[0017] Figure 3A According to some specific implementations Figure 1 The pattern displayed by the second device.
[0018] Figure 3B According to some specific implementations Figure 1 A virtual representation of a second device.
[0019] Figure 4 is a block diagram of an exemplary first device according to some specific implementations.
[0020] Figure 5 is a block diagram of an exemplary second device according to some specific implementations.
[0021] Figure 6 is a block diagram of an exemplary head mounted device (HMD) according to some implementations.
[0022] Figure 7is a flowchart representation of a method for providing an improved user interface for interacting with a virtual environment, according to some implementations.
[0023] Figure 8 is a flowchart representation of a method for providing an improved user interface for interacting with a virtual environment, according to some implementations.
[0024] Fig. 9 is a flow chart representation of a method for tracking the position of a device using a light-based indicator to encode device motion or synchronization data.
[0025] As is common practice, the various features shown in the drawings may not be drawn to scale. Therefore, the sizes of the various features may be arbitrarily expanded or reduced for clarity. In addition, some drawings may not depict all components of a given system, method, or device. Finally, throughout the specification and drawings, similar reference numerals may be used to represent similar features. DETAILED DESCRIPTION
[0026] Many details are described in order to provide a thorough understanding of the exemplary implementations shown in the accompanying drawings. However, the accompanying drawings only illustrate some exemplary aspects of the present disclosure and should not be considered limiting. One of ordinary skill in the art will appreciate that other effective aspects or variations do not include all of the specific details described herein. In addition, well-known systems, methods, components, devices, and circuits are not described in detail in order to avoid obscuring more relevant aspects of the exemplary implementations described herein.
[0027] Figure 1 1 is a block diagram of an exemplary physical environment 100 including a user 110, a physical first device 120, and a physical second device 130. In some embodiments, the physical first device 120 is configured to present content such as a CGR environment to the user 110. A computer generated reality (CGR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic system. In CGR, a subset of a person's physical movements or a representation thereof is tracked, and in response, one or more features of one or more virtual objects simulated in the CGR environment are adjusted in a manner that complies with at least one physical law. For example, a CGR system can detect a person's head rotation, and in response, adjust the graphical content and sound field presented to the person in a manner similar to the way such views and sounds change in a physical environment. In some cases (e.g., for accessibility reasons), the adjustment of one or more features of one or more virtual objects in the CGR environment can be performed in response to a representation of physical movement (e.g., a voice command).
[0028] A person may sense and / or interact with a CGR object using any of their senses, including vision, hearing, touch, taste, and smell. For example, a person may sense and / or interact with an audio object that creates a 3D or spatial audio environment that provides a perception of a point audio source in 3D space. As another example, an audio object may enable audio transparency that selectively introduces ambient sounds from the physical environment with or without computer-generated audio. In some CGR environments, a person may sense and / or interact only with an audio object.
[0029] Examples of CGR include virtual reality and mixed reality. A virtual reality (VR) environment refers to a simulated environment designed to be based entirely on computer-generated sensory input to one or more senses. A VR environment includes virtual objects that a person can sense and / or interact with. For example, trees, buildings, and computer-generated images representing avatars of people are examples of virtual objects. A person can sense and / or interact with virtual objects in a VR environment through a simulation of the person's presence within the computer-generated environment, and / or through a simulation of a subset of the person's physical movements within the computer-generated environment.
[0030] In contrast to VR environments, which are designed to be based entirely on computer-generated sensory input, a mixed reality (MR) environment refers to a simulated environment designed to include sensory input from the physical environment, or representations thereof, in addition to computer-generated sensory input (e.g., virtual objects). On the virtuality continuum, a mixed reality environment is anything between a fully physical environment at one end and a virtual reality environment at the other end, but not including both ends.
[0031] In some MR environments, computer-generated sensory input can respond to changes in sensory input from the physical environment. In addition, some electronic systems used to render the MR environment can track position and / or orientation relative to the physical environment to enable virtual objects to interact with real objects (i.e., physical items from the physical environment or representations thereof). For example, the system can cause motion so that virtual trees appear stationary relative to the physical ground.
[0032] Examples of mixed reality include augmented reality and augmented virtuality. An augmented reality (AR) environment refers to a simulated environment in which one or more virtual objects are superimposed on a physical environment or a representation thereof. For example, an electronic system for presenting an AR environment may have a transparent or translucent display through which a person can directly view the physical environment. The system may be configured to present virtual objects on a transparent or translucent display so that a person uses the system to perceive virtual objects superimposed on a physical environment. Alternatively, the system may have an opaque display and one or more imaging sensors that capture images or videos of a physical environment that are representations of the physical environment. The system combines the image or video with the virtual object and presents the composition on an opaque display. People use the system to indirectly view the physical environment through an image or video of the physical environment and perceive virtual objects superimposed on the physical environment. As used herein, a video of a physical environment displayed on an opaque display is referred to as a "transparent video," meaning that the system uses one or more image sensors to capture images of the physical environment and uses those images when presenting an AR environment on an opaque display. Further alternatively, the system may have a projection system that projects virtual objects into a physical environment, for example as a hologram or on a physical surface, so that a person using the system perceives the virtual objects superimposed on the physical environment.
[0033] An augmented reality environment also refers to a simulated environment in which a representation of a physical environment is transformed by computer-generated sensory information. For example, in providing a pass-through video, the system may transform one or more sensor images to apply a selected perspective (e.g., a viewpoint) that is different from the perspective captured by the imaging sensor. For another example, a representation of a physical environment may be transformed by graphically modifying (e.g., enlarging) a portion thereof so that the modified portion may be a representative but not true version of the original captured image. For another example, a representation of a physical environment may be transformed by graphically eliminating or blurring a portion thereof.
[0034] An augmented virtual (AV) environment refers to a simulated environment in which a virtual or computer-generated environment is combined with one or more sensory inputs from a physical environment. The sensory input may be a representation of one or more features of the physical environment. For example, an AV park may have virtual trees and virtual buildings, but the faces of people are realistically reproduced from images taken of physical people. For another example, a virtual object may take the shape or color of a physical object imaged by one or more imaging sensors. For another example, a virtual object may take a shadow that conforms to the position of the sun in the physical environment.
[0035] There are many different types of electronic systems that enable people to sense and / or interact with various CGR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on people's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smart phones, tablets, and desktop / laptop computers. A head-mounted system can have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system can be configured to accept an external opaque display (e.g., a smart phone). A head-mounted system can incorporate one or more imaging sensors for capturing images or videos of a physical environment, and / or one or more microphones for capturing audio of a physical environment. Instead of an opaque display, a head-mounted system can have a transparent or translucent display. A transparent or translucent display can have a medium through which light representing an image is directed to a person's eyes. The display can utilize digital light projection, OLED, LED, uLED, silicon-based liquid crystal, laser scanning light source, or any combination of these technologies. The medium can be an optical waveguide, a hologram medium, an optical combiner, an optical reflector, or any combination thereof. In one embodiment, a transparent or translucent display can be configured to selectively become opaque. A projection-based system can employ retinal projection technology that projects a graphic image onto a person's retina. The projection system can also be configured to project a virtual object into a physical environment, such as a hologram or onto a physical surface.
[0036] exist Figure 1 , the physical first device 120 is shown as an HMD. Those skilled in the art will recognize that an HMD is only one form factor suitable for implementing the physical first device 120. Other form factors suitable for implementing the physical first device 120 include smartphones, AR glasses, smart glasses, desktop computers, laptops, tablet computers, computing devices, etc. In some specific implementations, the physical first device 120 includes a suitable combination of software, firmware, or hardware. For example, the physical first device 120 may include an image sensor (e.g., image sensor 122) and a display. In some specific implementations, the physical first device 120 includes a display on the inward-facing surface of the physical first device 120.
[0037] In some implementations, multiple cameras are used in the physical first device 120 and the physical second device 130 to capture image data of the physical environment 100. In addition, the image sensor 122 may be located in addition to Figure 1In some implementations, the image sensor 122 includes a high-quality, high-resolution RGB video camera, such as a 10 megapixel (e.g., 3072×3072 pixel count) camera with a frame rate of 60 frames per second (FPS) or higher, a horizontal field of view (HFOV) greater than 90 degrees, and a working distance of 0.1 meters (m) to infinity.
[0038] In some implementations, the image sensor 122 is an infrared (IR) camera with an IR illumination source or a light detection and ranging (LIDAR) transmitter and receiver / detector that, for example, captures depth or range information of objects and surfaces in the physical environment 100. The range information can be used, for example, to position virtual content synthesized into the image of the physical environment 100 at the correct depth. In some implementations, the range information can be used to adjust the depth of real objects in the environment when displayed; for example, nearby objects can be re-rendered to be smaller in the display to help the user 110 avoid the objects when moving around the environment.
[0039] In some implementations, the physical first device 120 and the physical second device 130 are communicatively coupled via one or more wired or wireless communication channels (e.g., BLUETOOTH, IEEE 802.11x, IEEE802.16x, IEEE 802.3x, etc.). Although this example and other examples discussed herein show a single physical first device 120 and a single physical second device 130 in a real-world physical environment 100, the techniques disclosed herein are applicable to multiple devices and other real-world environments. In addition, the functions of the physical first device 120 can be performed by multiple devices, and similarly, the functions of the physical second device 130 can be performed by multiple devices.
[0040] In some implementations, the physical first device 120 is configured to present a CGR environment to the user 110. In some implementations, the physical first device 120 includes a suitable combination of software, firmware, or hardware. In some implementations, the user 110 wears the physical first device 120 on his / her head, such as an HMD. In this way, the physical first device 120 may include one or more displays for displaying images. The physical first device 120 may surround the field of view of the user 110, such as an HMD. Figure 4 and Figure 6 The physical first device 120 is described in more detail.
[0041] In some implementations, the physical first device 120 presents a CGR experience to the user 110 while the user 110 is physically present within the physical environment 100 and virtually present within the CGR environment. In some implementations, while presenting the CGR environment to the user 110, the physical first device 120 is configured to present CGR content and enable optical see-through of at least a portion of the physical environment 100. In some implementations, while presenting the CGR environment, the physical first device 120 is configured to present CGR content and enable video see-through of the physical environment 100.
[0042] In some implementations, the image sensor 122 is configured to obtain image data corresponding to the physical environment (e.g., the physical environment 100) in which the physical first device 120 is located. In some implementations, the image sensor 122 is part of an image sensor array that is configured to capture a light field image corresponding to the physical environment (e.g., the physical environment 100) in which the physical first device 120 is located.
[0043] In some specific implementations, such as Figure 1 As shown, the physical second device 130 is a handheld electronic device (e.g., a smart phone or a tablet computer) that includes a physical display 135. In some implementations, the physical second device 130 is a laptop or desktop computer. In some implementations, the physical second device 130 has a trackpad, and in some implementations, the physical second device 130 has a touch-sensitive display (also referred to as a "touch screen" or "touch screen display").
[0044] In some implementations, the physical second device 130 has a graphical user interface ("GUI"), one or more processors, memory, and one or more modules, a program or instruction set stored in the memory for performing a variety of functions. In some implementations, the user 110 interacts with the GUI through finger contacts and gestures on the touch-sensitive surface. In some implementations, these functions include image editing, drawing, presentation, word processing, web page creation, disk editing, spreadsheet creation, playing games, making and receiving calls, video conferencing, sending and receiving emails, instant messaging, fitness support, digital photography, digital video recording, web browsing, digital music playback, or digital video playback. Executable instructions for performing these functions may be included in a computer-readable storage medium or other computer program product configured for execution by one or more processors.
[0045] In some implementations, presenting content representing virtual content includes identifying a placement location of a virtual object or virtual surface that corresponds to a real-world physical object (e.g., physical second device 130) or a real-world physical surface. In some implementations, the placement location of the virtual object or virtual surface that corresponds to the real-world physical object or the real-world physical surface is based on a spatial location of a physical surface in the physical environment 100 relative to the physical first device 120 or the physical second device 130. In some implementations, the spatial location is determined using an image sensor 122 of the physical first device 120, and in some implementations, the spatial location is determined using an image sensor external to the physical first device 120.
[0046] In some implementations, the physical first device 120 or the physical second device 130 creates and tracks a correspondence between the real world space (e.g., the physical environment 100) in which the user 110 is located and the virtual space including the virtual content. Therefore, the physical first device 120 or the physical second device 130 can use the world and camera coordinate system 102 (e.g., the y-axis points upward, the z-axis points to the user 110, and the x-axis points to the right of the user 110). In some implementations, the configuration can change the origin and orientation of the coordinate system relative to the real world. For example, each device can define its own local coordinate system.
[0047] In some implementations, each device combines information from the motion sensing hardware of the corresponding device with an analysis of the scene visible to the camera of the corresponding device to create a correspondence between real and virtual spaces, for example, through visual inertial odometry (VIO). For example, each device (e.g., physical first device 120 and physical second device 130) can identify significant features (e.g., plane detection) in virtual environment 100, track differences in the positions of those features across video frames, and compare this information with the motion sensing data. For example, by increasing the number of significant features in the scene image, the relative position of physical second device 130 with respect to physical first device 120 can be further accurately determined.
[0048] In some implementations, in order to prevent errors (e.g., drift) caused by small errors in inertial measurement, a tracking system using a fixed reference point is used to determine relative inertial motion. For example, small errors between the inertial measurement of the physical first device 120 and the inertial measurement of the physical second device 130 may accumulate over time. These errors may affect the ability of the physical first device 120 to accurately present the virtual representation of the physical second device 130 to the user, for example, although the physical second device 130 and the user's arm 35 remain in a relatively constant position in the physical environment 100, the virtual representation of the physical second device 130 or the user's arm (e.g., physical appendage 115) may appear to float slowly toward the user 110. For example, these errors may be estimated and compared to a threshold. If the error exceeds the threshold, the physical first device 120 may then use a fixed reference point to determine relative inertial motion.
[0049] In some implementations, the physical first device 120 or the physical second device 130 is configured to use one or more cameras (e.g., image sensor 122) to identify fixed reference points in an image (e.g., fixed points in the physical environment 100) and track the fixed reference points in additional images. For example, when it is determined that the estimated error associated with the position and orientation of the physical second device 130 relative to the physical first device 120 is greater than a threshold, the physical first device 120 may run a positioning algorithm that uses reference points to track movement in space (e.g., simultaneous localization and mapping (SLAM)). In some implementations, the inertial measurement device may perform inertial measurements at a frequency higher than that at which the tracking system performs tracking measurements. Therefore, the inertial measurements from the inertial measurement device may be primarily used by the processing system to determine the movement of the user 110, the physical first device 120, the physical second device 130, or a part of the user's body, and may be corrected at given intervals based on the tracking data. In some implementations, other tracking systems such as transmitters at fixed locations in the physical environment 100 or the physical second device 130 are used. For example, a sensor on the physical first device 120 or the physical second device 130 may detect a signal from a transmitter and determine a location of the user 110 , the physical first device 120 , or the physical second device 130 within the physical environment 100 based on the transmitted signal.
[0050] In addition, the coordinate system of the physical first device 120 (e.g., coordinate system 102) can be synchronized with the coordinate system of the physical second device 130 (e.g., coordinate system 102). Such synchronization can also compensate for situations where one of the two devices cannot effectively track the physical environment 100. For example, unpredictable lighting conditions can result in a reduced ability to track a scene, or excessive motion (e.g., too far, too fast, or too violently shaken) can result in blurred images or too large a distance for tracking features between video frames, reducing tracking quality.
[0051] Figure 2 Shows Figure 1 1, which includes a display 125 of a physical first device 120. In some implementations, the physical first device 120 (e.g., an HMD) presents a virtual scene 205 to the user 110 via the display 125. For example, if the virtual scene 205 represents an ocean beach, visual sensory content corresponding to the ocean beach can be presented on the display 125 of the physical first device 120. In some implementations, a virtual appendage 215 (e.g., a representation of the user's physical presence (e.g., the physical appendage 115)) can be presented in the virtual scene 205. Therefore, in some implementations, the user 110 can still see a representation of their physical presence in the virtual scene 205.
[0052] In some implementations, the physical first device 120 can determine the position or orientation of the physical second device 130, the user 110, or the physical appendage 115 by collecting image data using the image sensor 122 of the physical first device 120. In addition, in some implementations, the virtual scene 205 can include the virtual second device 230 and the virtual display 235 of the virtual second device 230, for example, a virtual representation of the physical second device 130 and the physical display 135 of the physical second device 130. For example, the user 110 can reach out with an arm (e.g., the physical appendage 115) holding the physical second device 130. Therefore, the virtual scene 205 can include the virtual appendage 215, as well as the virtual second device 230.
[0053] Figure 3A Shown by Figure 1In some implementations, the user 110 cannot view the physical display 135 of the physical second device 130 because the user 110 is immersed in the virtual scene 205. Therefore, in some implementations, the physical second device 130 displays the marker 310 on the physical display 135 of the physical second device 130 to facilitate the physical first device 120 to track the physical second device 130. In some implementations, the marker 310 is used as a reference point for the physical first device 120 to accurately track the position and rotation of the physical second device 130. In some implementations, the marker 310 is displayed on the front-facing display of the physical first device 120, and the marker 310 is used as a reference for the physical second device 130 to accurately track the position and rotation of the physical first device 120. For example, the display of the marker 310 can allow the physical second device 130 to estimate the required posture degrees of freedom (translation and rotation) to determine the posture of the marker 310. Thus, by displaying a marker 310 (e.g., a known pattern) on one device and tracking the marker with another device, the ability of one device to track the other device is enhanced, e.g., drift caused by errors in the inertial measurement of the inertial measurement device can be corrected / minimized. For example, tracking can be enhanced by combining the pose of the marker 310 with the inertial measurement of the inertial measurement device.
[0054] In some implementations, the marker 310 is an image containing texture / information that allows the image to be detected and makes it possible to determine the pose of the image relative to the camera. In some implementations, the marker 310 is a pattern, and in some implementations, the marker is a singular indicator. For example, the marker 310 may include a grid, a cross-hatching, a quadrant identifier, a screen border, etc. In some implementations, the marker 310 is predetermined and stored on the physical second device 130. In some implementations, the marker 310 is transmitted to the physical first device 120, and in some implementations, the marker 310 is determined by the physical first device 120 and transmitted to the physical second device 130. In some implementations, the marker 310 is transmitted to the physical second device 130, and in some implementations, the marker 310 is determined by the physical second device 130 and transmitted to the physical first device 120.
[0055] In some implementations, the marker 310 is displayed only when the screen is visible to another device. For example, when the physical display 135 is visible to the physical first device 120, the marker 310 may be displayed only on the physical display 135 of the physical second device 130. In some implementations, the other device detects an obstruction of the marker 310. For example, the obstruction of the physical display 135 may be visually detected by collecting image data using the image sensor 122. As another example, the obstruction of the marker 310 may be detected based on a touch sensor. For example, the touch screen of the physical display 135 may detect an obstruction (e.g., a finger placed on the display of the marker). In some implementations, when an obstruction of the marker 310 is detected, the marker 310 is displayed only on certain portions of the display. For example, if the user 100 blocks a portion of the physical display 135 (e.g., with a finger), the obstruction may be detected (e.g., visually or based on a touch sensor), and the marker 310 may be displayed on the unobstructed portion of the physical display 135.
[0056] Figure 3B Shown Figure 1 In some implementations, the virtual second device 230 includes a virtual display 235. In some implementations, the physical second device 130 is used as a controller for the virtual experience, for example, the touch screen input of the physical display 135 is detected by the physical second device 130 and sent as input to the physical first device 120. For example, the user 110 can interact with the virtual scene 205 via the input interface of the physical second device 130. Therefore, the physical second device 130 can be presented in the virtual scene 205 as a virtual second device 230, which includes a virtual display 235. In some implementations, the virtual display 235 can present a virtual controller 320, which includes one or more controls, selectable buttons, or any other combination of interactive or non-interactive objects. For example, the user 110 can navigate the virtual scene 205 by interacting with the physical second device 130 based on the virtual representation of the physical second device 130 (e.g., the virtual second device 230).
[0057] In some implementations, the virtual representation of the physical second device 130 can be a two-dimensional area that increases the amount of data that can be presented at a particular time (e.g., a virtual representation of an object), thereby improving the virtual experience of the user 110. In addition, the virtual second device 230 may have dimensions proportional to the input device (e.g., a physical input device). For example, the user 110 may interact with the physical second device 130 more effectively because the input provided by the user 110 through the physical second device 130 visually corresponds to an indication of the input in the virtual second device 230. Specifically, the user 110 may be able to view the virtual second device 230 while physically interacting with the physical second device 130, and the user 110 may expect that their input through the virtual second device 230 will correspond to similar input (or interaction) at the physical second device 130. In addition, because each location on the virtual display 235 of the virtual second device 230 may correspond to a single location on the physical display 135 of the physical second device 130, the user 110 may use the virtual controller 320 presented on the virtual display 235 of the virtual second device 230 to navigate the virtual scene 205 (e.g., up to and including the boundaries of the virtual representation) .
[0058] Figure 4 is a block diagram of an example of a physical first device 120 according to some implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some specific implementations, the physical first device 120 includes one or more processing units 402 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 406, one or more communication interfaces 408 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, SPI, I2C, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 410, one or more displays 412, one or more internally or externally facing image sensors 414, a memory 420, and one or more communication buses for interconnecting these components and various other components.
[0059] In some implementations, one or more communication buses include circuits that interconnect and control communication between system components. In some implementations, one or more I / O devices and sensors 406 include at least one of the following: IMU, accelerometer, magnetometer, gyroscope, thermometer, one or more physiological sensors (e.g., blood pressure monitor, heart rate monitor, blood oxygen sensor, blood glucose sensor, etc.), one or more microphones, one or more speakers, haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.
[0060] In some implementations, one or more displays 412 are configured to present a user interface to the user 110. In some implementations, one or more displays 412 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light emitting field effect transistor (OLET), organic light emitting diode (OLED), surface conduction electron emitter display (SED), field emission display (FED), quantum dot light emitting diode (QD-LED), microelectromechanical system (MEMS), retinal projection system or similar display type. In some implementations, one or more displays 412 correspond to diffraction, reflection, polarization, holographic or waveguide displays. In one example, the physical first device 120 includes a single display. In another example, the physical first device 120 includes a display for each eye of the user 110. In some implementations, one or more displays 412 are capable of presenting a CGR environment.
[0061] In some implementations, the one or more image sensor systems 414 are configured to obtain image data corresponding to at least a portion of the physical environment 100. For example, the one or more image sensor systems 414 can include one or more RGB cameras (e.g., with a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), a monochrome camera, an IR camera, an event-based camera, etc. In various implementations, the one or more image sensor systems 414 also include an illumination source that emits light on the physical environment 100, such as a flash or lighting.
[0062] The memory 420 includes a high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid-state memory devices. In some specific implementations, the memory 420 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices or other non-volatile solid-state storage devices. The memory 420 optionally includes one or more storage devices located away from the one or more processing units 402. The memory 420 includes a non-transitory computer-readable storage medium. In some specific implementations, the memory 420 or the non-transitory computer-readable storage medium of the memory 420 stores the following programs, modules and data structures or their subsets, including an optional operating system 430 and a user interface module 440.
[0063] The operating system 430 includes processes for handling various basic system services and for performing hardware-related tasks. In some specific implementations, the user interface module 440 is configured to present a user interface that utilizes inertial and image tracking to identify the location of the physical second device 130 and presents a virtual representation of the physical second device 130 via one or more displays 412. To this end, in various specific implementations, the user interface module 440 includes an inertial tracking unit 442, an image tracking unit 444, a drift correction unit 446, and a rendering unit 448.
[0064] In some implementations, the inertial tracking unit 442 is configured to obtain inertial data and use the inertial data to determine the position or location of the physical first device 120. In some implementations, the inertial tracking unit may also obtain inertial data from the physical second device 130 and use the inertial data to determine the position of the physical second device 130. In some implementations, the inertial tracking unit may determine the relative position and rotation of the physical second device 130 with respect to the physical first device 120. To this end, in various implementations, the inertial tracking unit 442 includes instructions or logic therefor as well as heuristics and metadata therefor.
[0065] In some implementations, the image tracking unit 444 is configured to obtain image data and use the data to identify the location of the physical first device 120. For example, the image tracking 444 unit may track changes in the image to identify the movement of the physical first device 120. In some implementations, the image tracking unit 444 may identify landmarks or reference points in the image data to identify the location of the physical first device 120. In some implementations, the physical first device 120 may receive image data from the physical second device 130 and use the received image data to determine the location of the physical first device 120 or the physical second device 130. In some implementations, the image tracking unit 444 may identify a marker 310 displayed by the physical second device 130 to identify the relative position and rotation of the physical second device 130. To this end, in various implementations, the image tracking unit 444 includes instructions or logic therefor as well as heuristics and metadata thereof.
[0066] In some implementations, the drift correction unit 446 is configured to use the image tracking data to correlate the inertial tracking data and determine the position and rotation correction of the physical first device 120, the physical second device 130, or the relative position of the physical first device 120 and the physical second device 130. To this end, in various implementations, the drift correction unit 446 includes instructions or logic therefor as well as heuristics and metadata thereof.
[0067] In some implementations, the presentation unit 448 is configured to present content via the one or more displays 412. In some implementations, the content includes a virtual representation of the physical second device 130 (e.g., the virtual second device 230), wherein the virtual representation of the physical second device 130 is presented based on the determined relative position of the physical second device 130. To this end, in various implementations, the presentation unit 448 includes instructions or logic therefor as well as heuristics and metadata thereof.
[0068] Although inertial tracking unit 442, image tracking unit 444, drift correction unit 446, and rendering unit 448 are shown as residing on a single device (e.g., physical first device 120), it should be understood that in other specific implementations, any combination of these units may be located in separate computing devices.
[0069] also, Figure 4 More as a functional description of various features present in a particular implementation, as opposed to a schematic diagram of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately can be combined, and some items can be separated. For example, Figure 4Some functional modules shown separately in the figure can be implemented in a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software or firmware selected for a specific specific implementation.
[0070] Figure 5 1 is a block diagram of an example of a physical second device 130 according to some implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some specific implementations, the physical second device 130 includes one or more processing units 502 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 506, one or more communication interfaces 508 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, BLUETOOTH, ZIGBEE, SPI, I2C, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 510, one or more displays 512, one or more internal or external-facing image sensor systems 514, a memory 520, and one or more communication buses 504 for interconnecting these components and various other components.
[0071] In some implementations, one or more communication buses 504 include circuits that interconnect and control communications between system components. In some implementations, one or more I / O devices and sensors 506 include at least one of the following: an IMU, an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time of flight, etc.), etc.
[0072] In some implementations, one or more displays 512 are configured to present a user interface to the user 110. In some implementations, one or more displays 512 correspond to holographic, DLP, LCD, LCoS, OLET, OLED, SED, FED, QD-LED, MEMS, retinal projection system or similar display types. In some implementations, one or more displays 512 correspond to diffraction, reflection, polarization, holographic or waveguide displays. In one example, the physical second device 130 includes a single display. In some implementations, one or more displays 412 can present a CGR environment.
[0073] In some implementations, the one or more image sensor systems 514 are configured to obtain image data corresponding to at least a portion of the physical environment 100. For example, the one or more image sensor systems 514 can include one or more RGB cameras (e.g., with a CMOS image sensor or a CCD image sensor), monochrome cameras, IR cameras, event-based cameras, etc. In various implementations, the one or more image sensor systems 514 also include an illumination source that emits light on the physical environment 100, such as a flash or lighting.
[0074] The memory 520 includes a high-speed random access memory, such as DRAM, SRAM, DDR RAM or other random access solid-state memory devices. In some specific implementations, the memory 520 includes a non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices or other non-volatile solid-state storage devices. The memory 520 optionally includes one or more storage devices located away from the one or more processing units 502. The memory 520 includes a non-transitory computer-readable storage medium. In some specific implementations, the memory 520 or the non-transitory computer-readable storage medium of the memory 520 stores the following programs, modules and data structures or their subsets, including an optional operating system 530 and a user interface module 540.
[0075] The operating system 530 includes processes for handling various basic system services and for performing hardware-related tasks. In some specific implementations, the user interface module 540 is configured to present the marker 310 via one or more displays 512, which facilitates image tracking of the physical first device 120. To this end, in various specific implementations, the user experience module 540 includes an inertial tracking unit 542, a marker display unit 544, a drift correction unit 546, and a controller unit 548.
[0076] In some implementations, the inertial tracking unit 542 is configured to obtain inertial data and use the inertial data to determine the position or location of the physical second device 130. In some implementations, the inertial tracking unit 542 may also obtain inertial data from the physical first device 120 and use the inertial data to determine the position of the physical first device 120. In some implementations, the inertial tracking unit 542 may determine the relative position and rotation of the physical second device 130 relative to the physical first device 120. To this end, in various implementations, the inertial tracking unit 542 includes instructions or logic therefor as well as heuristics and metadata therefor.
[0077] In some implementations, the marker display unit 544 is configured to display the marker 310 on the physical display 135 of the physical second device 130. In some implementations, the marker 310 displayed on the physical second device 130 helps the physical first device 120 detect the physical second device 130 in the physical environment 100. For example, the physical first device 120 can collect image data including the marker 310 displayed on the physical second device 130, and identify the location of the physical second device 130 by detecting the marker 310 in the image data. To this end, in various implementations, the image tracking unit 544 includes instructions or logic therefor as well as heuristics and metadata thereof.
[0078] In some implementations, the drift correction unit 546 is configured to use the image tracking data to correlate the inertial tracking data and determine the position and rotation correction of the physical first device 120, the physical second device 130, or the relative position of the physical first device 120 and the physical second device 130. To this end, in various implementations, the drift correction unit 446 includes instructions or logic therefor as well as heuristics and metadata therefor.
[0079] In some implementations, the controller unit 548 is configured to receive input at the physical second device 130 from the user 110, where the input is associated with a virtual interface presented to the user 110 by the physical first device 120. For example, the user 110 may navigate the virtual scene 205 in the physical environment 100 by making controller selections on the touch screen of the physical second device 130, where the user input corresponds to a virtual representation of the second device presented by the physical first device 120. To this end, in various implementations, the controller unit 548 includes instructions or logic therefor as well as heuristics and metadata therefor.
[0080] Although the inertial tracking unit 542, marker display unit 544, drift correction unit 546, and controller unit 548 are illustrated as residing on a single device (e.g., the physical second device 130), it should be understood that in other implementations, any combination of these units may be located in separate computing devices.
[0081] also, Figure 5 It is more of a functional description of various features present in a particular implementation, as opposed to a schematic diagram of the implementations described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Figure 5 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are allocated among them will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, or firmware selected for a specific specific implementation.
[0082] Figure 6 A block diagram of an exemplary physical first device 120 (e.g., a head mounted device) according to some implementations is shown. The physical first device 120 includes a housing 601 (or enclosure) that houses various components of the physical first device 120. The housing 601 includes (or is coupled to) eye pads (not shown) disposed at the proximal (of the user 110) end of the housing 601. In various implementations, the eye pads are plastic or rubber pieces that comfortably and snugly hold the physical first device 120 (e.g., a head mounted device) in place on the face of the user 110 (e.g., around the eyes of the user 110).
[0083] The housing 601 houses a display 610 that displays an image and emits light toward or onto the eyes of the user 110. In various implementations, the display 610 emits light through an eyepiece having one or more lenses 605 that refract the light emitted by the display 610 so that the display appears to the user 110 as a virtual distance farther than the actual distance from the eye to the display 610. In order to enable the user 110 to focus on the display 610, in various implementations, the virtual distance is at least greater than the minimum focal length of the eye (e.g., 7 cm). In addition, in order to provide a better user experience, in various implementations, the virtual distance is greater than 1 meter.
[0084] The housing 601 also houses a tracking system that includes one or more light sources 622, a camera 624, and a controller 680. The one or more light sources 622 emit light onto the eyes of the user 110, which is reflected as a light pattern (e.g., a flashing circle) that can be detected by the camera 624. Based on the light pattern, the controller 680 can determine the eye tracking characteristics of the user 110. For example, the controller 680 can determine the gaze direction or blink state (eyes open or closed) of the user 110. As another example, the controller 680 can determine the pupil center, pupil size, or focus point. Therefore, in various specific implementations, light is emitted by the one or more light sources 622, reflected from the eyes of the user 110, and detected by the camera 624. In various specific implementations, the light from the eyes of the user 110 reflects from the hot mirror or passes through the eyepiece before reaching the camera 624.
[0085] The display 610 emits light within a first wavelength range, and the one or more light sources 622 emit light within a second wavelength range. Similarly, the camera 624 detects light within the second wavelength range. In various implementations, the first wavelength range is a visible wavelength range (e.g., a wavelength range within the visible spectrum of approximately 400-700 nm), and the second wavelength range is a near infrared wavelength range (e.g., a wavelength range within the near infrared spectrum of approximately 700-1400 nm).
[0086] In various implementations, eye tracking (or specifically, determined gaze direction) is used to enable user 110 to interact (e.g., user 110 selects an option on display 610 by looking at him), provide perforated rendering (presenting a higher resolution in the area of display 610 that user 110 is looking at and a lower resolution elsewhere on display 610), or correct distortion (e.g., for images to be provided on display 610).
[0087] In various implementations, one or more light sources 622 emit light toward the eyes of the user 110, which is reflected in the form of multiple flashes.
[0088] In various implementations, the camera 624 is a frame / shutter based camera that generates images of the eyes of the user 110 at a particular point in time or at multiple points in time at a frame rate. Each image includes a matrix of pixel values corresponding to pixels of the image, the pixels corresponding to positions of the camera's light sensor matrix.
[0089] In various implementations, camera 624 is an event camera that includes multiple light sensors (e.g., a light sensor matrix) at multiple corresponding locations that generates event messages indicating a specific location of a particular light sensor in response to the particular light sensor detecting a change in light intensity.
[0090] Figure 7700 is a flowchart representation of a method 700 for interacting with a virtual environment according to some specific implementations. In some specific implementations, the method 700 is performed by a device (e.g., Figure 1 , Figure 2 , Figure 4 and Figure 6 The method 700 is performed on a physical first device 120 of the present invention, such as an HMD, a mobile device, a desktop computer, a laptop computer, or a server device. In this example, the method 700 is performed on a device (e.g., the physical first device 120) having one or more displays for displaying an image, so some or all of the features of the method 700 can be performed on the physical first device 120 itself. In other specific implementations, the method 700 is performed on more than one device, for example, the physical first device 120 can wirelessly receive images from an external camera or transmit images to a separate device. In some specific implementations, the method 700 is performed by a processing logic component (including hardware, firmware, software, or a combination thereof). In some specific implementations, the method 700 is performed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., a memory).
[0091] The virtual scene 205 includes an image (e.g., a sequence of frames or other images) displayed on at least a portion of the display space of the display of the HMD and observed through one or more lenses of the HMD. For example, content such as a movie, a sequence of images depicting 3D content, a series of VR images, or a series of other CGR images can be presented on the HMD or provided to the HMD for presentation. The image may include any content displayed on some or all of the display space. Each image may replace some or all of the previous images in the sequence, for example, depending on the frame rate. In some specific implementations, the image completely replaces the previous image in the sequence in the display space of the display. In some specific implementations, the image replaces only a portion of the display space, and some or all of the remaining portion of the display space is occupied by content from the previous image in the sequence or seen through a perspective display.
[0092] At block 710, method 700 uses an image sensor of a first device having a first display to obtain an image of a physical environment, the image including a marker displayed on a second display of a second device. For example, the first device may be an HMD including a camera, and the second device may be a handheld device having a touch screen display. In some implementations, the marker (e.g., a unique pattern) may be displayed on the touch screen display of the handheld device, and the image of the physical environment including the marker may be obtained by the camera of the HMD.
[0093] At box 720, method 700 determines the relative position and orientation of the second device relative to the first device based on the marker. For example, the size, shape, angle, or other observable characteristics of the marker in the image can be analyzed to determine the relative position and orientation of the marker relative to the first device (e.g., a device with a camera from which the image is obtained). In another example, the marker includes a plurality of points in a pattern, and the relative distances between the points in the image of the marker are used to determine the relative position and orientation of the marker relative to the first device. In these examples, since the position of the marker on the second device is known, the relative position and orientation of the second device relative to the first device can be determined accordingly.
[0094] Once the relative position and orientation are determined based on a marker or otherwise using image data, the relative position and orientation of the device can be updated via inertial measurements. In some implementations, one or both of the first device and the second device include a relative inertial measurement system that determines relative inertial motion based on inertial measurements. For example, inertial measurements from the first device and the second device can be synchronized to determine the relative motion of the first device and the second device. In some implementations, the most recently received inertial measurements can be used to determine relative movement, for example, inertial measurements can be measured at time intervals and previous time intervals to determine relative movement. Relative inertial motion can be used to determine the relative position and orientation of the second device relative to the first device. However, based on the frequency of the inertial measurements, inaccuracies such as drift effects can result.
[0095] In some implementations, inaccuracies associated with inertial measurements are minimized by utilizing image data. In some implementations, method 700 identifies the location of a marker in the image obtained at block 710. For example, the first device may identify the marker in the image and calculate the relative rotation or position of the second device based on the location of the marker in the image. In addition, method 700 may determine the relative position and orientation of the second device relative to the first device based solely on the location of the marker in the image data, for example, without using inertial measurements.
[0096] At block 730, method 700 displays a representation of the second device on the first display based on the relative position and orientation of the second device, the representation including virtual content positioned to replace the pattern. For example, the user can see the representation of the second device via the first device (e.g., HMD). In some implementations, the representation of the second device can display any combination of controllers or virtual interactive and non-interactive content. Thus, the user can use the controller or selectable content to interact with or guide the CGR experience.
[0097] Figure 88 is a flowchart representation of a method 800 for interacting with a virtual environment according to some specific implementations. In some specific implementations, the method 800 is performed by a device (e.g., Figure 1 -3 and Figure 5 In this example, the method 800 is performed on at least one device (e.g., Figure 1 -3 and Figure 5 The method 800 is performed on a physical second device 130 having at least one or more displays for displaying images, so some or all of the features of the method 800 may be performed on the physical second device 130 itself. In other specific implementations, the method 800 is performed on more than one device, for example, the physical second device 130 may wirelessly receive images from an external camera or transmit images to a separate device. In some specific implementations, the method 800 is performed by a processing logic component (including hardware, firmware, software, or a combination thereof). In some specific implementations, the method 800 is performed by a processor that executes code stored in a non-transitory computer-readable medium (e.g., a memory).
[0098] At block 810, method 800 uses an image sensor of a first device having a first display to obtain an image of a physical environment, the image including a marker displayed on a second display. For example, the first device can be a handheld device having a touch screen display and a camera (e.g., physical second device 130), and the second display can be a front-facing display of an HMD (e.g., physical first device 120). In some implementations, the marker (e.g., a unique pattern) can be displayed on the second display of the HMD, and the image of the physical environment including the marker can be obtained by the camera of the handheld device.
[0099] At block 820, method 800 determines the relative position and orientation of the HMD relative to the first device based on the marker. For example, the size, shape, angle, or other observable characteristics of the marker in the image may be analyzed to determine the relative position and orientation of the marker relative to the first device (e.g., a device having a camera from which the image is obtained). In another example, the marker includes a plurality of points in a pattern, and the relative distances between the points in the image of the marker are used to determine the relative position and orientation of the marker relative to the first device. In these examples, since the position of the marker on the second display is known, the relative position and orientation of the second device relative to the first device can be determined accordingly.
[0100] Once the relative position and orientation are determined based on a marker or otherwise using image data, the relative position and orientation of the device can be updated via inertial measurements. In some implementations, one or both of the first device and the HMD include a relative inertial measurement system that determines relative inertial motion based on inertial measurements. For example, inertial measurements from the first device and the second HMD can be synchronized to determine the relative motion of the first device and the HMD. In some implementations, the most recently received inertial measurements can be used to determine relative movement, i.e., inertial measurements can be measured at time intervals and previous time intervals to determine relative movement. Relative inertial motion can be used to determine the relative position and orientation of the HMD relative to the first device. However, based on the frequency of the inertial measurements, inaccuracies such as drift effects can result.
[0101] In some implementations, inaccuracies associated with inertial measurements are minimized by utilizing image data. In some implementations, method 800 identifies the location of a marker in the image obtained at block 810. For example, the first device may identify the marker in the image and calculate the relative rotation or position of the HMD based on the location of the marker in the image. In addition, method 800 may determine the relative position and orientation of the HMD relative to the first device based solely on the location of the marker in the image data, for example, without using inertial measurements.
[0102] In some embodiments, a light source (e.g., a series of LEDs, a pixel-based display, etc.) on a second device generates light at a given moment, and the light encoding can be used to synchronize motion data (e.g., accelerometer data, IMU data, etc.) generated by the second device, where an image of the second device is captured by the first device.
[0103] Fig. 9 1 is a flowchart representation of a method 900 for tracking the movement of a device using a light-based indicator to encode device movement or synchronization data. In some implementations, the method 900 is performed by a device (e.g., a physical first device 120), such as an HMD, a mobile device, a desktop computer, a laptop computer, or a server device. In some implementations, the method 700 is performed by a processing logic component (including hardware, firmware, software, or a combination thereof). In some implementations, the method 900 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory).
[0104] At box 910, method 900 obtains an image of a physical environment using an image sensor of a first device. The image includes a description of a second device. The description of the second device includes a description of a light-based indicator provided by a light source on the second device. In some embodiments, the light-based indicator has a plurality of LEDs that generate a binary pattern of light that encodes motion data (e.g., accelerometer data or IMU data) or time data associated with motion data generated via the second device. For example, the light-based indicator of the second device can be a plurality of LEDs that generate a binary pattern of light that encodes current motion data generated at the second device. As another example, such LEDs may generate a binary pattern of light that encodes time data associated with motion data generated via the second device, for example, the time at which a motion sensor on the device captures data relative to the time at which the binary pattern is provided. In other embodiments, the second device includes a pixel-based display that displays a pattern that encodes motion or time data of the second device. In other embodiments, the second device includes an IR light source that generates a pattern of IR light that encodes motion or time data of the second device.
[0105] At box 920, the method synchronizes motion data generated via the second device with processing performed by the first device based on the description of the light-based indicator. In some specific implementations, the motion data includes sensor data about the movement or positioning of the second device. For example, accelerometer data from a second device (e.g., a controller device) can be synchronized with an image of the second device captured by a first device (e.g., a head-mounted device (HMD)). In some specific implementations, the capture time of the image can be precisely synchronized with the capture time of the accelerometer data. Therefore, any computer vision-based determination (e.g., relative positioning or second device and other objects) made using the image can be associated with or otherwise synchronized with the motion data of the second device tracked by the sensor of the second device at the same time. Synchronization can be achieved without the need for wireless communication between devices (e.g., without Bluetooth, wifi, etc.). Using a light-based indicator to provide the motion data itself or the time data for synchronizing motion data received through other channels can be simpler, more time and energy efficient, and more accurate than trying to synchronize multiple devices using wifi or Bluetooth communication without such light-based indicator communication. In some implementations, the two devices are able to accurately synchronize their respective clocks (e.g., determine clock offsets) without the need for wifi or bluetooth time synchronization communications. In some implementations, the second device includes a light source that generates light at a given moment that is captured in a single image and encodes motion data or time synchronization data. In some implementations, such data is transmitted at multiple moments, for example, providing the current position of the second device at time 1, providing the current position of the second device at time 2, and so on.
[0106] In some implementations, the second device is a smartphone, tablet computer, or other mobile device. In some implementations, the second device is a pencil, pen, touchpad, or other handheld or hand-controlled device including a light source. In some implementations, the second device is a watch, bracelet, armband, ankleband, waistband, headband, hat, ring, clothing, or other wearable device including a light source.
[0107] In some implementations, the second device includes a light source that emits light in the visible spectrum. In some implementations, the light source emits light in the IR or other invisible parts of the spectrum. In some implementations, the light source emits light in both the visible and invisible parts of the spectrum.
[0108] In some implementations, the light source displays a light-based indicator by displaying a binary pattern, a barcode, a QR code, a 2D code, a 3D code, a graphic (e.g., an arrow with a direction and a size / length indicating a magnitude). In one implementation, the light source includes eight LEDs configured to emit an eight-bit binary pattern that indicates a numerical value associated with a particular position or movement, such as the direction and / or magnitude of movement from a previous time to a current time. In some implementations, the light-based indicator encodes multiple aspects of the movement data, such as movement in each of the three degrees of freedom (e.g., x, y, z) of a motion sensor. In some implementations, the flashing of the light-based indicator or other time-based indication provides additional information useful in synchronizing the motion data of the second device with the image of the second device.
[0109] At block 930, the method generates a control signal based on synchronization of the motion data with the image. For example, if the motion data of the second device is associated with movement of the second device intended to move an associated cursor displayed on the first device, the method may generate an appropriate signal to cause such movement of the cursor.
[0110] Numerous specific details are set forth herein to provide a comprehensive understanding of the claimed subject matter. However, those skilled in the art will appreciate that the claimed subject matter may be practiced without these specific details. In other instances, methods, devices, or systems known to those of ordinary skill are not described in detail so as not to obscure the claimed subject matter.
[0111] Unless otherwise specifically noted, it should be understood that throughout the specification, discussions utilizing terms such as "process," "compute," "calculate," "determine," and "identify" refer to the actions or processes of a computing device, such as one or more computers or similar electronic computing devices, that manipulate or transform data represented as physical electronic or magnetic quantities within a memory, register, or other information storage device, transmission device, or display device of a computing platform.
[0112] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include multi-purpose microprocessor-based computer systems that access stored software that programs or configures the computing system from a general-purpose computing device to a dedicated computing device that implements one or more specific implementations of the subject matter of the present invention. Any suitable programming, scripting, or other type of language or combination of languages may be used to implement the teachings contained herein in software for programming or configuring a computing device.
[0113] The specific implementation of the method disclosed herein can be performed in the operation of such a computing device. The order of the boxes presented in the above examples can be changed, for example, the boxes can be reordered, combined or divided into sub-boxes. Some boxes or processes can be executed in parallel.
[0114] The use of "suitable for" or "configured to" herein is meant to be open and inclusive language that does not exclude devices that are suitable for or configured to perform additional tasks or steps. In addition, the use of "based on" is meant to be open and inclusive, as a process, step, calculation, or other action "based on" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbers included herein are for ease of explanation only and are not intended to be limiting.
[0115] It will also be understood that, although the terms "first", "second", etc. may be used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, a first node may be referred to as a second node, and similarly, a second node may be referred to as a first node, which changes the meaning of the description, as long as all occurrences of the "first node" are consistently renamed and all occurrences of the "second node" are consistently renamed. Both the first node and the second node are nodes, but they are not the same node.
[0116] The terms used herein are only for describing specific implementations and are not intended to limit the claims. As used in the description of this specific implementation and the appended claims, the singular forms of "a", "an" and "the" are intended to also cover the plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "or" used herein refers to and covers any and all possible combinations of one or more items in the associated listed items. It will also be understood that the term "comprises" or "comprising" when used in this specification specifies the presence of stated features, integers, steps, operations, elements or parts, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts or their groupings.
[0117] As used herein, the term “if” may be interpreted to mean “when the antecedent is true” or “when the antecedent is true” or “in response to determining” or “upon determining” or “in response to detecting” that the antecedent is true, depending on the context. Similarly, the phrase “if it is determined that [the antecedent is true]” or “if [the antecedent is true]” or “when [the antecedent is true]” is interpreted to mean “upon determining that the antecedent is true” or “in response to determining” or “upon determining” that the antecedent is true or “when detecting that the antecedent is true” or “in response to detecting” that the antecedent is true, depending on the context.
[0118] The foregoing description and summary of the present invention should be understood to be illustrative and exemplary in every aspect, rather than restrictive, and the scope of the present invention disclosed herein is determined not only by the detailed description of the exemplary specific implementations, but by the full breadth allowed by the patent law. It should be understood that the specific implementations shown and described herein are only illustrative of the principles of the present invention, and that various modifications can be implemented by those skilled in the art without departing from the scope and essence of the present invention.
Claims
1. A method for interaction, include: At a first device comprising a processor, a computer readable storage medium, and a first display: receiving data corresponding to a tracked position or orientation of the second device from a second device, wherein the tracked position or the orientation of the second device is based on a fixed reference point identifying a physical environment; determining a relative position and orientation of the second device relative to the first device based on the received data; as well as A control signal is generated based on the relative position and orientation of the second device.
2. The method of claim 1 , wherein the control signal generated is based on an input on the second device, wherein the first device uses the relative position and orientation of the second device to enable the second device to function as a three-dimensional (3D) controller, a 3D pointer, or a user interface input device. The method of claim 1 , wherein the generated control signal modifies a user interface element displayed by the first device.
4. The method of claim 1, wherein the first device is a head mounted device and the second device comprises a touch screen. 5 . The method of claim 1 , further comprising displaying real-world content on the first display, wherein the real-world content includes a representation of the physical environment. 6 . The method of claim 1 , further comprising adjusting the determined relative position and orientation of the second device with respect to the first device over time.
7. The method of claim 1, wherein the position and the orientation of the second device are determined via simultaneous positioning and mapping.
8. The method according to claim 1, further comprising: include: detecting that an estimated error associated with the relative position and orientation of the second device with respect to the first device is greater than a threshold; as well as According to detecting that the estimation error is greater than the threshold: The determined relative position and orientation of the second device with respect to the first device is adjusted over time based on the additional data.
9. The method according to claim 1, further comprising: include: A representation of the second device is displayed on the first display based on the relative position and orientation of the second device, the representation including virtual content.
10. The method of claim 9, wherein the virtual content includes controls corresponding to interactions.
11. The method according to claim 10, further comprising: include: obtaining data indicating a touch event corresponding to the control on the second device; as well as The interaction is initiated in response to the touch event.
12. A first device, include: a non-transitory computer-readable storage medium; as well as one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the first device to perform operations including: receiving data corresponding to a tracked position or orientation of the second device from a second device, wherein the tracked position or the orientation of the second device is based on a fixed reference point identifying a physical environment; determining a relative position and orientation of the second device with respect to the first device based on the received data; and A control signal is generated based on the relative position and orientation of the second device.
13. A first device according to claim 12, wherein the generated control signal is based on an input on the second device, and wherein the operation further comprises using the relative position and orientation of the second device to enable the second device to be used as a three-dimensional 3D controller, a 3D pointer or a user interface input device.
14. The first device of claim 12, wherein the generated control signal modifies a user interface element displayed by the first device.
15. The first device of claim 12, wherein the first device is a head mounted device and the second device comprises a touch screen.
16. The first device of claim 12, wherein the position and the orientation of the second device are determined via simultaneous positioning and mapping.
17. A non-transitory computer-readable storage medium storing program instructions, wherein the program instructions can be executed by one or more processors of a first device to perform an operation, wherein the operation include: receiving data corresponding to a tracked position or orientation of the second device from a second device, wherein the tracked position or the orientation of the second device is based on a fixed reference point identifying a physical environment; determining a relative position and orientation of the second device relative to the first device based on the received data; as well as A control signal is generated based on the relative position and orientation of the second device.
18. A non-transitory computer-readable storage medium according to claim 17, wherein the control signal generated is based on an input on the second device, and wherein the operation further includes using the relative position and orientation of the second device to enable the second device to be used as a three-dimensional 3D controller, 3D pointer or user interface input device.
19. A method for interacting, include: At a processor of a first device in a physical environment: providing content at the first device, the content including features controlled based on user input provided by motion of a second device; obtaining motion data from a second device, wherein the second device is different from the first device, the motion data being based on identifying a fixed reference point in image data of the physical environment, the image data being obtained using an outward-facing image sensor on the second device, and the image data comprising at least one image of the physical environment; and Based on the obtained motion data, the user input is identified.
20. The method of claim 19, wherein the first device is a head mounted device.
21. The method of claim 19, wherein the second device is a handheld controller.
22. The method of claim 19, wherein the motion data comprises a position and orientation of the second device in the physical environment over a period of time.
23. The method of claim 19, wherein the motion data is generated by the first device based on the image data received from the second device.
24. The method of claim 19, wherein the motion data is generated by the second device based on the image data and sent to the first device.
25. A first device, include: a non-transitory computer-readable storage medium; as well as one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the first device to perform operations including: providing content at the first device, the content including features controlled based on user input provided by motion of a second device; obtaining motion data from a second device, wherein the second device is different from the first device, the motion data being based on identifying fixed reference points in image data of a physical environment, the image data being obtained using an outward-facing image sensor on the second device, and the image data comprising at least one image of the physical environment; and Based on the obtained motion data, the user input is identified.
26. A non-transitory computer-readable storage medium storing program instructions, wherein the program instructions can be executed by one or more processors of a first device to perform an operation, wherein the operation include: providing content at the first device, the content including features controlled based on user input provided by motion of a second device; obtaining motion data from a second device, wherein the second device is different from the first device, the motion data being based on identifying fixed reference points in image data of a physical environment, the image data being obtained using an outward-facing image sensor on the second device, and the image data comprising at least one image of the physical environment; and Based on the obtained motion data, the user input is identified.
Citation Information
Patent Citations
System for tracking a handheld device in virtual reality
CN107646098A
Touchscreen hover detection in an augmented and / or virtual reality environment
CN107743604A