Generating a ground truth dataset for virtual reality experiences
Patent Information
- Application Number
- CN202180047128.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-30
- Filing Date
- 2021-06-09
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2041-06-09
Smart Images

Figure CN115956259B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 046,150, filed June 30, 2020, the contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The examples set forth in this disclosure relate to the field of virtual reality (VR). More specifically, but not as a limitation, the invention describes systems and methods for generating foundational real-world datasets that can be used for device research and development, for training machine learning and computer vision algorithms, and for generating VR experiences. Background Technology
[0004] Many types of computers and electronic devices available today (such as mobile devices (e.g., smartphones, tablets, and laptops) and wearable devices (e.g., smart glasses, digital eyewear, headbands, head-mounted displays)) include a variety of cameras, sensors, wireless transceivers, input systems (e.g., touch-sensitive surfaces, indicators), peripherals, displays, and graphical user interfaces (GUIs) that allow users to interact with the displayed content.
[0005] Virtual reality (VR) technology generates complete virtual environments that include realistic images, sometimes displayed on VR headsets or other head-mounted displays. VR experiences allow users to move around in the virtual environment and interact with virtual objects. Augmented reality (AR) is a VR technology that combines real-world objects from the physical environment with virtual objects and displays this combination to the user. This combination gives the impression that virtual objects truly exist in the environment, especially when the virtual objects look and behave like real objects. Cross-reality (XR) is generally understood as a general term referring to systems that include or combine elements from AR, VR, and MR (mixed reality) systems.
[0006] Advanced VR technologies, such as computer vision and object tracking, can be used to create perceptually rich and immersive experiences. Computer vision algorithms extract 3D data about the physical world from data captured in digital images or videos. Object recognition and tracking algorithms can be used to detect objects in digital images or videos, estimate their orientation or pose, and track their movement over time.
[0007] Software as a Service (SaaS) is a software access and delivery model in which software, applications, data, services, and other systems are hosted on remote servers. Authorized users typically access the software through a web browser. SaaS platforms allow users to configure, run, test, and manage applications. SaaS platform providers offer servers, storage devices, operating systems, middleware, databases, access to datasets, and other services related to the hosted applications. Attached Figure Description
[0008] The features of the various examples described will be readily understood from the following detailed embodiments with reference to the accompanying drawings. Each feature is indicated by a reference numeral in the specification and several views of the drawings. When multiple similar features exist, a single reference numeral can be assigned to each similar feature, using a lowercase letter to indicate the specific feature.
[0009] Unless otherwise stated, the various features shown in the figures are not drawn to scale. The dimensions of the individual features may be enlarged or reduced for clarity. Several figures depict one or more specific embodiments and are presented by way of example only and should not be construed as limiting. The following figures are included in the figures:
[0010] Figure 1A is a side view (right) of an exemplary hardware configuration for an eye-worn device suitable for a system for generating basic real datasets;
[0011] Figure 1B is a partial cross-sectional perspective view of the right corner of the eye-wearing device in Figure 1A, depicting the right visible light camera and circuit board.
[0012] Figure 1C is a side view (left) of an exemplary hardware configuration of the eye-wearing device of Figure 1A, showing the left visible light camera;
[0013] Figure 1D is a partial cross-sectional perspective view of the left corner of the eye-wearing device in Figure 1C, depicting the left visible light camera and circuit board;
[0014] Figures 2A and 2B are rear views of an exemplary hardware configuration of an eye-wearing device used in the basic real dataset generation system;
[0015] Figure 3 is a graphical depiction of a 3D scene, the left original image captured by the left visible light camera, and the right original image captured by the right visible light camera;
[0016] Figure 4 is a functional block diagram of an exemplary basic real dataset generation system that includes wearable devices (e.g., VR headsets, eyewear), client devices, and server systems connected via various networks.
[0017] Figure 5 is a graphical representation of an exemplary hardware configuration of a client device used in the system for generating the base real dataset of Figure 4;
[0018] Figure 6 is a schematic illustration of a user in an exemplary environment used to describe real-time location and map building;
[0019] Figure 7 is a schematic diagram of an exemplary user interface with optional components suitable for use with the basic real dataset generation system of Figure 4;
[0020] Figure 7A is an illustration of an exemplary environmental component of the user interface of Figure 7;
[0021] Figure 7B is an illustration of an exemplary sensor component of the user interface of Figure 7;
[0022] Figure 7C is an illustration of an exemplary trajectory component of the user interface of Figure 7;
[0023] Figure 7D is an illustration of an exemplary offset, configuration, and submission component of the user interface of Figure 7;
[0024] Figure 8A is an illustration of an exemplary physical environment on a display, which is suitable for use with the underlying real dataset generation system of Figure 4;
[0025] Figure 8B is an illustration of an exemplary virtual environment on a display, which is suitable for use with the underlying real dataset generation system of Figure 4;
[0026] Figure 8C is an illustration of an exemplary composite view on a display, which includes a portion of the virtual environment of Figure 8B, which is presented as an overlay relative to the physical environment of Figure 8A, and is suitable for use with the underlying real dataset generation system of Figure 4.
[0027] Figure 9 is an illustration of an exemplary trace projected as an overlay relative to an exemplary map view associated with a virtual environment, which is suitable for use with the underlying real dataset generation system of Figure 4.
[0028] Figure 10A is a schematic illustration of an exemplary virtual path, waypoint, and gaze direction;
[0029] Figure 10B is an illustration of an exemplary view of multiple waypoints and range icons displayed on a map view for reference;
[0030] Figure 11 is a schematic diagram of an exemplary application of the parallelization utility;
[0031] Figure 12 is a graphical representation of an exemplary server;
[0032] Figure 13 is a flowchart depicting an exemplary method for displaying a sensor array model;
[0033] Figure 14 is a flowchart depicting an exemplary method for generating the underlying real dataset;
[0034] Figure 15 is a flowchart depicting an exemplary method for presenting a preview associated with a virtual path; and
[0035] Figure 16 is a schematic diagram of an exemplary polynomial interpolation module suitable for use with the exemplary underlying real dataset generation system of Figure 4. Detailed Implementation
[0036] Various specific implementations and details are described with reference to examples used to generate a base real-world dataset for assembling, configuring, testing, and developing headsets and related components for AR, VR, and XR experiences or combinations thereof. For example, a first recording device with one or more cameras and one or more inertial measurement units records images and real motion data along a path through a physical environment. A SLAM application calculates the headset trajectory based on the recorded data. A polynomial interpolation module generates a continuous-time trajectory (CTT) function. A set of analog sensors is assembled and used in a virtual environment. The CTT function is used to generate a base real-world output dataset representing the assembled set of motion-simulating sensors in the virtual environment. Because the CTT function is generated using real motion data, the output dataset produces a lifelike and realistic VR experience. Furthermore, this method can be used to generate multiple output datasets at various sampling rates, which can be used to test various new or different sets of analog sensors (e.g., as configured on new or different VR headsets).
[0037] An exemplary method includes configuring a first recording device comprising one or more cameras and at least one inertial measurement unit (IMU). The method then includes recording a master reference set captured by the first recording device during motion traversing a path through a physical environment. The master reference set includes a series of real-world images captured by the cameras and a series of real-world motion data captured by the IMU. A SLAM application uses the master reference set to compute a headset trajectory comprising a series of camera poses. A polynomial interpolation module generates a continuous-time trajectory (CTT) function that approximates the headset trajectory and the series of real-world motion data. In some specific implementations, the CTT function is a Chebyshev polynomial function with one or more fitting coefficients.
[0038] The method includes identifying a first virtual environment and assembling a simulated sensor configuration (e.g., a simulated VR headset) comprising one or more simulated cameras and one or more simulated IMU sensors. Using a CTT function (generated to fit real motion data), the method includes generating a base real output dataset representing the simulated sensor configuration moving along a first virtual path. This virtual path is closely correlated with motion along a real path captured by a first recording device.
[0039] The underlying real-world output dataset comprises a series of virtual images (based on poses calculated using the CTT function for the first virtual environment) and a series of analog IMU data (also calculated using the CTT function). Because the CTT function is generated to fit real-world motion data, the output dataset of the analog sensor configuration is closely correlated with the real-world motion data captured by the first recording device.
[0040] The following detailed description includes systems, methods, techniques, instruction sequences, and computer program products, illustrating examples set forth in this invention. Numerous details and examples are included to provide a thorough understanding of the disclosed subject matter and its associated teachings. However, those skilled in the art will understand how the associated teachings can be applied without such details. Aspects of the disclosed subject matter are not limited to the specific devices, systems, and methods described, as the associated teachings can be applied or practiced in various ways. The terminology and naming used herein are for descriptive purposes only and are not intended to be limiting. Generally, well-known examples of instructions, protocols, structures, and techniques are not necessarily shown in detail.
[0041] As used herein, the terms “coupled” or “connected” refer to any logical, optical, physical, or electrical connection (including links, etc.) through which electrical or magnetic signals generated or provided by one system element are transmitted to another coupled or connected system element. Unless otherwise stated, coupled or connected elements or devices are not necessarily directly connected to each other and may be separated by intermediate components, elements, or communication media, one or more of which may modify, manipulate, or carry electrical signals. The term “on” means that the element is directly supported by the element or indirectly supported by another element integrated into or supported by the element.
[0042] The term "proximal" is used to describe an object or part of an object that is located near, to the left of, or next to an object or person; or it is closer to other parts of the object that can be described as "distal." For example, the end of an object that is closest to an object can be called the proximal end, while the roughly opposite end can be called the distal end.
[0043] For purposes of illustration and discussion, the orientation of wearable or eye-worn devices, client or mobile devices, associated components, and any other devices incorporating a camera, inertial measurement unit (IMU), or both, as shown in any of the accompanying figures, is given by way of example only. In operation, the eye-worn device may be oriented in any other direction suitable for the particular application of the eye-worn device, such as up, down, sideways, or any other orientation. Furthermore, for the purposes of this document, any directional terms such as front, back, inside, outside, towards, left, right, sideways, longitudinal, up, down, high, low, top, bottom, side, horizontal, vertical, and diagonal are used by way of example only and do not limit the orientation or orientation of any camera or inertial measurement unit (IMU) as constructed or otherwise described herein.
[0044] In the context of this invention, virtual reality (VR) refers to and includes all types of systems that generate simulated experiences using virtual objects in any environment. For example, generating a virtual environment as part of an immersive VR experience. Augmented reality (AR) experiences include virtual objects presented in a real physical environment. According to their common meaning, as used herein, virtual reality (VR) includes both VR and AR experiences. Therefore, the terms VR and AR are used interchangeably herein; unless the context clearly indicates otherwise, the use of either term includes the other.
[0045] Generally, the concept of "base reality" refers to information that is as close as possible to objective reality. In the context of this invention, a base reality dataset includes the data necessary to generate a VR experience. For example, a base reality dataset includes images, object locations, sounds, sensations, and other data associated with the environment, as well as the user's movement, trajectory, position, and visual orientation within that environment.
[0046] Real-world, reliable data creates more vivid and convincing VR experiences. Some VR systems use datasets tailored for users wearing specific headsets and moving along finite trajectories within a specific environment. These trajectories describe the user's apparent movement within the environment, including the path taken and the user's apparent motion. In some systems, the user's apparent motion may be limited to a subset of simple linear and angular movements, sometimes referred to as motion primitives. In some cases, the result appears as if a virtual camera is smoothly moving along a path, looking more like a flying aircraft or drone than a walking or running person. Apparent motion in many VR systems appears synthetic because the trajectories are artificially smoothed.
[0047] Other objects, advantages, and novel features of the example will be set forth in part in the detailed description below, and in part will become apparent to those skilled in the art upon examination of the following description and the accompanying drawings, or may be learned by production or operation of the example. The objects and advantages of this subject matter may be realized and achieved by means of the methods, means, and combinations particularly pointed out in the appended claims.
[0048] Now refer in detail to the accompanying drawings and the examples discussed below.
[0049] Figure 4 is a functional block diagram of an exemplary basic real dataset generation system 1000, which includes a wearable device 101, a client device 410, and a server 498 connected via various networks 495 (such as the Internet). The client device 401 can be any computing device, such as a desktop computer or a mobile phone. As used herein, the wearable device 101 includes any of a variety of devices with displays for use with a VR system, such as head-mounted displays, VR headsets, and the eye-wearing device 100 described herein. While the features and functions of the wearable device 101 are described herein with reference to the eye-wearing device 100, such features and functions may be present on any type of wearable device 101 suitable for use with a VR system.
[0050] Figure 1A is a side view (right) of an exemplary hardware configuration of a wearable device 101 in the form of an eye-wearing device 100, which includes a touch-sensitive input device or touchpad 181. As shown, the touchpad 181 may have subtle and barely perceptible boundaries; alternatively, the boundaries may be clearly visible or include raised or otherwise tactile edges that provide feedback to the user about the position and boundaries of the touchpad 181. In other embodiments, the eye-wearing device 100 may include a touchpad on the left side.
[0051] The surface of touchpad 181 is configured to detect finger touches, taps, and gestures (e.g., movement touches) for use with the GUI displayed on the image display of the eye-wearing device, thereby allowing users to navigate and select menu options in an intuitive way, which improves and simplifies the user experience.
[0052] Detection of finger input on touchpad 181 enables several functions. For example, touching anywhere on touchpad 181 can cause the GUI to display or highlight an item on a display screen, which can be projected onto at least one of optical components 180A, 180B. Double-clicking on touchpad 181 selects an item or icon. Sliding or swiping a finger in a specific direction (e.g., from front to back, from back to front, from top to bottom, or from bottom to top) allows an item or icon to slide or scroll in that direction; for example, to move to the next item, icon, video, image, page, or slideshow. Sliding a finger in another direction allows sliding or scrolling in the opposite direction; for example, to move to the previous item, icon, video, image, page, or slideshow. Touchpad 181 can be located virtually anywhere on the eye-wearing device 100.
[0053] In one example, a recognized finger gesture clicked on touchpad 181 initiates the selection or pressing of graphical user interface elements in the image displayed on the image displays of optical components 180A and 180B. Adjustments to the image presented on the image displays of optical components 180A and 180B based on the recognized finger gesture can be a primary action of selecting or submitting graphical user interface elements on the image displays of optical components 180A and 180B for further display or execution.
[0054] As shown in the figure, the eye-wearing device 100 includes a right visible light camera 114B. As further described herein, two cameras 114A and 114B capture image information of the scene from two separate viewpoints. The two captured images can be used to project a 3D display onto an image display for viewing using 3D glasses.
[0055] The eye-wearing device 100 includes a right optical component 180B having an image display for presenting images, such as depth images. As shown in Figures 1A and 1B, the eye-wearing device 100 includes a right visible light camera 114B. The eye-wearing device 100 may include a plurality of visible light cameras 114A, 114B forming a passive three-dimensional camera, such as a stereo camera, wherein the right visible light camera 114B is located at the right corner 110B. As shown in Figures 1C-D, the eye-wearing device 100 also includes a left visible light camera 114A.
[0056] Left and right visible light cameras 114A and 114B are sensitive to wavelengths within the visible light range. Each of the visible light cameras 114A and 114B has a different forward field of view, which overlaps to enable the generation of a three-dimensional depth image; for example, the right visible light camera 114B depicts a right field of view 111B. Typically, a "field of view" is a portion of a scene that is visible in a specific location and orientation in space via the camera. Fields of view 111A and 111B have an overlapping field of view 304 (Figure 3). When the visible light cameras capture images, objects or object features outside of fields of view 111A and 111B are not recorded in the original image (e.g., a photograph or picture). The field of view describes the angular range or amplitude of electromagnetic radiation of a given scene picked up by the image sensors of the visible light cameras 114A and 114B in the captured image of that scene. The field of view can be expressed as the angular size of the view frustum; i.e., the viewing angle. The viewing angle can be measured horizontally, vertically, or diagonally.
[0057] In the exemplary configuration, one or both of the visible light cameras 114A and 114B have a 100° field of view and a resolution of 480 × 480 pixels. The “coverage angle” describes the angular range within which the lens of the visible light camera 114A, 114B, or the infrared camera 410 (see Figure 2A) can effectively image. Typically, a camera lens produces an image circle large enough to completely cover the camera's film or sensor, possibly including some degree of vignetting (e.g., the image darkens towards the edges compared to the center). If the camera lens's coverage angle does not extend across the sensor, the image circle will be visible, typically with strong vignetting towards the edges, and the effective field of view will be limited to the coverage angle.
[0058] Examples of such visible light cameras 114A and 114B include high-resolution complementary metal-oxide-semiconductor (CMOS) image sensors and digital VGA cameras (video graphics arrays) with resolutions of 640p (e.g., 640 × 480 pixels, totaling 0.3 megapixels), 720p, or 1080p. Other examples of visible light cameras 114A and 114B include those that capture high-definition (HD) still images and store these images at a resolution of 1642 × 1642 pixels (or greater); or those that record high-definition video at a high frame rate (e.g., thirty to sixty frames per second or more) and store the recording at a resolution of 1216 × 1216 pixels (or greater).
[0059] The eye-wearing device 100 can capture image sensor data from visible light cameras 114A and 114B, as well as geolocation data digitized by an image processor, for storage in memory. The visible light cameras 114A and 114B capture corresponding left and right raw images in a two-dimensional spatial domain. These raw images include a pixel matrix in a two-dimensional coordinate system, which includes an X-axis for horizontal positioning and a Y-axis for vertical positioning. Each pixel includes color attribute values (e.g., red pixel light value, green pixel light value, or blue pixel light value); and positioning attributes (e.g., X-axis coordinates and Y-axis coordinates).
[0060] To capture stereoscopic images for later display as a 3D projection, an image processor 412 (shown in Figure 4) may be coupled to visible light cameras 114A, 114B to receive and store visual image information. The image processor 412, or another processor, controls the operation of the visible light cameras 114A, 114B to act as stereoscopic cameras simulating human binocular vision and may add timestamps to each image. The timestamps on each pair of images allow the images to be displayed together as part of a 3D projection. The 3D projection produces an immersive and realistic experience, which is desirable in various contexts including virtual reality (VR) experiences and video games.
[0061] Figure 1B is a cross-sectional perspective view of the right corner 110B of the eye-wearing device 100 of Figure 1A, depicting the right visible light camera 114B of the camera system and the circuit board. Figure 1C is a side view (left) of an exemplary hardware configuration of the eye-wearing device 100 of Figure 1A, showing the left visible light camera 114A of the camera system. Figure 1D is a cross-sectional perspective view of the left corner 110A of the eye-wearing device of Figure 1C, depicting the left visible light camera 114A of the 3D camera and the circuit board.
[0062] Except for the connection and coupling located on the left side 170A, the structure and arrangement of the left visible light camera 114A are substantially similar to those of the right visible light camera 114B. As shown in the example of Figure 1B, the eye-wearing device 100 includes the right visible light camera 114B and a circuit board 140B, which may be a flexible printed circuit board (PCB). The right hinge 126B connects the right corner 110B to the right temple 125B of the eye-wearing device 100. In some examples, the right visible light camera 114B, the flexible PCB 140B, or other components such as electrical connectors or contacts may be located on the right temple 125B or the right hinge 126B.
[0063] The right corner portion 110B includes a corner body 190 and a corner cap; the temple cap is omitted in the cross-section of Figure 1B. Inside the right corner portion 110B are various interconnected circuit boards, such as PCBs or flexible PCBs, which include components for the right visible light camera 114B, a microphone, and low-power wireless circuitry (e.g., for use via Bluetooth).TM Controller circuits for short-range wireless network communication and high-speed wireless circuits (e.g., for wireless LAN communication via Wi-Fi).
[0064] The right visible light camera 114B is coupled to or disposed on the flexible PCB 140B and covered by a visible light camera overlay lens, which is aimed through an opening formed in the frame 105. For example, the right edge 107B of the frame 105, as shown in FIG. 2A, is connected to the right corner 110B and includes an opening for the visible light camera overlay lens. The frame 105 includes a front side configured to face outwards and away from the user's eye. The opening for the visible light camera overlay lens is formed on and extends through the front or outward side of the frame 105. In the example, the right visible light camera 114B has an outward-facing field of view 111B (shown in FIG. 3), the line of sight or viewing angle of which is related to the right eye of the user of the eyewear device 100. The visible light camera overlay lens may also be adhered to the front side or outward-facing surface of the right corner 110B, wherein the opening forms an outward-facing coverage angle, but in a different outward direction. Coupling may also be achieved indirectly via an intermediary member.
[0065] As shown in Figure 1B, the flexible PCB 140B is disposed within the right corner portion 110B and coupled to one or more other components housed in the right corner portion 110B. Although shown as being formed on a circuit board in the right corner portion 110B, the right visible light camera 114B may be formed on a circuit board in the left corner portion 110A, temples 125A, 125B, or frame 105.
[0066] Figures 2A and 2B are rear perspective views of an exemplary hardware configuration of the eye-wearing device 100, including two different types of image displays. The size and shape of the eye-wearing device 100 are designed to be configured for wear by a user; in this example, it is in the form of glasses. The eye-wearing device 100 may take other forms and may be incorporated into any other wearable device 101, such as a VR headset or helmet, a head-mounted display, or other devices suitable for use with a VR system.
[0067] In the example of eyeglasses, the eye-wearing device 100 includes a frame 105 comprising a left edge 107A connected to the right edge 107B via a nose bridge 106 adapted for support by the user's nose. The left and right edges 107A, 107B include corresponding apertures 175A, 175B that hold corresponding optical elements 180A, 180B, such as lenses and display devices. The term "lens" as used herein is intended to include a sheet of transparent or translucent glass or plastic having a curved or flat surface that causes light to converge / diverge or to cause little or no convergence or divergence.
[0068] Although shown as having two optical elements 180A, 180B, the eyewear device 100 may include other arrangements, such as a single optical element (or it may not include any optical elements 180A, 180B), depending on the application of the eyewear device 100 or the intended user. As previously described, the eyewear device 100 includes a left corner portion 110A adjacent to the left side face 170A of the frame 105 and a right corner portion 110B adjacent to the right side face 170B of the frame 105. The corner portions 110A, 110B are integrated into the corresponding sides 170A, 170B of the frame 105 (as shown) or implemented as separate components attached to the corresponding sides 170A, 170B of the frame 105. Alternatively, the corner portions 110A, 110B may be integrated into the temples (not shown) attached to the frame 105.
[0069] In one example, the image display of optical components 180A, 180B includes an integrated image display. As shown in FIG2A, each optical component 180A, 180B includes a suitable display matrix 177, such as a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, or any other such display. Each optical component 180A, 180B also includes one or more optical layers 176, which may include lenses, optical coatings, prisms, mirrors, waveguides, optical strips, and other optical components and any combination thereof. Optical layers 176A, 176B, ..., 176N (shown as 176A-N in FIG2A) may include prisms having suitable dimensions and construction and including a first surface for receiving light from the display matrix and a second surface for emitting light toward the user's eye. The prisms of optical layers 176A-N extend over all or part of apertures 175A, 175B, which are formed in the left and right edges 107A, 107B to allow the user to see the second surface of the prism when viewing through the corresponding left and right edges 107A, 107B. The first surface of the prisms of optical layers 176A-N faces upward from the frame 105, and the display matrix 177 covers the prisms such that photons and light emitted by the display matrix 177 illuminate the first surface. The size and shape of the prisms are designed such that light is refracted within the prisms and directed by the second surface of the prisms of optical layers 176A-N to the user's eye. In this respect, the second surface of the prisms of optical layers 176A-N may be convex to direct light to the center of the eye. The prisms may be selectively sized and shaped to magnify the image projected by the display matrix 177, and light passing through the prisms such that the image viewed from the second surface is larger than the image emitted from the display matrix 177 in one or more dimensions.
[0070] In one example, optical layers 176A-N may include a transparent LCD layer (keeping the lens open) unless and until a voltage is applied to make the layer opaque (closing or blocking the lens). An image processor 412 on the eyewear device 100 may execute a program to apply voltage to the LCD layer to create an active shutter system, thereby adapting the eyewear device 100 for viewing visual content displayed as a three-dimensional projection. Technologies other than LCDs may be used in the active shutter mode, including other types of reactive layers that respond to voltage or another type of input.
[0071] In another example, the image display device for optical components 180A, 180B includes a projected image display as shown in FIG2B. Each optical component 180A, 180B includes a laser projector 150, which is a tri-color laser projector using a scanning mirror or a galvanometer. During operation, a light source (such as the laser projector 150) is positioned in or above one of the temples 125A, 125B of the eyewear device 100. In this example, optical component 180B includes one or more optical strips 155A, 155B, ... 155N (shown as 155A-N in FIG2B) spaced apart across the width of the lens of each optical component 180A, 180B, or across the depth of the lens between the front and rear surfaces of the lens.
[0072] As photons projected by the laser projector 150 travel through the lens of each optical component 180A, 180B, they encounter optical strips 155A-N. When a particular photon encounters a particular optical strip, it is either redirected to the user's eye or passed to the next optical strip. A combination of modulation of the laser projector 150 and modulation of the optical strips can control a particular photon or beam of light. In the example, the processor controls the optical strips 155A-N by emitting mechanical, acoustic, or electromagnetic signals. Although shown as having two optical components 180A, 180B, the eye-wearing device 100 may include other arrangements, such as single or three optical components, or each optical component 180A, 180B may be arranged in a different configuration, depending on the application of the eye-wearing device 100 or the intended user.
[0073] As further shown in Figures 2A and 2B, the eye-wearing device 100 includes a left corner portion 110A adjacent to the left side surface 170A of the frame 105 and a right corner portion 110B adjacent to the right side surface 170B of the frame 105. The corner portions 110A and 110B may be integrated into the corresponding sides 170A and 170B of the frame 105 (as shown) or implemented as separate components attached to the corresponding sides 170A and 170B of the frame 105. Alternatively, the corner portions 110A and 110B may be integrated into the temples 125A and 125B attached to the frame 105.
[0074] In another example, the eye-wearing device 100 shown in Figure 2B may include two projectors, a left projector 150A (not shown) and a right projector 150B (shown as projector 150). The left optical assembly 180A may include a left display matrix 177A (not shown) or left optical strips 155'A, 155'B, ..., 155'N (155', A to N, not shown), configured to interact with light from the left projector 150A. Similarly, the right optical assembly 180B may include a right display matrix 177B (not shown) or right optical strips 155"A, 155"B, ..., 155"N (155", A to N, not shown), configured to interact with light from the right projector 150B. In this example, the eye-wearing device 100 includes a left display and a right display.
[0075] Figure 3 is a graphical depiction of a 3D scene 306, a left raw image 302A captured by a left visible light camera 114A, and a right raw image 302B captured by a right visible light camera 114B. As shown, the left field of view 111A may overlap with the right field of view 111B. The overlapping field of view 304 represents the portion of the image captured by both cameras 114A and 114B. The term "overlap" in relation to field of view means that the pixel matrix in the generated raw image overlaps by thirty percent (30%) or more. "Substantially overlapping" means that the pixel matrix in the generated raw image or the pixel matrix in the infrared image of the scene overlaps by fifty percent (50%) or more. As described herein, the two raw images 302A and 302B can be processed to include a timestamp that allows the images to be displayed together as part of a 3D projection.
[0076] To capture a stereoscopic image, a pair of raw red-green-blue (RGB) images of the real scene 306 are captured at a given time, as shown in Figure 3: a left raw image 302A captured by the left camera 114A and a right raw image 302B captured by the right camera 114B. When the pair of raw images 302A, 302B are processed (e.g., by image processor 412), a depth image is generated. The generated depth image can be viewed on the optical components 180A, 180B of the eye-wearing device, on another display (e.g., image display 580 on client device 401), or on a screen.
[0077] The generated depth image is in a three-dimensional spatial domain and may include a vertex matrix in a three-dimensional positional coordinate system, which includes an X-axis for horizontal positioning (e.g., length), a Y-axis for vertical positioning (e.g., height), and a Z-axis for depth (e.g., distance). Each vertex may include color attributes (e.g., red pixel light value, green pixel light value, or blue pixel light value); positional attributes (e.g., X-coordinate, Y-coordinate, and Z-coordinate); texture attributes; reflectance attributes; or combinations thereof. Texture attributes quantify the perceptual texture of the depth image, such as the spatial arrangement of colors or intensities in the vertex regions of the depth image.
[0078] In one example, the base real dataset generation system 1000 (Figure 4) includes a wearable device 101, such as the eye-worn device 100 described herein, which includes a frame 105, a left temple 110A extending from the left side 170A of the frame 105, and a right temple 125B extending from the right side 170B of the frame 105. The eye-worn device 100 may further include at least two visible light cameras 114A, 114B having overlapping fields of view. In one example, the eye-worn device 100 includes a left visible light camera 114A having a left field of view 111A, as shown in Figure 3. The left camera 114A is attached to the frame 105 or the left temple 110A to capture a left raw image 302A from the left side of scene 306. The eye-worn device 100 further includes a right visible light camera 114B having a right field of view 111B. The right camera 114B is attached to the frame 105 or the right temple 125B to capture the right raw image 302B from the right side of the scene 306.
[0079] Figure 4 is a functional block diagram of an exemplary basic real dataset generation system 1000, which includes a wearable device 101 (e.g., an AR headset or eyewear device 100), a client device 401 (e.g., a desktop computer or mobile phone), and a server system 498 connected via various networks 495 (such as the Internet). The basic real dataset generation system 1000 includes a low-power wireless connection 425 and a high-speed wireless connection 437 between the eyewear device 100 and the client device 401. The client device 401 can be any computing device, such as a desktop computer or mobile phone. As used herein, the wearable device 101 includes any of a variety of devices with a display for use with a VR system, such as a head-mounted display, a VR headset, and the eyewear device 100 described herein. While the features and functions of the wearable device 101 are described herein with reference to the eyewear device 100, such features and functions may be present on any type of wearable device 101 suitable for use with a VR system.
[0080] As shown in Figure 4, as described herein, the eye-wearing device 100 includes one or more visible light cameras 114A, 114B that capture still images, video images, or both. Cameras 114A, 114B may have direct memory access (DMA) to high-speed circuitry 430 and function as stereo cameras. Cameras 114A, 114B can be used to capture initial depth images, which can be rendered into three-dimensional (3D) models, which are texture-mapped images of a red-green-blue (RGB) imaged scene. Device 100 may also include a depth sensor 213 that uses infrared signals to estimate the location of an object relative to device 100. In some examples, depth sensor 213 includes one or more infrared emitters 215 and an infrared camera 410.
[0081] The eye-wear device 100 further includes two image displays for each optical component 180A, 180B (one associated with the left side 170A and one associated with the right side 170B). The eye-wear device 100 also includes an image display driver 442, an image processor 412, low-power circuitry 420, and high-speed circuitry 430. The image displays for each optical component 180A, 180B are used to present images, including still images, video images, or still and video images. The image display driver 442 is coupled to the image displays for each optical component 180A, 180B to control the display of the images.
[0082] The eye-wearing device 100 also includes one or more speakers 440 (e.g., one associated with the left side of the eye-wearing device and another associated with the right side of the eye-wearing device). The speakers 440 may be included in the frame 105, temple 125, or corner 110 of the eye-wearing device 100. The one or more speakers 440 are driven by an audio processor 443 under the control of a low-power circuit 420, a high-speed circuit 430, or both. The speakers 440 are used to present audio signals, including, for example, a beat track. The audio processor 443 is coupled to the speakers 440 to control the presentation of sound.
[0083] The components for the eye-wearing device 100 shown in Figure 4 are located on one or more circuit boards, such as printed circuit boards (PCBs) or flexible printed circuit boards (FPCs) located in the edges or temples. Alternatively or additionally, the depicted components may be located in the corners, frames, hinges, or bridge of the eye-wearing device 100. The left and right visible light cameras 114A, 114B may include digital camera elements, such as complementary metal-oxide-semiconductor (CMOS) image sensors, charge-coupled devices, lenses, or any other corresponding visible or light-capturing elements that can be used to capture data, including still images or videos of scenes with unknown objects.
[0084] As shown in Figure 4, the high-speed circuit 430 includes a high-speed processor 432, a memory 434, and a high-speed wireless circuit 436. In this example, an image display driver 442 is coupled to the high-speed circuit 430 and operated by the high-speed processor 432 to drive the left and right image displays of each optical component 180A, 180B. The high-speed processor 432 can be any processor capable of managing the high-speed communication and operation of any general-purpose computing system required by the eye-wear device 100. The high-speed processor 432 includes the processing resources required to manage high-speed data transmission over a high-speed wireless connection 437 to a wireless local area network (WLAN) using the high-speed wireless circuit 436.
[0085] In some examples, the high-speed processor 432 executes an operating system, such as the LINUX operating system or other such operating system of the eye-wear device 100, and the operating system is stored in memory 434 for execution. Among other duties, the high-speed processor 432, which executes the software architecture of the eye-wear device 100, also manages data transmissions utilizing the high-speed wireless circuit 436. In some examples, the high-speed wireless circuit 436 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, also referred to herein as Wi-Fi. In other examples, the high-speed wireless circuit 436 may implement other high-speed communication standards.
[0086] Low-power circuitry 420 includes a low-power processor 422 and a low-power wireless circuitry 424. The low-power wireless circuitry 424 and high-speed wireless circuitry 436 of the eye-wear device 100 may include a short-range transceiver (Bluetooth). TM Or Bluetooth Low Energy (BLE) and wireless wide area network, local area network or wide area network transceivers (e.g. cellular or Wi-Fi). Client device 401, including transceivers communicating via low-power wireless connection 425 and high-speed wireless connection 437, can be implemented using the architectural details of eye-wearing device 100, just like other components of network 495.
[0087] Memory 434 includes any storage device capable of storing various data and applications, including camera data generated by the left and right visible light cameras 114A, 114B, the infrared camera 410, the image processor 412, and images generated for display by the image display driver 442 for each optical component 180A, 180B. While memory 434 is shown as integrated with high-speed circuitry 430, in other examples, memory 434 may be a separate, independent component of the eye-wearing device 100. In some such examples, electrical wiring may provide a connection from a chip, including either a high-speed processor 432 or a low-power processor 422 of the image processor 412, to memory 434. In other examples, the high-speed processor 432 may manage addressing of memory 434 such that the low-power processor 422 will activate the high-speed processor 432 whenever a read or write operation involving memory 434 is required.
[0088] As shown in Figure 4, the high-speed processor 432 of the eye-wearing device 100 can be coupled to the camera system (visible light cameras 114A, 114B), the image display driver 442, the user input device 491, and the memory 434. As shown in Figure 5, the CPU 530 of the client device 401 can be coupled to the camera system 570, the mobile display driver 582, the user input layer 591, and the memory 540A.
[0089] Server system 498 may be one or more computing devices as part of a service or network computing system, such as computing devices including a processor, memory, and network communication interfaces for communicating with client devices and wearable devices 101 (such as eye-wearing devices 100 described herein) via network 495.
[0090] The output components of the eye-wear device 100 include visual elements, such as left and right image displays (e.g., displays such as liquid crystal displays (LCDs), plasma display panels (PDPs), light-emitting diode (LED) displays, projectors, or waveguides) associated with each lens or optical component 180A, 180B as shown in Figures 2A and 2B. The eye-wear device 100 may include user-facing indicators (e.g., LEDs, speakers, or vibration actuators) or outward-facing signals (e.g., LEDs, speakers). The image display for each optical component 180A, 180B is driven by an image display driver 442. In some exemplary configurations, the output components of the eye-wear device 100 further include additional indicators, such as audible elements (e.g., speakers), tactile elements (e.g., actuators, such as vibration motors for generating tactile feedback), and other signal generators. For example, the device 100 may include a group of user-facing indicators and a group of outward-facing signals. The user-facing indicator group is configured to be seen or otherwise perceived by a user of the device 100. For example, device 100 may include an LED display positioned so that a user can see it, one or more speakers positioned to generate sounds that a user can hear, or actuators providing tactile feedback that a user can feel. Outward-facing signal arrays are configured to be seen or otherwise perceived by an observer near device 100. Similarly, device 100 may include LEDs, speakers, or actuators configured and positioned to be perceived by an observer.
[0091] The input components of the eye-wearing device 100 may include alphanumeric input components (e.g., a touchscreen or touchpad configured to receive alphanumeric input, a photographic optical keyboard, or other alphanumeric-configured elements), point-based input components (e.g., a mouse, touchpad, trackball, joystick, motion sensor, or other pointing instrument), haptic input components (e.g., a push-button switch, a touchscreen or touchpad that senses the position, force, or position and force of a touch or touch gesture, or other haptic-configured elements), and audio input components (e.g., a microphone). The client device 401 and server system 498 may include alphanumeric, point-based, haptic, audio, and other input components.
[0092] In some examples, the eye-wearing device 100 includes motion-sensing components referred to as inertial measurement units (IMUs) 472. These motion-sensing components can be microelectromechanical systems (MEMS) with micro-moving parts, typically small enough to be part of a microchip. In some exemplary configurations, the IMU 472 includes an accelerometer, a gyroscope, and a magnetometer. The accelerometer senses the linear acceleration (including acceleration due to gravity) of the device 100 relative to three orthogonal axes (x, y, z). The gyroscope senses the angular velocity of the device 100 about three rotational axes (pitch, roll, yaw). Together, the accelerometer and gyroscope provide positioning, orientation, and motion data about the device relative to six axes (x, y, z, pitch, roll, yaw). If a magnetometer is present, it senses the heading of the device 100 relative to the magnetic north pole. The positioning of device 100 can be determined by position sensors such as GPS unit 473, one or more transceivers for generating relative positioning coordinates, altitude sensors or barometers, and other orientation sensors. Such positioning system coordinates can also be received from client device 401 via low-power wireless circuit 424 or high-speed wireless circuit 436 through wireless connections 425 and 437.
[0093] IMU472 may include, or cooperate with, a digital motion processor or program that acquires raw data from components and calculates multiple useful values regarding the positioning, orientation, and motion of device 100. For example, acceleration data acquired from an accelerometer may be integrated to obtain velocity relative to each axis (x, y, z); and integrated again to obtain the positioning of device 100 (represented in linear coordinates x, y, and z). Angular velocity data from a gyroscope may be integrated to obtain the positioning of device 100 (represented in spherical coordinates). The program used to calculate these effective values may be stored in memory 434 and executed by the high-speed processor 432 of the eye-wearing device 100.
[0094] The eye-worn device 100 may optionally include additional peripheral sensors, such as biometric sensors, characteristic sensors, or display elements integrated with the eye-worn device 100. For example, peripheral device elements may include any I / O components, including output components, motion components, positioning components, or any other such components described herein. For example, biometric sensors may include components that detect facial expressions (e.g., gestures, facial expressions, vocal expressions, body posture, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), or identify a person (e.g., identification based on voice, retina, facial features, fingerprints, or electrophysiological signals such as electroencephalogram data).
[0095] Client device 401 may be a desktop computer, laptop computer, mobile device, smartphone, tablet computer, access point, or any other such device capable of connecting to eye-wearing device 100 using both low-power wireless connection 425 and high-speed wireless connection 437. Client device 401 connects to server system 498 and network 495. Network 495 may include any combination of wired and wireless connections.
[0096] As shown in Figure 4, the basic real dataset generation system 1000 includes a client device 401, such as a desktop computer or mobile device, coupled to a wearable device 101 via a network 495. The basic real dataset generation system 1000 includes multiple memory elements for storing instructions and a processor element for executing these instructions. The processor 432 executes the instructions associated with the basic real dataset generation system 1000 to configure the wearable device 101 (e.g., an eye-wearing device 100) to cooperate with the client device 401. The basic real dataset generation system 1000 may utilize the memory 434 of the eye-wearing device 100 or the memory elements 540A, 540B, 540C of the client device 401 (Figure 5). Furthermore, the basic real dataset generation system 1000 may utilize the processor elements 432, 422 of the wearable or eye-wearing device 100 or the central processing unit (CPU) 530 of the client device 401 (Figure 5). Furthermore, the basic real dataset generation system 1000 can further utilize the memory and processor elements of the server system 498, as described herein. In this respect, the memory and processing capabilities of the basic real dataset generation system 1000 can be shared or distributed across the wearable device 101 or eye-wearing device 100, the client device 401, and the server system 498.
[0097] Figure 12 is a graphical representation of an exemplary server system 498, on which instructions 1208 (e.g., software, programs, applications, applets, or other executable code) can be executed to cause the server system 498 to perform any or more of the methods discussed herein. For example, instructions 1208 may cause the server system 498 to perform any or more of the methods described herein. Instructions 1208 transform a general, unprogrammed server system 498 into a specific server system 498 programmed to perform the described and illustrated functions in the manner described. The server system 498 may operate as a standalone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the server system 498 may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Server system 498 may include, but is not limited to, server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web appliances, network routers, network switches, bridges, or any machine capable of sequentially or otherwise executing instructions 1208 specifying actions to be taken by server system 498. Furthermore, although only a single server system 498 is shown, the term "server" should also be understood to include a collection of machines that individually or jointly execute instructions 1208 to perform any one or more of the methods discussed herein.
[0098] Server system 498 may include server processor 1202, server memory element 1204, and I / O unit 1242, which may be configured to communicate with each other via bus 1244. In the example, processor 1202 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 1206 and processor 1210 that execute instructions 1208. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 5 illustrates multiple processors 1202, server system 498 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0099] Server memory 1204 includes main memory 1212, static memory 1214, and storage units 1216, all of which are accessible by processor 1202 via bus 1244. Main server memory 1204, static memory 1214, and storage units 1216 store instructions 1208 embodying any one or more of the methods or functions described herein. Instructions 1208 may also reside wholly or partially within main memory 1212, static memory 1214, machine-readable medium 1218 within storage unit 1216 (e.g., non-transitory machine-readable storage medium), at least one of processor 1202 (e.g., processor cache), or any suitable combination thereof during execution by server system 498.
[0100] Furthermore, the machine-readable medium 1218 is non-transient (in other words, it does not possess any transient signals) because it does not embody propagation signals. However, labeling the machine-readable medium 1218 as "non-transient" should not be interpreted as meaning that the medium cannot be moved; the medium should be considered as capable of being transported from one physical location to another. Additionally, since the machine-readable medium 1218 is tangible, it can be a machine-readable device.
[0101] I / O component 1242 may include a wide variety of components to receive input, provide output, generate output, send information, exchange information, capture measurement results, and so on. The specific I / O component 1242 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine may not include such a touch input device. It should be understood that I / O component 1242 may include many other components not shown in FIG. 5. In various examples, I / O component 1242 may include output component 1228 and input component 1230. Output component 1228 may include visual components (e.g., displays such as plasma display panels (PDP), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistive feedback mechanisms), other signal generators, and so on. Input component 1230 may include alphanumeric input components (e.g., keyboard, touchscreen configured to receive alphanumeric input, camera optical keyboard or other alphanumeric input components), point-based input components (e.g., mouse, touchpad, trackball, joystick, motion sensor or other pointing instrument), haptic input components (e.g., physical button, touchscreen that provides touch location, touch force or touch gesture, or other haptic input components), audio input components (e.g., microphone), etc.
[0102] In more examples, I / O component 1242 may include biometric component 1232, motion component 1234, environmental component 1236, or positioning component 1238, as well as a wide variety of other components. For example, biometric component 1232 includes components for detecting facial expressions (e.g., gestures, facial expressions, vocalizations, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and recognizing people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 1234 includes accelerometer components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), and the like. Environmental components 1236 include, for example, illuminance sensor components (e.g., a photometer), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., a barometer), sound sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., an infrared sensor that detects nearby objects), gas sensors (e.g., a gas detection sensor that detects the concentration of hazardous gases for safety or measures pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. Positioning components 1238 include position sensor components (e.g., a GPS receiver component), altitude sensor components (e.g., an altimeter or barometer that detects air pressure from which altitude can be derived), orientation sensor components (e.g., a magnetometer), etc.
[0103] Communication can be implemented using a wide variety of technologies. I / O component 1242 further includes communication component 1240, operable to couple server system 498 to network 495 via coupling 1224, or to one or more remote devices 1222 via coupling 1226. For example, communication component 1240 may include a network interface component or another suitable device interfacing with network 495. In further examples, communication component 1240 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, etc. Components (e.g.) (low power consumption) Components, and other communication components for providing communication via other means. Device 1222 may be another machine, such as client device 401 as described herein, or any of various peripheral devices (e.g., peripheral devices coupled via USB).
[0104] Furthermore, the communication component 1240 may detect identifiers or include components operable to detect identifiers. For example, the communication component 1240 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, DataMatrix codes, Dataglyph, MaxiCode, PDF417, Ultra codes, UCCRSS-2D barcodes, and other optical codes), or a sound detection component (e.g., a microphone for identifying tagged audio signals). Additionally, various information can be exported via the communication component 1240, such as location via Internet Protocol (IP) geolocation, etc. The location of signal triangulation, the location of NFC beacon signals that can be detected to indicate a specific location, etc.
[0105] Various memories (e.g., server memory 1204, main memory 1212, static memory 1214, and processor memory 1202) and storage units 1216 may store one or more sets of instructions and data structures (e.g., software) embodying any one or more of the methods or functions described herein, or used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 1208) cause various operations to implement the disclosed examples when executed by processor 1202.
[0106] Instruction 1208 may be sent or received over network 495 using a transmission medium via a network interface device (e.g., a network interface component included in communication component 1240) and using any of a variety of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instruction 1208 may be sent to or received from device 1222 using a transmission medium via coupling 1226 (e.g., peer-to-peer coupling).
[0107] As shown in Figure 4, in some exemplary embodiments, the server storage 1204 of the server system 498 is coupled to one or more databases, including a virtual environment database 450, a headset database 452, a trajectory database 454, and an AR dataset repository 456. In some embodiments, these databases may also be accessed by wearable device 101 (e.g., eye-wearing device 100) and / or client device 401 and are sometimes stored locally on the wearable device and / or client device, as described herein.
[0108] Virtual environment database 450 stores data about various virtual environments, much of which is available and accessible through a website known as the Virtual Environment Marketplace. Headset database 452 stores data about various commercially available wearable devices 101 suitable for use with VR systems, as well as data about arrays of sensors (i.e., intrinsic sensors) coupled to or otherwise provided with such wearable devices 101. In some embodiments, headset database 452 also includes discrete data files about non-intrinsic or additional sensors, cameras, inertial measurement units (IMUs), and other sensors (i.e., extrinsic sensors) that, if selected, can be used with such wearable devices 101. Track database 454 stores data about various pre-recorded tracks suitable for use with AR systems.
[0109] In some implementations, the underlying real dataset generation system 1000 includes multiple virtual machines (VMs) to allocate the applications, data, and tasks to be completed. System 1000 can generate, access, and control multiple managed agents and schedulers assigned to perform certain tasks. Results can be collected and stored in one or more virtual buckets or user-accessible folders. Users can specify the output format type best suited for their specific use. The completed AR dataset is stored in the AR dataset repository 456.
[0110] Server 498 is coupled to one or more applications, including a trajectory recording application 460, a polynomial interpolation module 461, a SLAM application 462, an image rendering application 464, a parallelization utility 466, and a rolling shutter simulation engine 468. In some implementations, these applications may be accessed by client device 401 or wearable device 101 and are sometimes stored locally on the client device and / or wearable device, as described herein.
[0111] The trajectory recording application 460 allows users to record their own trajectories by, for example, wearing a first recording device 726 (e.g., an eye-worn device 100 as described herein) and moving through the physical environment, while the first recording device 726 captures the data needed to generate trajectory data for use with the dataset generation system 1000 described herein.
[0112] In some implementations, the polynomial interpolation module 461 generates a continuous-time trajectory (CTT) function 640, as described herein, for interpolating the trajectory data. The CTT function generates a single curve approximating the headphone trajectory and the actual motion data captured by the first recording device 726.
[0113] SLAM application 462 processes the raw data captured by track recording application 460. In some exemplary embodiments, image rendering application 464 selects and processes visual images for rendering on a display. The available visual images are typically part of a virtual environment file stored in a virtual environment database 450. Parallelization utility 466 parses the data into subsets, called groups, which are processed in parallel to reduce processing time. In some embodiments, rolling shutter simulation engine 468 processes one or more images to simulate a rolling shutter effect, which, if selected, is applied to a sequence of video images.
[0114] Figure 5 is a high-level functional block diagram of an exemplary client device 401. Client device 401 includes a flash memory 540A storing programs to be executed by a CPU 530 to perform all or a subset of the functions described herein. In some specific implementations, client device 401 is a mobile device, such as a smartphone.
[0115] Client device 401 may include camera 570, which includes at least two visible light cameras (first and second visible light cameras with overlapping fields of view) or at least one visible light camera with substantially overlapping fields of view and a depth sensor. Flash memory 540A may further include stored images and videos captured by camera 570.
[0116] As shown in the figure, client device 401 includes an image display 580, a mobile display driver 582 for controlling the image display 580, and a display controller 584. In the example of Figure 5, the image display 580 includes a user input layer 591 (e.g., a touchscreen) that is overlaid on top of the screen used by the image display 580 or otherwise integrated into the screen.
[0117] Any type of client device 401 may include a touchscreen. The structure and operation of touchscreen devices are provided by way of example; the subject matter described herein is not intended to be limited thereto. For the purposes of this discussion, FIG5 therefore provides a block diagram illustration of an exemplary client device 401 having a user interface including a touchscreen input layer 891 for receiving input (touch via hand, stylus or other tool, multi-touch or gesture, etc.) and an image display 580 for displaying content.
[0118] As shown in Figure 5, client device 401 includes at least one digital transceiver (XCVR) 510 for digital wireless communication via a wide-area wireless mobile communication network, shown as a WWANXCVR. Client device 401 also includes additional digital or analog transceivers, such as those for communication via NFC, VLC, DECT, ZigBee, Bluetooth, etc. TM Or a short-range transceiver (XCVR) 520 for short-range network communication via Wi-Fi. For example, the short-range XCVR 520 may take the form of any available bidirectional wireless local area network (WLAN) transceiver that is compatible with one or more standard communication protocols implemented in a wireless local area network (e.g., one of the Wi-Fi standards compliant with IEEE 802.11).
[0119] To generate location coordinates for locating client device 401, client device 401 may include a Global Positioning System (GPS) receiver. Alternatively or additionally, client device 401 may utilize either or both of a short-range XCVR 520 and a WWAN XCVR 510 to generate location coordinates for positioning, for example, based on a cellular network, Wi-Fi, or Bluetooth. TM The positioning systems can generate very accurate location coordinates, especially when used in combination. These location coordinates can be transmitted to the eye-wearing device 100 via one or more network connections through the XCVR510, 520.
[0120] In some examples, client device 401 includes a collection of motion-sensing components called an inertial measurement unit (IMU) 572 for sensing the positioning, orientation, and motion of client device 401. The motion-sensing components can be microelectromechanical systems (MEMS) with micro-moving parts, which are typically small enough to be part of a microchip. In some exemplary configurations, the inertial measurement unit (IMU) 572 includes an accelerometer, a gyroscope, and a magnetometer. The accelerometer senses the linear acceleration (including acceleration due to gravity) of client device 401 relative to three orthogonal axes (x, y, z). The gyroscope senses the angular velocity of client device 401 about three rotational axes (pitch, roll, yaw). Together, the accelerometer and gyroscope can provide positioning, orientation, and motion data about the device relative to six axes (x, y, z, pitch, roll, yaw). If a magnetometer is present, it senses the heading of client device 401 relative to magnetic north pole.
[0121] The IMU572 may include, or cooperate with, a digital motion processor or program that acquires raw data from components and calculates multiple useful values regarding the positioning, orientation, and motion of the client device 401. For example, acceleration data acquired from an accelerometer may be integrated to obtain velocity relative to each axis (x, y, z); and integrated again to obtain the positioning of the client device 401 (represented in linear coordinates x, y, and z). Angular velocity data from a gyroscope may be integrated to obtain the positioning of the client device 401 (represented in spherical coordinates). The program used to calculate these useful values may be stored in one or more memory elements 540A, 540B, 540C and executed by the CPU 540 of the client device 401.
[0122] Transceivers 510 and 520 (i.e., network communication interfaces) conform to one or more of the various digital wireless communication standards utilized by modern mobile networks. Examples of WWAN transceivers 510 include (but are not limited to) transceivers configured to operate according to Code Division Multiple Access (CDMA) and 3rd Generation Partnership Project (3GPP) network technologies, including, for example, but not limited to, 3GPP Type 2 (or 3GPP2) and LTE, sometimes referred to as "4G". For example, transceivers 510 and 520 provide bidirectional wireless communication of information, including digitized audio signals, still images and video signals, web page information for display and web-related input, and various types of message communication to / from client device 401.
[0123] Client device 401 further includes a microprocessor serving as a central processing unit (CPU); as shown in Figure 4, CPU 530. A processor is a circuit having elements constructed and arranged to perform one or more processing functions (typically various data processing functions). Although discrete logic components can be used, these examples utilize components that form a programmable CPU. A microprocessor includes, for example, one or more integrated circuit (IC) chips that incorporate electronic components that perform the functions of the CPU. For example, CPU 530 may be based on any known or available microprocessor architecture, such as Reduced Instruction Set Computing (RISC) using the ARM architecture, as is commonly used today in mobile devices and other portable electronic devices. Of course, other arrangements of the processor circuitry can be used to form CPU 530 or processor hardware in smartphones, laptops, and tablets.
[0124] By configuring the client device 401 to perform various operations, for example, according to instructions or programs executable by the CPU 530, the CPU 530 acts as a programmable host controller for the client device 401. Such operations may include, for example, various general operations of the client device, as well as operations related to programs for applications on the client device. Although the processor can be configured using hardwired logic, a typical processor in a client device is a general-purpose processing circuit configured by executing programs.
[0125] Client device 401 includes a memory or storage system for storing programs and data. In this example, the memory system may include flash memory 540A, random access memory (RAM) 540B, and other memory components 540C as needed. RAM 540B serves as a short-term storage device for instructions and data processed by CPU 530, for example, as working data processing memory. Flash memory 540A typically provides long-term storage.
[0126] Therefore, in the example of client device 401, flash memory 540A is used to store programs or instructions executed by CPU 530. Depending on the type of device, client device 401 may run a mobile operating system, through which specific applications are executed. Examples of mobile operating systems include Google Android, Apple iOS (for iPhone or iPad devices), Windows Mobile, Amazon Fire OS, RIM BlackBerry OS, etc.
[0127] The processor 432 within the eye-wearing device 100 can construct a map of the environment surrounding the eye-wearing device 100, determine the position of the eye-wearing device within the mapped environment, and determine the relative position of the eye-wearing device with respect to one or more objects in the mapped environment. The processor 432 can construct the map and use a Simultaneous Localization and Mapping (SLAM) algorithm applied to data received from one or more sensors to determine location and positioning information. In the context of VR and AR systems, the SLAM algorithm is used to construct and update a map of the environment while tracking and updating the position of the device (or user) within the mapped environment. Mathematical solutions can be approximated using various statistical methods, such as particle filters, Kalman filters, extended Kalman filters, and covariance intersection.
[0128] Sensor data includes images received from one or both of cameras 114A and 114B, distances received from a laser rangefinder, positioning information received from GPS unit 473, or a combination of two or more such sensor data, or data from other sensors that provide data for determining positioning information.
[0129] Figure 6 depicts an exemplary physical environment 600 and elements useful when using SLAM application 462 as described herein and other tracking applications (e.g., Natural Feature Tracking (NFT)). A user 602 of wearable device 101 (such as eye-wearing device 100) is present in the exemplary physical environment 600 (an interior room in Figure 6). Processor 432 of eye-wearing device 100 uses captured images to determine its location relative to one or more objects 604 within environment 600, constructs a map of environment 600 using the coordinate system (x, y, z) of environment 600, and determines its location within the coordinate system. Additionally, processor 432 determines the head pose (roll, pitch, and yaw) of eye-wearing device 100 within the environment by using two or more location points (e.g., three location points 606a, 606b, and 606c) associated with a single object 604a or by using one or more location points 606 associated with two or more objects 604a, 604b, and 604c. The processor 432 of the eye-wearing device 100 can locate and display virtual objects (such as the key 608 shown in Figure 6) within the virtual environment 600 for viewing during an augmented reality (AR) experience. Similarly, the processor 432 of the eye-wearing device 100 can locate and display virtual objects and the virtual environment relative to the user for viewing during an immersive VR experience.
[0130] Figure 7 is a schematic diagram of an exemplary user interface 700, which includes several optional components suitable for use with the underlying real-world dataset generation system 1000 described herein. In some specific implementations, the user interface 700 is presented on the display 580 of a user's client device 401 accessing the server 498. In this respect, the SaaS model can be used to access and deliver the system 1000. As shown, the user interface 700 includes an environment component 710, a sensor component 720, a trajectory component 730, an offset component 740, a configuration component 750, and a submission component 790.
[0131] For the underlying real-world dataset generation system 1000 described herein, the selection of the virtual environment, headset, sensor configuration, and trajectory is independent of other selections. In other words, selections are persistent; changing one selection will not change other selections. Any combination of selections can be submitted for processing. In this respect, the underlying real-world dataset built from the combination of selections is decoupled from any particular selection. As used herein, the term "selection" refers to and includes processes such as selection, configuration, identification, assembly, etc., which can be performed by a person such as an authorized user or by any element of the system configured to perform selections. Therefore, the use of the term "selection" to describe exemplary processes may or may not include human actions.
[0132] Figure 7A is an illustration of an exemplary environment component 710 that provides an interface through which a user can select a virtual environment 715 from a list 711 of virtual environments stored in the virtual environment database 450 described herein. The list 711 may be presented as a drop-down menu, as shown, including a slider and other relevant features typically provided with a menu interface. In some exemplary embodiments, the environment component 710 displays a sample image 712 that includes a representative scene from the selected virtual environment 715.
[0133] In the context of this disclosure, selecting a component (e.g., a virtual environment, a headset, a sensor, a sensor configuration, a trajectory) means or includes selecting one or more files, records, data storage, applications, subroutines, or other elements associated with the selected component.
[0134] Figure 7B is an illustration of an exemplary sensor component 720 of a user interface 700, providing an interface through which a user can select one or more sensor components from a menu 721. In some specific embodiments, menu 721 includes a headset and a list of sensor components stored in a headset database 452 and available, as described herein. Menu 721 may be presented as a drop-down menu, as shown, and may include a slider and other relevant features typically provided with menu interfaces.
[0135] In some implementations, menu 721 includes a list of one or more available headphones or headphone configurations (e.g., Alpha, Beta, Gamma as shown in the figure) that include a predetermined set of sensor components. For example, selecting Beta will include all or most of the sensor components associated with the Beta headphone configuration. The wearable device 101 described herein may be listed on menu 721 for available headphones.
[0136] As shown in the figure, menu 721 also includes a list of one or more discrete sensor components or sets of sensor components (e.g., Alpha Global Shutter, BetaRGB+Depth, Gamma Design Option 2) that can be selected individually; separate from any headset. For example, selecting Alpha Global Shutter will exclude other sensor components associated with the Alpha headset.
[0137] In some implementations, sensor component 720 includes an upload utility 722, as shown, through which a user can upload a file containing data about one or more sensor components or a set of sensor components in a particular arrangement. This file may include a calibration file, which may be available in JSON format. This file may represent a new version of existing headset documentation or a hypothetical version of a headset under development.
[0138] In some implementations, sensor component 720 displays candidate sensor components 725 in a selection window, as shown in the figure. When a component is selected, sensor component 720 stacks the selected sensor elements 765 together to assemble an analog sensor configuration 760, as described herein. Sensor component 720 presents the selected sensor elements 765 together on the display in the form of a graphical sensor array model 760-M, as shown in the figure.
[0139] Selected sensor elements 765 for a particular analog sensor configuration 760 may include one or more analog cameras, such as the camera systems and arrangements described herein, and may further include one or more analog motion sensors, such as the inertial measurement units (IMUs) described herein.
[0140] The analog sensor configuration 760 is described as "analog" because it may include one or more sensor elements in a hypothetical, under-development, or otherwise unavailable arrangement in a single product, such as a commercially available VR headset. In this respect, sensor component 720 facilitates the assembly of the analog sensor configuration 760 for research, development, testing, product trials, and similar endeavors, without having to build a physical prototype headset with physical sensor components.
[0141] Referring again to Figure 7B, the sensor array model 760-M includes selected sensor elements 765 displayed together in a representative graphical arrangement. In some embodiments, the sensor array model 760-M is interactive, allowing selection of one or more sensor elements 765, as described in flowchart 1300 (Figure 13). A selection element (such as cursor 761 shown in Figure 7B) can be used to select a specific sensor element 765. In response, the sensor component 720 presents information about the selected sensor element 765 in a menu frame 763 on the display. As shown, this information includes a sensor name 766 and a list of one or more parameters 767. Additionally, cursor 761 can be used to select a point near model 760 that serves as a virtual anchor point 762. Using cursor 761 as a handle, the sensor array model 760 can be rotated on the display.
[0142] Figure 13 is a flowchart 1300 describing a method for displaying a sensor array model 760-M. One or more steps shown and described may be performed simultaneously, sequentially, in a different order than those shown and described, or in combination with additional steps. Some steps may be omitted, or may be repeated in some applications.
[0143] At frame 1302, sensor component 720 presents a sensor array model 760-M on the display. In some embodiments, sensor array model 760-M is presented in a dimensional representation of the selected sensor elements 765 in configuration and arrangement. For example, the sensor array model 760-M shown in Figure 7B is a three-dimensional wireframe model showing four selected sensor elements 765 arranged in an upper pair and a lower pair. Sensor array model 760-M approximately corresponds to the relative positioning of the selected sensor elements 765 selected as part of the simulated sensor configuration 760.
[0144] At box 1304, cursor 761 can select a point near sensor array model 760-M, and in response, sensor component 720 displays instantaneous anchor point 762, as shown in Figure 7B.
[0145] At frame 1306, sensor component 720 uses cursor 761 as a handle to rotate sensor array model 760-M in response to movement of instantaneous anchor point 762. In this respect, rotating sensor array model 760-M facilitates observation of one or more optional sensors 765 and their relative positioning.
[0146] At box 1308, cursor 761 can select a specific sensor shown on sensor array model 760-M, and in response to the selection, sensor component 720 highlights the selected sensor element 765.
[0147] At box 1310, sensor component 720 presents information about the selected sensor element 765 in menu frame 763 on the display. This information includes the sensor name 766 and a list of one or more parameters 767, as shown in Figure 7B.
[0148] Figure 7C is an illustration of an exemplary trajectory component 730 that provides an interface through which a user can select a trajectory 735 from a collection 731 of available trajectories in a trajectory database 454. In some specific embodiments, the trajectory component 730 presents a trajectory upload option 723 on a display, through which a user can upload a replacement trajectory file, which may be provided as a ZIP file. For example, a user may record a master reference set 610, as described herein, and store it in memory and / or in the trajectory database 454. The upload option 732 can be used to retrieve the recorded master reference set 610.
[0149] The predefined trajectory 735 typically includes complete data about the route along the defined path, including, for example, the location, orientation, and movement (e.g., the inertial attitude vector for each pose) at each location along the defined path. With the predefined trajectory 735, SLAM applications typically do not need to extract data from camera images or IMU data, as described herein, to generate the headset trajectory.
[0150] Figure 7D is an illustration of an exemplary offset component 740, an exemplary configuration component 750, and an exemplary submission component 790. In some specific embodiments, the offset component 740 provides an interface through which a user can input an initial offset, which includes a translation coordinate 741, a rotation coordinate 742, or a combination thereof.
[0151] When generating output dataset 990, the starting point associated with the selected trajectory file (or the recorded master reference set 610) is adjusted based on the initial offset coordinates 741, 742. For example, the trajectory or path may start from the origin (zero, zero, zero) and be horizontally oriented relative to the rotation coordinates. The initial offset coordinates 741, 742 adjust the starting point so that processing begins based on the offset translation coordinates 741 (x, y, z) and the offset rotation coordinates 742 (roll, pitch, yaw).
[0152] In some implementations, configuration component 750 provides an interface through which a user can input motion blur settings 751 and time offset settings 752, as shown in the figure. Each setting is configured with a selector and a slider. These settings 751, 752 are applied when system 1000 generates the underlying real dataset, causing one or more images to be rendered according to settings 751, 752. Motion blur setting 751 can be expressed in microseconds (e.g., fifty microseconds, as shown in the figure) and can be applied during image rendering to adjust the simulated exposure time of a portion of the image. Time offset setting 752 can be expressed in microseconds (e.g., -12 microseconds, as shown in the figure) and represents the time offset between cameras to be applied during image rendering.
[0153] In some implementations, configuration component 750 provides an interface through which users can input various other configuration settings, including, for example, weather conditions, lens effects (e.g., flash, vignetting), external lighting conditions (e.g., sun direction, intensity, movement), internal lighting conditions (e.g., lights and fixtures), character animation, facial landmarks, and output formats (e.g., ZIP, raw data, VIO-ready).
[0154] In some implementations, the submission component 790 includes a simple button, as shown in the figure. Pressing "Submit" will send the job to the process described herein for generating the underlying real dataset, based on selection, upload, or a combination thereof.
[0155] Figure 14 is a flowchart 1400 illustrating the steps in an exemplary method 990 for generating a base real output dataset. Those skilled in the art will understand that one or more steps shown and described may be performed simultaneously, sequentially, in a different order than those shown and described, or in combination with additional steps. Some steps may be omitted, or may be repeated in some implementations. For example, the process described in flowchart 1400 may be performed to generate an output dataset 990 as described herein, as well as to generate multiple additional or sample output datasets as described herein.
[0156] Box 1402 describes exemplary steps for configuring a first recording device 726 coupled to one or more cameras 620 and one or more inertial measurement unit (IMU) sensors 672. For example, the eye-worn device 100 described herein may operate as the first recording device 726. In other examples, the first recording device 726 is a mobile phone, tablet, VR headset, or other portable device including at least one camera 620 and IMU 672. Camera 620 may include any of a variety of camera systems, including cameras similar to those described herein. IMU sensor 672 may include any of a variety of motion sensors, including sensors similar to inertial measurement units (IMUs) 472, 572 described herein. In some specific implementations, the first recording device 726 is worn by a user while moving through a physical environment 600.
[0157] Box 1404 in Figure 14 illustrates exemplary steps of recording a master reference set 610 using a first recording device 726. In some embodiments, the first recording device 726 captures the master reference set 610 in motion as it traverses a real path 612 through the physical environment 600. The master reference set 610 includes (a) a series of real images 912 captured by camera 620 and (b) a series of real motion data 673 captured by IMU 672. In some embodiments, the first recording device 726 includes or is coupled to a trajectory recording application 460 to facilitate the recording process.
[0158] In some specific implementations, a series of real images 912 are captured by one or more high-resolution digital video cameras at an image sampling rate of 624 (e.g., thirty to sixty frames per second or more). The captured real images 912 include not only photographic values but also various data (e.g., depth data) stored as part of the image. For example, an image with depth data is stored as a three-dimensional matrix. Each vertex of this matrix includes color attributes (e.g., for each pixel), location attributes (X, Y, Z coordinates), texture, reflectivity, and other attributes or combinations thereof. The location attributes facilitate the ability of SLAM applications 462 to determine or extract values from each digital image, such as the camera position and camera orientation (i.e., camera pose) associated with each image, and the relative motion of the camera across the series of images.
[0159] In some implementations, a series of real-world motion data 673 is captured by IMU 672 at an inertial sampling rate 674 (e.g., one hundred inertial samples per second (Hz), one kiloHz, or more). IMU 672 captures the motion of a user wearing the first recording device 726 with exquisite detail. The user's motion in the environment includes not only purposeful macroscopic movements, such as walking on various terrains, but also microscopic movements, such as minute changes in head position, unconscious body movements, and almost imperceptible movements, such as vibrations produced by breathing.
[0160] During recording, in some specific implementations, a portion of a series of real images 912 may be displayed on a monitor as they are captured in real time, as shown in FIG8C. For reference, FIG8A shows an exemplary view 801 of the real physical environment 600 (e.g., an office, outdoor space) from the perspective of a user wearing or carrying the first recording device 726 during recording. FIG8B shows an exemplary virtual environment 802 from the same angle, as if the user could see it. FIG8C shows an exemplary composite view 800 in which a portion 803 of the virtual environment 802 is presented as an overlay relative to the physical environment 600. The composite view 800 provides a preview of the virtual environment 802 while allowing the user to move safely through the physical environment 600 along a desired path.
[0161] In some specific implementations, the example trajectory component 730 shown in Figure 7C can be used to select a predefined trajectory file 735 instead of performing the process of recording the master reference set 610.
[0162] Box 1406 in Figure 14 illustrates exemplary steps in calculating the headset trajectory 910 based on the recorded master reference set 610. The calculated headset trajectory 910 includes a series of camera poses 630. In some specific implementations, the headset trajectory 910 is calculated using a SLAM application 462, as described herein, to determine or extract values from a series of real images 912 and generate a series of camera poses 630, which includes the camera position and camera orientation associated with each real image 912. In this respect, the camera poses 630 represent the relative motion of the camera over time across the series of images 912.
[0163] Figure 16 is a schematic diagram including a series of real images 912 and a series of real motion data 673 used to calculate the headset trajectory 910 and camera pose 630. As shown in the upper graph, smaller data points 673C represent real motion data 673 changing over time. Larger data points 630C represent camera pose 630 changing over time. As shown, there are time points without camera pose data (i.e., blank intervals between larger data points). Similarly, there are time points without motion data (i.e., blank intervals between smaller data points). In some specific implementations, the process of calculating the headset trajectory 910 is performed by a SLAM application 462.
[0164] Box 1408 of Figure 14 illustrates exemplary steps for generating a continuous-time trajectory (CTT) function 640 for interpolating the calculated headphone trajectory 910 and a series of real motion data 673. In some implementations, a polynomial interpolation module 461 is used to generate the CTT function 640.
[0165] As shown in Figure 16, the headset trajectory 910 (including camera pose 630) and real motion data 673 represent the input data used for the exemplary polynomial interpolation module 461. Referring again to the upper graph, the smaller data point 673C represents the real motion data 673 changing over time. The larger data point 630C represents the camera pose 630 changing over time. As shown, there are time points without camera pose data (i.e., blank intervals between larger data points). Similarly, there are time points without motion data (i.e., blank intervals between smaller data points). Some applications may interpolate data between points (e.g., in each blank interval) and apply Newton's laws of motion to estimate the missing data. In practice, this approach would consume excessive time and computational resources. Furthermore, this approach introduces errors because some input data contains noise or other interference. In a related aspect, at block 1406 of this exemplary method, when the SLAM application 462 is used to compute the headset trajectory 910, Newton's laws of motion are applied to the data.
[0166] The CTT function 640 is shown in the lower plot, as illustrated, and generates a single curve that changes over time (e.g., the first curve 641, subsequent curves 642, etc., as shown). For any time (t), the CTT function 640 returns output data including position, orientation, velocity, angular velocity, acceleration, and angular acceleration, and in some implementations returns the values necessary to calculate this output data.
[0167] In some specific implementations, the CTT function 640 is a polynomial equation that includes one or more fitting coefficients. In one example, the CTT function 640 includes one or more Chebyshev polynomial functions, where the value of x at time (t) is represented by the following equation:
[0168] x(t)=∑ k c k b k (t)
[0169] Where b k (t) is the k-th basis function, c k These are the fitting coefficients of the basis function. In this example, the basis function b... k (t) includes the Chebyshev polynomial for a given time (t) (shown on the x-axis of the lower graph in Figure 16). The Cheyshev polynomial is recursively defined as:
[0170] b0(t) = 1,
[0171] b1(t) = t, and
[0172] b n+1 (t)=2tb n (t)-b n-1 (t)
[0173] Fitting coefficient c k It varies depending on the input data. For example, for the input data during the first time interval, the fitting coefficient C is calculated. k This ensures that the first curve 641 closely approximates the input data. For subsequent time intervals, a subsequent set of fitting coefficients C is calculated. k This makes the subsequent curve 642 closely approximate the input data during the subsequent time interval.
[0174] Referring again to Figure 16, the input data includes captured real motion data 673 and a calculated series of camera poses 630. As shown, motion data 673 is used twice: first, as input for calculating the headset trajectory 910, and second, as input for generating the CTT function 640. As shown in the upper graph, motion data point 673C typically occurs more frequently than camera pose data point 630C in this example. In this respect, the process of generating the CTT function 640 uses Chebyshev polynomials to mathematically constrain the interpolation of data points falling within the blank intervals between camera pose data points 630C.
[0175] The CTT function 640 returns values for multiple time intervals, which, when combined, produce the entire trajectory (e.g., a continuous curve, as shown in the lower part of Figure 16). The continuous curve closely approximates the input data (e.g., a smaller data point 673C representing real motion data 673 and a larger data point 630C representing camera pose 630). As shown, the first curve 641 is associated with a first time interval, the subsequent curve 642 with a subsequent time interval, and so on.
[0176] The CTT function 640 generates a continuous curve, as shown in Figure 16, illustrating that mathematically, for any given time (t), the CTT function 640 returns output data values (e.g., position, orientation, velocity, etc.). For example, even if a series of camera poses 630 in the headset trajectory 910 are calculated at a relatively low frequency (e.g., only thirty-six times per second), the CTT function 640 returns output data values that can be used to estimate the camera pose of the output dataset 990 more frequently (or less frequently) (e.g., seven hundred and twenty times per second). In a related aspect, the CTT function 640 is also capable of returning a series of output data values (e.g., position, orientation, velocity, etc.) for any duration at any specific or selected time interval (e.g., ten times, one hundred times, one thousand times per second). In this respect, for example, the CTT function 640 (associated with a single headset trajectory 910 and a series of motion data 673) can be used to generate multiple different output datasets (e.g., position and angular velocity, one hundred values per second, for a duration of three minutes; orientation and acceleration, fifty values per second, for a duration of ninety seconds; etc.).
[0177] Box 1410 of Figure 14 illustrates exemplary steps for identifying a first virtual environment 900. The virtual environment 900 typically includes a virtual map 913 and a vast library of virtual images, image assets, metadata, artifacts, sensory data files, depth maps, and other data used to generate VR experiences.
[0178] In some implementations, the process of identifying the first virtual environment 900 includes receiving a selection from a list 711 of virtual environments through an interface such as the exemplary environment component 710 shown in FIG7A.
[0179] Box 1412 of Figure 14 illustrates exemplary steps of assembling an analog sensor configuration 760 comprising one or more analog cameras 728 and one or more analog motion sensors 729. The analog camera 728 may include any of a variety of camera systems, including those described herein. The analog motion sensor 729 may include any of a variety of motion sensors, including the inertial measurement unit (IMU) described herein. The sensor configuration 760 is referred to as analog because the process of assembling it involves and includes the virtual selection and assembly of one or more components 728, 729 to create a virtual model of a device (e.g., a virtual VR headset) that includes or supports the selected components of the analog sensor configuration 760.
[0180] In some specific implementations, the process of assembling the analog sensor configuration 760 includes selecting the sensor element 765 through an interface such as the exemplary sensor component 720 shown in FIG7B.
[0181] In some exemplary embodiments, the process of selecting the analog sensor configuration 760 includes selecting an initial offset position 741 and an initial orientation 742 relative to a virtual map 913 associated with the selected first virtual environment 900. Figure 7D is an illustration of an exemplary offset component 740 providing an interface for selecting and inputting the initial offset, which includes translation coordinates 741, rotation coordinates 742, or a combination thereof. In this respect, the initial offset position 741 and initial orientation 742 are set, and the CTT function 640 is used to generate a new or different sample output dataset, referred to as the offset output dataset.
[0182] Box 1414 of Figure 14 illustrates exemplary steps of generating an output dataset 990 for a selected analog sensor configuration 760 using the CTT function 640. The resulting output dataset 990 represents output data associated with the selected analog sensor configuration 760 (e.g., analog camera 728 and analog IMU sensor 729), as if the analog sensor configuration 760 were moving along a virtual path 960 through an identified first virtual environment 900. The virtual path 960 is closely related to path 612 through the real environment 600 because data collected by the first recording device 726 along path 612 was used when generating the CTT function 640.
[0183] As described above, at box 1408, the CTT function 640 is generated based on data captured by sensors on the first recording device 726 in the real environment. Now, at box 1414, the CTT function 640 is used to generate an output dataset 990, not for the sensor configuration of the first recording device 726, but for the simulated sensor configuration 760 associated with the virtual environment 900.
[0184] For example, the first recording device 726 may include a single IMU 672 that captures real motion data 673 at an inertial sampling rate 674 of sixty samples per second as the first recording device 726 moves along path 612 through the real physical environment 600. An exemplary analog sensor configuration 760 may include four analog IMU sensors 729 located at specific locations on an analog device (e.g., a virtual reality VR headset). In use, the CTT function 640 can be used to generate an output dataset 990 that includes values (e.g., position, orientation, acceleration) from each of the four analog IMU sensors 729, as if the analog sensor configuration 760 were moving along a virtual path 960 through the identified first virtual environment 900. Because the CTT function 640 is generated using real data captured along a real path, the output dataset 990 of the analog sensor configuration 760 produces a lifelike and realistic series of images and movements.
[0185] In some implementations, the process of generating the output dataset 990 involves selecting and applying a specific sampling rate (e.g., N values per second), referred to herein as the modified sampling rate. As shown in Figure 16, the CTT function 640 generates a continuous curve that varies over time. When generating the output dataset 990, the modified sampling rate can be selected and applied to any or all output values (e.g., N position coordinates per second, N orientation vectors per second, N velocity values per second). In this respect, the modified sampling rate can be arbitrary; that is, a rate not based on or associated with another sampling rate (e.g., the sampling rate associated with the main reference set 610) or not associated with another sampling rate.
[0186] In some embodiments, the modified sampling rate may be equal to the image sampling rate 624 of the camera 620 on the first recording device 726, the inertial sampling rate 674 of the IMU 672 on the first recording device 726, the divisor of the image sampling rate 624 and the inertial sampling rate 674, or a primary sampling rate 634 associated with a series of camera poses 630 as part of the calculated headset trajectory 910. In some embodiments, selecting the divisor involves mathematically comparing the image sampling rate 624 (e.g., thirty-six image frames per second) and the inertial sampling rate 674 (e.g., seven hundred and twenty IMU samples per second) to identify one or more divisors (e.g., seventy-two, three hundred and sixty).
[0187] In related aspects, the process of selecting and applying a modified sampling rate can be used to generate multiple sampled output datasets, each characterized by a specific modified sampling rate, all of which are generated using the same CTT function 640.
[0188] The process of generating the output dataset 990 includes rendering a series of virtual images 970 for each of the one or more analog cameras 728 in the analog sensor configuration 760. The process of generating the output dataset 990 also includes generating a series of analog IMU data 980 for each of the one or more analog IMU sensors 729 in the analog sensor configuration 760.
[0189] In some implementations, the image rendering application 464 facilitates the process of rendering a series of virtual images 970 associated with a first environment 900 for each of one or more analog cameras 728. Each virtual image 970 includes a view of the first environment 900 from the perspective of a user located at a virtual waypoint along a virtual path 960, as described herein. The virtual path 960 may begin at a location associated with the starting point of the generated continuous time trajectory 640, or it may begin at a selected initial offset position 741. Similarly, user orientation along the virtual path 960 may be associated with initial orientation data of the generated continuous time trajectory 640, or it may begin at a selected initial orientation 742.
[0190] The process of generating the output dataset 990 also includes generating a series of analog IMU data 980 for each of the one or more analog IMU sensors 729 in the analog sensor configuration 760. In this respect, the series of analog IMU data 980 typically corresponds to a series of camera poses 630. Because the CTT function 640 is generated using real motion data 673, the motion described by the series of analog IMU data 980 in the output dataset 990 is very closely related to the motion captured by the first recording device 726 as it traverses the real path in the physical environment 600.
[0191] Figure 9 is an illustration of an exemplary trace 930 presented on a display as an overlay of an exemplary map view 713 associated with a virtual environment 900. In some embodiments, the virtual environment 900 is identified at block 1410 of the exemplary method shown in Figure 14. In other embodiments, the virtual environment 900 is selected from list 711 via an interface such as exemplary environment component 710 shown in Figure 7A. Data associated with the selected virtual environment 900 typically includes a virtual map 913.
[0192] Referring again to box 1404 of Figure 14, as the first recording device 726 traverses the real path 612 through the physical environment 600, it captures a master reference set 610 in motion. In some implementations, the master reference set 610 is used to calculate a virtual path 960 through the virtual environment 900. Therefore, the virtual path 960 is closely related to the real path 612. In this regard, the trace 930 shown in Figure 9 corresponds to at least a portion of the virtual path 960.
[0193] The map view 713 in Figure 9 corresponds to at least a portion of the virtual map 913 associated with the selected virtual environment 900. In some embodiments, the size and shape of the map view 713 are designed to correspond to the size and shape of the trace 930. As shown in Figure 9, the trace 930 view allows the user to preview the position and shape of the virtual path 960 relative to the selected first virtual environment 900 before submitting data for further processing (e.g., generating output dataset 990, as described herein). For example, if the user observes that the trace 930 shown in Figure 9 is unacceptable (e.g., the trace 930 passes through a physical virtual object such as a wall or structure), the user may choose to use the first recording device 726 to record another master reference set 610.
[0194] Figure 15 is a flowchart 1500 depicting an exemplary method for presenting a preview associated with virtual path 960. Those skilled in the art will understand that one or more steps shown and described may be performed simultaneously, sequentially, in a different order than those shown and described, or in combination with additional steps. Some steps may be omitted, or may be repeated in some implementations.
[0195] Box 1502 describes exemplary steps for presenting trace 930 on a display as an overlay relative to a map view 713 associated with virtual environment 900, as shown in FIG9. In some embodiments, the process of presenting trace 930 includes preliminary steps of calculating a virtual path 960 through the selected virtual environment 900, wherein the virtual path 960 is based on data regarding a real path 612 traversing the physical environment 600 (and stored in the master reference set 610). In some embodiments, the process of calculating the virtual path 960 involves a point-by-point projection of the real path 612 onto a virtual map 913 associated with the selected virtual environment 900.
[0196] In some implementations, the process of presenting trace 930 includes identifying a map view 713 on which trace 930 will be presented. Map view 713 is part of a virtual map 913 associated with the selected virtual environment 900. In some implementations, identifying map view 713 includes selecting dimensions and shapes corresponding to the size and shape of trace 930. For example, the dimensions and shape of map view 713 shown in Figure 9 are designed such that trace 930 occupies approximately one-third of the total area of map view 713. The dimensions and shape of map view 713 are designed to have boundaries such that trace 930 is approximately presented near the center of map view 713.
[0197] Box 1504 of Figure 15 illustrates another exemplary step of presenting the trace 930 on the display as an overlay relative to the alternative exemplary map view 713. The exemplary map view 713 shown in the upper right corner of Figure 10B is relatively small compared to the exemplary map view 713 in Figure 9. More specifically, box 1504 describes an exemplary step of presenting a range icon 940 along the preview trace 930 on the map view 713, wherein the range icon 940 corresponds to a first virtual waypoint 962a along the calculated virtual path 960. In some embodiments, the range icon 940 is oriented in a basic northward direction, thereby providing a recognizable map-based visual cue. In some embodiments, the range icon 940 is oriented in a direction corresponding to the gaze vector 965a of the first waypoint, thereby providing a visual cue regarding the gaze direction associated with the first virtual waypoint 962a. An exemplary virtual path 960 with waypoints and gaze vectors is shown in Figure 10A.
[0198] Figure 10A is a schematic diagram of an exemplary virtual path 960 relative to a portion of a virtual map 913 associated with a selected virtual environment 900. As shown, a frame 901 comprising multiple cells 902 has been applied to the virtual map 913. The virtual path 960 traverses or passes through the vicinity of a first cell 902a. A first virtual waypoint 962a is located within the first cell 902a. As shown in the detailed view, the portion of the virtual path 960 that passes through the first cell 920a includes a segment referred to as the first virtual waypoint 962a. A first waypoint gaze vector 965a is associated with the first virtual waypoint 962a. Based on the generated output dataset 990, the first waypoint gaze vector 965a represents the viewing direction of the analog sensor configuration 760 moving along the virtual path 960.
[0199] Each cell 902 in frame 901 is characterized by a set of viewing directions 904 and a set of viewing heights 906, as shown in reference sample cells 902s in Figure 10A. In some embodiments, the set of viewing directions 904 includes eight directions (e.g., basic and ordinal directions, as shown). In some embodiments, the set of viewing heights 906 is characterized by one or more heights above the ground or other virtual surface in the virtual environment 600. For example, a first viewing height 906-1 may be characterized by a height of three feet (e.g., possibly representing a kneeling or crouching posture), and a second viewing height 906-2 may be characterized by a height of six feet (e.g., representing standing or walking). In some embodiments, each viewing height includes the same set of viewing directions.
[0200] Box 1506 of Figure 15 illustrates an exemplary step of storing a set of gaze views 719 for each of a plurality of cells 902 associated with a frame 901 applied to a virtual map 913, wherein each gaze view 719 is characterized by a viewing direction 904 and a viewing height 906.
[0201] In some implementations, the process of storing a set of gaze views 719 for each cell 902 involves selecting images associated with each viewing direction 904 and each viewing height 906. For example, these images can be selected from data associated with a selected virtual environment 900 stored in the virtual environment database 450. Referring again to the sample cells 902s in Figure 10A, this set of gaze views 719 for sample 902s includes a set of eight images for each viewing height 906-1, 906-2, where these eight images are associated with one of the eight viewing directions 904. In some implementations, this set of gaze views 719 is intended to provide an overall preview rather than a detailed rendering. In this respect, the set of gaze views 719 may be at a relatively low resolution (e.g., compared to a high-resolution rendered image) and may be stored in a cache for fast retrieval.
[0202] Figure 10B is an illustration of an exemplary gaze preview 950, which includes an exemplary set of waypoint gaze views 719-1 to 719-8, and a range icon 940 located in the upper right corner along the preview trail 930 presented on the map view 713.
[0203] Box 1508 of Figure 15 illustrates exemplary steps for identifying a first cell 902a associated with a first virtual waypoint 962a. In some implementations, the process includes identifying the cells traversed by the virtual path 960. As shown in Figure 10A, the virtual path 960 passes through the first cell 902a. The first virtual waypoint 962a is located within the first cell 902a, as shown in the detail view. The first virtual waypoint 962a includes a first waypoint gaze vector 965a, as shown. In some implementations, the first waypoint gaze vector 965a represents the viewing orientation of the analog sensor configuration 760 moving along the virtual path 960 according to the generated output dataset 990. The first waypoint gaze vector 965a can be oriented in any direction relative to the first virtual waypoint 962a. For example, the first waypoint gaze vector 965a is not necessarily perpendicular to the plane of the first virtual waypoint 962a, as shown in this example.
[0204] Box 1510 of Figure 15 illustrates exemplary steps for selecting a first viewing direction 904a that is most closely aligned with the first waypoint gaze vector 965a. Recall that each cell 902 is characterized by a set of viewing directions 904 (e.g., eight directions, as shown in the reference sample cells 902s). The process of selecting the first viewing direction 904a involves comparing the first waypoint gaze vector 965a associated with cell 902a with that set of viewing directions 904 for cell 902a. In the illustrated example, the first viewing direction 904a (e.g., northeast) is most closely aligned with the first waypoint gaze vector 965a.
[0205] Box 1512 of Figure 15 illustrates exemplary steps of presenting a first gaze view 719-1 associated with a first viewing direction 904a on a display. For example, as shown in Figure 10B, the first gaze view 719-1 represents a view seen from a first virtual waypoint 962a in the direction of the first viewing direction 904a (e.g., northeast in the illustrated example).
[0206] In some specific implementations, the process of presenting the first gaze view 719-1 associated with the first viewing direction 904a further includes presenting the first gaze view 719-1 from a perspective of one of the viewing heights (e.g., the viewing heights of the front view 906-1, 906-2, etc.).
[0207] Box 1514 of Figure 15 illustrates exemplary steps for presenting one or more secondary viewing views 719-2 to 719-8 on a display. In some specific embodiments, the secondary viewing views are presented sequentially clockwise according to the set of viewing directions 904. For example, where the first viewing view 719-1 includes an image associated with the first viewing direction 904a (e.g., northeast), the second viewing view 719-2 includes an image facing east, the third viewing view 719-3 includes an image facing southeast, and so on. In this respect, the first viewing view 719-1 serves as the starting point for presenting the secondary viewing views.
[0208] In some specific implementations, the process of presenting one or more secondary gaze views further includes presenting one or more secondary gaze views from an angle at one of the viewing heights (e.g., viewing heights 906-1, 906-2, etc.).
[0209] The process of presenting a gaze view (e.g., a first gaze view 719-1 and one or more secondary gaze views 719-1 to 719-8; collectively referred to as 719-n) provides additional context about the first virtual waypoint 962a relative to other features of the virtual environment. Viewing gaze view 719-n on the display provides the user with a low-resolution preview from the user's perspective, located near the first virtual waypoint 962a along the virtual path 960. As shown in FIG10A, the first viewing direction 904a does not need to be precisely aligned with the gaze vector 965a of the first waypoint. Therefore, the corresponding gaze view represents a rough approximation of the view seen from the first virtual waypoint 962a. For example, if the user finds the gaze view 719-n presented on the display unacceptable, the user may choose to use the first recording device 726 to record another master reference set 610.
[0210] In a related aspect, the process of presenting the gaze view 719-n on the display provides the user with a low-resolution preview of various virtual objects within the selected virtual environment 900. In some implementations, the virtual path 960 includes depth information about the positions of virtual objects that are part of the selected virtual environment 900. For example, the selected virtual environment 900 typically includes various virtual objects such as surfaces, structures, vehicles, people, equipment, natural features, and other types of objects, both realistic and exotic, static and moving. Each virtual object may be located at a known position expressed in coordinates relative to the selected virtual environment 900. Generating the virtual path 960 includes calculating a depth map that includes the virtual positions of one or more virtual objects relative to the virtual path 960. Positioning the virtual objects on the depth map initially allows for accurate rendering and display of the virtual objects relative to a user traversing the virtual path 960.
[0211] In related aspects, in some specific implementations, the process of presenting a series of virtual images 970 includes estimating the position relative to a camera on a headset and the oriented character's eye position. Virtual environment data typically includes data regarding the character's eye position relative to one or more headsets. Eye position is particularly useful when rendering virtual images that include objects positioned very close to the character (e.g., the character's own arms, hands, and other body parts). This accurate rendering of close-up objects produces a more realistic and lifelike experience. In this exemplary implementation, the underlying real dataset generation system 1000 estimates the character's eye position relative to the camera, which produces a calculated series of poses, each persistently associated with the estimated eye position. By associating each pose with the estimated eye position, body part images are rendered at positions closely and persistently associated with the estimated character eye position, resulting in virtual images of body parts closely associated with the user's viewpoint (i.e., egocentric). In this example, the process of rendering a series of virtual images 970 includes rendering a series of character body part images persistently associated with the estimated character eye positions.
[0212] Rendering a series of virtual images (970) consumes significant computing resources and time. For example, rendering 300 frames of high-quality AR video (which would produce 10 seconds of video at 30 frames per second) could take 24 hours with existing systems. With existing systems, including depth images and multiple camera views, along a relatively long trajectory, and with additional processing using custom configurations, it could take up to 20 days to complete.
[0213] Figure 11 is a conceptual diagram 1100 illustrating a parallelization utility 466 associated with virtual path 960. In some specific implementations, image rendering application 464 utilizes parallelization utility 466, which parses data into subsets (referred to as groups) to be processed in parallel, thereby reducing processing time. For example, parallelization utility 466 determines the optimal number of groups based on virtual path 960 (e.g., groups zero to groups n, as shown). Each group processes a subset of virtual path 960, producing a portion of a series of rendered virtual images 970. Each group has a unique timestamp that facilitates the ordered assembly of results from each group after processing.
[0214] In some implementations, the image rendering application 464 utilizes a rolling shutter simulation engine 468, which processes one or more images to simulate a rolling shutter effect, which, if selected, is applied to a sequence of frames in the video data. A rolling shutter is an image processing method in which each frame of a video is captured by rapidly scanning the frame vertically or horizontally. This rapid scanning occurs over short periods of time at intervals called readout times (e.g., between thirty and one hundred microseconds). Images processed using the rolling shutter method typically consist of a series of rows a few pixels wide, each scanned at different times. The rows in a frame are recorded at different times; however, after processing, the entire frame is displayed simultaneously during playback (as if it represented a single moment). Longer readout times can cause noticeable motion artifacts and distortion. Images processed for feature films typically use shorter readout times to minimize artifacts and distortion.
[0215] In some implementations, the rolling shutter simulation engine 468 includes a row-by-row simulated scan of pixels in a subset of virtual images selected from a series of rendered virtual images 970. For example, when applied to, for instance, a first virtual image, the rolling shutter simulation engine 468 typically produces one or more rendered images that have so much distortion that a single pixel from the first virtual image is located at multiple pixel locations within the rendered image. To minimize distortion caused by this effect, in some implementations, the rolling shutter simulation engine 468 calculates the expected pose for each image in the subset of virtual images. For example, the series of camera poses 630 described herein includes one or more vectors associated with each point along path 612. To calculate the expected pose for the series of virtual images 970 in the output dataset 990, in some implementations, the rolling shutter simulation engine 468 interpolates the series of camera poses 630 over time, as applied to the series of virtual images 970 (e.g., including images immediately before and after the selected first virtual image). In some examples, the expected pose of the subsequently rendered image includes the expected pose of each row of pixels in the first virtual image. To improve processing speed, in some implementations, the rolling shutter simulation engine 468 accesses and applies parallelization utilities 466.
[0216] Use Cases
[0217] Based on an exemplary use case of the underlying real-dataset generation system 1000, the output dataset 990 represents a simulated sensor configuration during motion traversing a path 960 through a virtual environment 900. The output dataset 990 is closely related to a first recording device 726, which captures data about its movement as the device 726 moves through a real physical environment. Therefore, the output dataset 990 produces a lifelike and realistic VR experience, especially for users wearing headsets with equipment identical or similar to that on the first recording device 726.
[0218] Based on another exemplary use case of the underlying real dataset generation system 1000, the output dataset 990 is useful in training one or more machine learning algorithms that are part of many VR systems.
[0219] Machine learning refers to algorithms that improve gradually through experience. By processing a large number of different input datasets, machine learning algorithms can develop generalizations that improve on a particular dataset, and then use these generalizations to produce accurate outputs or solutions when processing new datasets. More broadly, machine learning algorithms consist of one or more parameters that adjust or change in response to new experience, thus gradually improving the algorithm; this is similar to a learning process.
[0220] In the context of computer vision and VR systems, mathematical models attempt to mimic the tasks performed by the human visual system, aiming to use computers to extract information from images and achieve accurate understanding of image content. Computer vision algorithms have been developed for various fields, including artificial intelligence, autonomous navigation, and VR systems, to extract and analyze data contained in digital images and videos.
[0221] According to this exemplary use case, output dataset 990 is used to generate a large number of different sample output datasets 995, which are useful for training machine learning algorithms as part of many VR systems. For example, as shown in Figure 7D, system 1000 includes an offset component 740 and a configuration component 750. In use, output dataset 990 can be generated in which the translation coordinates 741 of the offset component 740 are set to zero. In other words, the recorded trajectory 910 starts from the origin (zero, zero, zero). Then, using the same selected component, another sample output dataset 995 is generated except that a different set of translation coordinates 741 is selected, such as (two, six, six). It is identical to the first dataset in almost every respect except that the trajectory starts at a different location. Similarly, using additional selections, any or all of the different components of the configuration component 750 can be changed, possibly hundreds of times, thereby generating a large number of different sample output datasets 995. In fact, a single output dataset 990 can be used to generate a virtually unlimited number of related but slightly different sample output datasets 995. Because generating multiple sample output datasets 995 is time-consuming, the parallelization utility 466 can be used to parse the data into groups for parallel processing, as described herein. Using multiple sample output datasets 995 that include one or more common attributes (e.g., a specific headset) is particularly useful for training ML algorithms. Each sample output dataset can be used to train an ML algorithm that improves the operation of the VR system as a whole. For example, an ML algorithm trained using all the sample output datasets generated using the first recording device 726 will produce a particularly robust and well-trained ML algorithm for later users with VR headsets that are identical or similar to those on the first recording device 726.
[0222] Based on another exemplary use case of the underlying real-world dataset generation system 1000, the output dataset 990 is useful when testing new or different components in combination with other components. For example, system 1000 can be used to construct a hypothetical headset with an analog sensor configuration 760, as described herein, and then used to generate the output dataset 990. After processing, a user can watch a VR experience created using the output dataset 990 and evaluate the performance of the analog sensor configuration 760. For example, the analog sensor configuration 760 may produce an unclear, unstable, or otherwise unsatisfactory VR experience. In subsequent iterations, the analog sensor configuration 760 can be changed, and the user can watch and evaluate the resulting VR experience.
[0223] According to another exemplary use case, the base real dataset generation system 1000 generates multiple sample output datasets 995 by selecting and applying modified sampling rates during processing. In this respect, in some specific implementations, the process of generating output dataset 990 (at box 1414) includes selecting and applying a modified sampling rate, and then generating one or more sample output datasets 995 based on the modified sampling rate. For example, the CTT function 640 may generate a first output dataset including data points based on a first sampling rate (e.g., thirty-six samples per second). The first sampling rate may be changed to a modified sampling rate (e.g., seventy-two samples per second) to produce a new sample output dataset 995 that is generated using the same CTT function 640. In this application, the process of selecting and applying different modified sampling rates can be used to generate many additional sample output datasets 995, which are particularly useful for training the ML algorithms described herein, thereby improving the accuracy and operation of the VR system.
[0224] According to another exemplary use case, the underlying real dataset generation system 1000 generates another sample output dataset 995 by applying a modified flex angle during the process of rendering a series of virtual images. For a headset comprising two spaced-apart cameras, the flex angle describes the angular relationship between the orientation of the first camera and the orientation of the second camera. Although the camera orientations on a headset are generally expected to be parallel (i.e., the flex angle is zero), most headsets undergo flexing during use, resulting in a non-zero flex angle between the cameras. For example, subtle head movements during trajectory recording can lead to measurable flex angles, which in turn affect the images captured by the first camera relative to the second camera. During the process of calculating a series of camera poses 630 as part of a headset trajectory 910, as described herein, each pose vector is based on the camera orientation at each point along the path 612 used to capture each image. When two cameras are present, each camera orientation is based on the first camera orientation, the second camera orientation, and the flex angle. In this respect, the flex angle affects the process of calculating a series of pose vectors and poses. During processing, the base real dataset generation system 1000 can be used to select and apply modified curvature angles (e.g., different from measured or actual curvature angles), thereby generating one or more additional sample output datasets 995 based on the modified curvature angles. Applying various modified curvature angles will generate many additional sample output datasets 995, which are particularly useful for training the ML algorithm described herein, thereby improving the accuracy and operation of the VR system.
[0225] As described herein, any function of the wearable or eye-worn device 100, client device 401, and server system 498 can be embodied in one or more computer software applications or sets of programming instructions. According to some examples, a “function,” “application,” “instruction,” or “program” is a program that performs the functions defined in the program. Various programming languages can be used to develop one or more applications that are structured in various ways, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language). In a particular example, a third-party application (e.g., an entity other than a platform-specific vendor using Android) may be used. TM or iOS TM Applications developed using a Software Development Kit (SDK) can be included in mobile operating systems such as iOS. TM ANDROID TM , Mobile software running on a phone or another mobile operating system. In this example, a third-party application may invoke API calls provided by the operating system to facilitate the functionality described herein.
[0226] Therefore, machine-readable media can take many forms of tangible storage media. Non-volatile storage media include, for example, optical discs or disks, any storage device such as any computer device, such as client devices, media gateways, code converters, etc., that can be used to implement the figures shown. Volatile storage media include dynamic memory, such as the main memory of computer platforms. Tangible transmission media include coaxial cables; copper wires and optical fibers, including wires that form buses within a computer system. Carrier transmission media can take the form of electrical or electromagnetic signals, or sound or light waves, such as those generated during radio frequency (RF) and infrared (IR) data communications. Therefore, common forms of computer-readable media include, for example: floppy disks, floppy disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, DVDs or DVD-ROMs, any other optical media, punched card tapes, any other physical storage media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs, any other memory chips or cartridges, carrier waves for transmitting data or instructions, cables or links for transmitting such carrier waves, or any other medium from which a computer can read program code or data. Many of these forms of computer-readable media can be used to carry one or more sequences of one or more instructions to a processor for execution.
[0227] In addition to what has just been stated above, whether or not it is stated in the claims, the stated or described content is not intended or should not be construed as causing any part, step, feature, object, benefit, advantage or equivalent to be offered to the public.
[0228] It should be understood that, unless otherwise specified herein, the terms and expressions used herein have the general meaning consistent with those in the corresponding fields of investigation and research. Relational terms such as “first” and “second” are used only to distinguish one entity or action from another, and do not necessarily require or imply any actual such relationship or order between these entities or actions. The terms “comprising,” “including,” “containing,” “having,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes or comprises a list of elements or steps includes not only those elements or steps, but may also include other elements or steps not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element prefixed with “a” or “an” does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes that element.
[0229] Unless otherwise stated, any and all measurements, values, ratings, positions, quantities, dimensions, and other specifications set forth in this specification, including those in the appended claims, are approximate, not precise. Such quantities are intended to have a reasonable range consistent with the functions they relate to and the conventions in the fields to which they pertain. For example, unless otherwise expressly stated, parameter values, etc., can vary from said quantity or range by up to plus or minus ten percent.
[0230] Furthermore, as can be seen in the foregoing specific embodiments, various features have been combined in various examples for the purpose of simplifying this disclosure. The disclosed method should not be construed as reflecting an intention to require more features than expressly recited in each claim in the claimed examples. Rather, as reflected in the following claims, the claimed subject matter lies in fewer features than in any single disclosed example. Therefore, the following claims are hereby incorporated into the specific embodiments, wherein each claim exists independently as a separately claimed subject matter.
[0231] While examples considered to be best practices and other examples have been described above, it should be understood that various modifications may be made therein, and the subject matter disclosed herein can be implemented in various forms and examples, and is applicable to many applications, of which only some have been described herein. The appended claims are intended to claim protection for any and all modifications and variations falling within the true scope of the inventive concept.
Claims
1. A method for generating a basic real output dataset, comprising: Configure a first recording device, the first recording device including a camera and an inertial measurement unit; A master reference set is recorded by the first recording device during motion traversing a path through the physical environment, wherein the master reference set includes a series of real images captured by the camera and a series of real motion data captured by the inertial measurement unit; The headphone trajectory is calculated based on the recorded master reference set, wherein the headphone trajectory includes a series of camera poses; A polynomial interpolation module is used to generate a continuous-time trajectory function for interpolating the calculated headphone trajectory and the series of real motion data; Identify the first virtual environment; A simulated sensor configuration is virtually assembled via a user interface, wherein the assembly includes selecting and virtually combining one or more sensor elements such that the simulated sensor configuration includes one or more simulated cameras and one or more simulated inertial measurement unit sensors; as well as An output dataset is generated based on the generated continuous-time trajectory function, such that the output dataset represents the assembled analog sensor configuration in motion along a virtual path through the identified first virtual environment.
2. The method according to claim 1, wherein the process of generating the output dataset further comprises: For each of the one or more simulated cameras, a series of virtual images associated with the identified first virtual environment are rendered; as well as For each of the one or more analog inertial measurement unit (AIM) sensors, a series of analog AIM data is generated.
3. The method according to claim 1, wherein the process of generating the continuous-time trajectory function further comprises: Identify polynomial equations that include basis functions and one or more fitting coefficients, wherein the basis functions include one or more Chebyshev polynomials; For a first time interval, a first set of values for one or more fitting coefficients is calculated such that the polynomial equation having the calculated first set of values generates a first curve that approximates both the calculated headphone trajectory and the series of real motion data during the first time interval. as well as For subsequent time intervals, a subsequent set of values for the one or more fitting coefficients is calculated such that the polynomial equation with the calculated subsequent set of values generates a subsequent curve that approximates both the calculated headphone trajectory and the series of real motion data during the subsequent time intervals.
4. The method according to claim 1, further comprising: Receive the initial offset position and initial orientation relative to the virtual map associated with the first virtual environment; as well as Using the generated continuous-time trajectory function, an offset output dataset is generated, which represents the simulated sensor configuration in motion along an offset virtual path based on the initial offset position and the initial orientation.
5. The method of claim 1, further comprising: The trace corresponding to at least a portion of the generated virtual path is presented on the display as an overlay of a map view associated with a selected first virtual environment, wherein the virtual path is characterized by one or more virtual waypoints and waypoint gaze vectors associated with the series of camera poses. A frame is applied to the virtual path, the frame comprising multiple cells associated with a virtual map; For each of the plurality of cells, a set of gaze views is stored in the cache, each view being characterized by the viewing direction and viewing height relative to the first virtual environment; Select a first virtual waypoint along the virtual path, wherein the first virtual waypoint includes a first waypoint gaze vector; Identify the first cell that is closest to the first virtual waypoint; Select the first viewing direction that is closest to the gaze vector of the first waypoint; Retrieve the first gaze view associated with the first viewing direction from the cache and display it on the display; The display shows a segment of the virtual path and a range icon indicating the location along the segment based on the first virtual waypoint; as well as Present one or more secondary gaze views associated with the first cell.
6. The method of claim 2, wherein the process of rendering a series of virtual images further comprises: Based on the generated continuous-time trajectory function, determine the optimal number of groups for parallel processing; Each of the groups is processed in parallel operations, wherein each of the groups produces a portion of the series of virtual images and simulated inertial measurement unit data, each portion having a unique timestamp; as well as The generated parts are combined based on the unique timestamp.
7. The method of claim 1, wherein the main reference set comprises (a) the series of real images captured by the camera at an image sampling rate, and (b) the series of real motion data captured by the inertial measurement unit at an inertial sampling rate. The calculated headphone trajectory includes a series of camera poses based on the main sampling rate, and The process of generating the output dataset based on the generated continuous-time trajectory function further includes: The modified sampling rate is selected from a group consisting of any sampling rate, image sampling rate, inertial sampling rate, divisor of image sampling rate and inertial sampling rate, and main sampling rate.
8. The method of claim 2, wherein the process of rendering a series of virtual images further comprises: Adjust the simulated exposure time for the first subset of images based on the motion blur settings; Adjust the time offset between the one or more analog cameras used for the second image subset according to the time offset setting; as well as Apply a rolling shutter effect to a selected subset of images based on one or more rolling shutter parameters.
9. The method of claim 2, wherein the process of rendering a series of virtual images further comprises: Estimate the character's eye position relative to the one or more simulated cameras; as well as A series of images of the character's body parts are rendered based on the estimated position of the character's eyes.
10. The method of claim 1, wherein the identified first virtual environment comprises one or more virtual objects located at known locations relative to a virtual map, and wherein the process of generating the output dataset further comprises: Calculate a depth map, the depth map including the virtual positions of the one or more virtual objects relative to the virtual path.
11. The method of claim 1, wherein the process of assembling the analog sensor configuration comprises: A sensor array model is presented on the display, which represents the arrangement of one or more optional sensors associated with a candidate analog sensor configuration; The instantaneous anchor point placed by the cursor is displayed on the screen as a layer relative to the sensor array model; and In response to the movement of the cursor over the instantaneous anchor point, the sensor array model is rotated on the display.
12. A system for generating basic real-world datasets, comprising: A first recording device, comprising a camera and an inertial measurement unit; A trajectory recording application configured to generate a master reference set based on the first recording device in motion traversing a path through a physical environment, wherein the master reference set includes a series of real images captured by the camera and a series of real motion data captured by the inertial measurement unit; The headphone trajectory is calculated based on the recorded master reference set, wherein the headphone trajectory includes a series of camera poses; A polynomial interpolation module, configured to generate a continuous-time trajectory function for interpolating the calculated headphone trajectory and the series of real motion data; An environment component, the environment component being configured to recognize a first virtual environment; A sensor component configured to virtually assemble a simulated sensor configuration via a user interface, wherein the assembly includes selecting and virtually combining one or more sensor elements such that the simulated sensor configuration includes one or more simulated cameras and one or more simulated inertial measurement unit sensors; and The output dataset generated from the generated continuous-time trajectory function represents the assembled analog sensor configuration in motion along a virtual path through the identified first virtual environment.
13. The system for generating a basic real dataset according to claim 12, further comprising: An image rendering application configured to render a series of virtual images associated with an identified first virtual environment for each of the one or more analog cameras.
14. The system for generating a base real dataset according to claim 12, wherein the output dataset is generated based on a modified sampling rate selected from the group consisting of (a) an arbitrary sampling rate, (b) an image sampling rate associated with the camera of the first recording device, (c) an inertial sampling rate associated with the inertial measurement unit of the first recording device, (d) a divisor of the image sampling rate and the inertial sampling rate, and (e) a master sampling rate associated with the series of camera poses of the calculated headphone trajectory.
15. The system for generating a basic real dataset according to claim 12, wherein the polynomial interpolation module is further configured as follows: Identify polynomial equations that include basis functions and one or more fitting coefficients, wherein the basis functions include one or more Chebyshev polynomials; For a first time interval, a first set of values for the one or more fitting coefficients is calculated such that the polynomial equation having the calculated first set of values generates a first curve that approximates both the calculated headphone trajectory and the series of real motion data during the first time interval; and For subsequent time intervals, a subsequent set of values for the one or more fitting coefficients is calculated such that the polynomial equation with the calculated subsequent set of values generates a subsequent curve that approximates both the calculated headphone trajectory and the series of real motion data during the subsequent time intervals.
16. The basic real dataset generation system according to claim 12, wherein the sensor component is further configured as follows: The trace corresponding to at least a portion of the generated virtual path is presented on the display as an overlay of a map view associated with a selected first virtual environment, wherein the virtual path is characterized by one or more virtual waypoints and waypoint gaze vectors associated with the series of camera poses. A frame is applied to the virtual path, the frame comprising multiple cells associated with a virtual map; For each of the plurality of cells, a set of gaze views is stored in the cache, each view being characterized by the viewing direction and viewing height relative to the first virtual environment; Select a first virtual waypoint along the virtual path, wherein the first virtual waypoint includes a first waypoint gaze vector; Identify the first cell that is closest to the first virtual waypoint; Select the first viewing direction that is closest to the gaze vector of the first waypoint; Retrieve the first gaze view associated with the first viewing direction from the cache and display it on the display; The display shows a segment of the virtual path and a range icon indicating the location along the segment based on the first virtual waypoint; as well as Present one or more secondary gaze views associated with the first cell.
17. The underlying real dataset generation system of claim 12, further comprising a parallelization utility configured to: Based on the generated continuous-time trajectory function, determine the optimal number of groups for parallel processing; Each of the groups is processed in parallel, wherein each group produces a portion of the series of virtual images and simulated inertial measurement unit data, each portion having a unique timestamp; and The generated parts are combined based on the unique timestamp.
18. A non-transitory computer-readable medium storing program code, which, when executed, causes an electronic processor to perform the following steps: Configure a first recording device, the first recording device including a camera and an inertial measurement unit; A master reference set is recorded by the first recording device during motion traversing a path through the physical environment, wherein the master reference set includes a series of real images captured by the camera and a series of real motion data captured by the inertial measurement unit; The headphone trajectory is calculated based on the recorded master reference set, wherein the headphone trajectory includes a series of camera poses; A polynomial interpolation module is used to generate a continuous-time trajectory function for interpolating the calculated headphone trajectory and the series of real motion data; Identify the first virtual environment; A simulated sensor configuration is virtually assembled via a user interface, wherein the assembly includes selecting and virtually combining one or more sensor elements such that the simulated sensor configuration includes one or more simulated cameras and one or more simulated inertial measurement unit sensors; as well as An output dataset is generated based on the generated continuous-time trajectory function, such that the output dataset represents the assembled analog sensor configuration in motion along a virtual path through the identified first virtual environment.
19. The non-transitory computer-readable medium storing program code according to claim 18, wherein the process of generating the continuous-time trajectory function further comprises: Identify polynomial equations that include basis functions and one or more fitting coefficients, wherein the basis functions include one or more Chebyshev polynomials; For a first time interval, a first set of values for one or more fitting coefficients is calculated such that the polynomial equation having the calculated first set of values generates a first curve that approximates both the calculated headphone trajectory and the series of real motion data during the first time interval. as well as For subsequent time intervals, a subsequent set of values for the one or more fitting coefficients is calculated such that the polynomial equation with the calculated subsequent set of values generates a subsequent curve that approximates both the calculated headphone trajectory and the series of real motion data during the subsequent time intervals.
20. The non-transitory computer-readable medium storing program code according to claim 18, wherein the program code, when executed, causes the electronic processor to perform the following additional steps: The trace corresponding to at least a portion of the generated virtual path is presented on the display as an overlay of a map view associated with a selected first virtual environment, wherein the virtual path is characterized by one or more virtual waypoints and waypoint gaze vectors associated with the series of camera poses. A frame is applied to the virtual path, the frame comprising multiple cells associated with a virtual map; For each of the plurality of cells, a set of gaze views is stored in the cache, each view being characterized by the viewing direction and viewing height relative to the first virtual environment; Select a first virtual waypoint along the virtual path, wherein the first virtual waypoint includes a first waypoint gaze vector; Identify the first cell that is closest to the first virtual waypoint; Select the first viewing direction that is closest to the gaze vector of the first waypoint; Retrieve the first gaze view associated with the first viewing direction from the cache and display it on the display; The display shows a segment of the virtual path and a range icon indicating the location along the segment based on the first virtual waypoint; as well as Present one or more secondary gaze views associated with the first cell.
Citation Information
Patent Citations
Method and system for combining real shooting of full dome film with three-dimensional virtual scene
CN105488801A
Systems and methods for video processing and display
CN109076249A