Sensor components and head mounted displays

By using sensor components with multiple stacked sensor layers in the head-mounted display, processing image data and extracting features, the problems of slow processing speed and high power consumption in the prior art are solved, and more efficient and accurate data processing and lower power consumption are achieved.

CN115291388BActive Publication Date: 2025-05-13CTRL-LABS CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210638608.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-03-01
Filing Date
2018-07-24
Publication Date
2025-05-13
Estimated Expiration
2038-07-24

AI Technical Summary

Technical Problem

Sensor devices in existing head-mounted displays are prone to saturation when processing large amounts of data, resulting in reduced processing speeds and consume a lot of power and with delays.

Method used

Sensor components employing multiple stacked sensor layers, wherein the top sensor layer includes a pixel array for capturing images, the lower sensor layer is used to process image data, including an analog-to-digital conversion layer, logic circuitry for extracting features and determining object depth information, and convolutional neural networks for improving processing efficiency.

Benefits of technology

Through layered processing and feature extraction, the efficiency and accuracy of data processing are improved, power consumption and delay are reduced, and the tracking accuracy and robustness of the sensor device are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115291388B_ABST
    Figure CN115291388B_ABST
Patent Text Reader

Abstract

The present application relates to a sensor assembly and a head-mounted display. A sensor assembly and a head-mounted display for determining one or more features of a local area are proposed. The sensor assembly includes a plurality of stacked sensor layers. A first sensor layer located on the top of the sensor assembly in the plurality of stacked sensor layers includes an array of pixels. The top sensor layer can be configured to capture one or more images of light reflected from one or more objects in the local area. The sensor assembly also includes one or more sensor layers located below the top sensor layer. The one or more sensor layers can be configured to process data related to the captured one or more images. Multiple sensor assemblies can be integrated into an artificial reality system such as a head-mounted display.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of an application filed on July 24, 2018, with application number 201810821296.1 and invention name “Sensor assembly and head-mounted display”.

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims the benefit of U.S. Provisional Patent Application Serial No. 62 / 536,605, filed on July 25, 2017, which is hereby incorporated by reference in its entirety. Technical Field

[0004] The present disclosure relates generally to implementations of sensor devices, and in particular to a sensor system that includes a plurality of stacked sensor layers and that may be part of an artificial reality system. Background Art

[0005] Artificial reality systems such as head mounted display (HMD) systems employ complex sensor devices (cameras) for capturing features of objects in the surrounding area in order to provide a satisfactory user experience. A limited number of conventional sensor devices can be implemented in HMD systems and used, for example, for eye tracking, hand tracking, body tracking, scanning the surrounding area with a wide field of view, and the like. Most of the time, conventional sensor devices capture a large amount of information from the surrounding area. Due to processing a large amount of data, conventional sensor devices may be easily saturated, thereby negatively affecting the processing speed. In addition, due to performing computationally intensive operations, conventional sensor devices employed in artificial reality systems consume a lot of power while having very large delays. Summary of the invention

[0006] A sensor assembly is proposed herein for determining one or more features of a local area surrounding part or all of the sensor assembly. The sensor assembly includes a plurality of stacked sensor layers, i.e., sensor layers stacked one on top of each other. The first sensor layer located on the top of the sensor assembly among the plurality of stacked sensor layers may be implemented as a photodetector layer and include an array of pixels. The top sensor layer may be configured to capture one or more images of light reflected from one or more objects in the local area. The sensor assembly also includes one or more sensor layers located below the photodetector layer. The one or more sensor layers may be configured to process data related to the captured one or more images for determining the one or more features of the local area, such as depth information of the one or more objects, an image classifier, etc.

[0007] A head mounted display (HMD) may further integrate multiple sensor assemblies. The HMD displays content to a user wearing the HMD. The HMD may be part of an artificial reality system. The HMD further includes an electronic display, at least one illumination source, and an optical assembly. The electronic display is configured to emit image light. The at least one illumination source is configured to illuminate a local area with light, which is captured by at least one sensor assembly of the multiple sensor assemblies. The optical assembly is configured to direct the image light to an eye box of the HMD corresponding to a user's eye position. The image light may include depth information of the local area determined by the at least one sensor assembly based in part on processed data related to one or more captured images.

[0008] This application also involves the following aspects:

[0009] 1). A sensor assembly, comprising:

[0010] a first sensor layer of the plurality of stacked sensor layers, the first sensor layer comprising an array of pixels configured to capture one or more images of light reflected from one or more objects in a local area; and

[0011] One or more sensor layers of the plurality of stacked sensor layers, the one or more sensor layers being located in the sensor assembly below the first sensor layer, the one or more sensor layers being configured to process data associated with the captured one or more images.

[0012] 2). A sensor assembly according to 1), wherein the one or more sensor layers include an analog-to-digital conversion (ADC) layer having a logic circuit for converting an analog intensity value related to the reflected light captured by the pixel into a digital value.

[0013] 3). The sensor assembly according to 2) further comprises an interface connection between each pixel in the array and the logic circuit of the analog-to-digital conversion layer.

[0014] 4). The sensor assembly according to 1), wherein the first sensor layer comprises a non-silicon photodetection material.

[0015] 5). The sensor assembly according to 4), wherein the non-silicon photodetector material is selected from the group consisting of: quantum dot (QD) photodetection material and organic photonic film (OPF) photodetection material.

[0016] 6). The sensor assembly according to 1), wherein at least one of the one or more sensor layers comprises a logic circuit configured to extract one or more features from the captured one or more images.

[0017] 7). The sensor assembly of 1), wherein at least one of the one or more sensor layers comprises a logic circuit configured to determine depth information of the one or more objects based in part on processed data.

[0018] 8). The sensor assembly according to 1), wherein at least one of the one or more sensor layers comprises a convolutional neural network (CNN).

[0019] 9). A sensor assembly according to 8), wherein the plurality of network weights in the convolutional neural network are trained for at least one of: classification of one or more captured images; and identification of at least one feature in one or more captured images.

[0020] 10) A sensor assembly according to 9), wherein the trained network weights are applied to data associated with an image in one or more captured images to determine a classifier for the image.

[0021] 11). A sensor component according to 8), wherein the convolutional neural network comprises an array of memristors configured to store trained network weights.

[0022] 12) The sensor assembly according to 1) further comprises:

[0023] An optical assembly is positioned on top of the first sensor layer, the optical assembly being configured to direct reflected light toward the array of pixels.

[0024] 13). A sensor assembly according to 12), wherein the optical assembly comprises a layer of one or more chips stacked on top of the first sensor layer, each chip of the optical assembly being implemented as a glass chip and configured as a separate optical element of the optical assembly.

[0025] 14). A head mounted display (HMD), comprising:

[0026] an electronic display configured to emit image light;

[0027] an illumination source configured to illuminate a local area with light;

[0028] A plurality of sensor assemblies, each sensor assembly of the plurality of sensor assemblies comprising:

[0029] a first sensor layer of the plurality of stacked sensor layers, the first sensor layer comprising an array of pixels configured to capture one or more images of at least a portion of light reflected from one or more objects in a local area, and

[0030] one or more sensor layers of the plurality of stacked sensor layers, the one or more sensor layers being located below the sensor layer, the one or more sensor layers being configured to process data associated with the captured one or more images; and

[0031] an optical assembly configured to direct image light to an eye box of the head mounted display corresponding to a position of a user's eyes,

[0032] Wherein, the depth information of the local area is determined by at least one sensor component of the plurality of sensor components based in part on the processed data.

[0033] 15). The head mounted display according to 14), further comprising:

[0034] A controller is configured to dynamically activate a first subset of the plurality of sensor assemblies and deactivate a second subset of the plurality of sensor assemblies.

[0035] 16). A head-mounted display according to 15), wherein the controller is further configured to instruct at least one sensor component in the second subset to transmit information about one or more tracked features to at least one other sensor in the first component before being deactivated.

[0036] 17). The head mounted display according to 14), further comprising a controller, wherein the controller is coupled to a sensor component among the plurality of sensor components, wherein:

[0037] The sensor assembly is configured to send first data of a first resolution to the controller using a first frame rate, the first data being related to an image captured by the sensor assembly at a first moment in time,

[0038] The controller is configured to transmit information regarding one or more characteristics obtained based on the first data received from the sensor assembly using the first frame rate, and

[0039] The sensor assembly is further configured to send second data at a second resolution lower than the first resolution to the controller using a second frame rate higher than the first frame rate, the second data being related to another image captured by the sensor assembly at a second time.

[0040] 18). A head mounted display according to 14), wherein each sensor assembly further comprises an interface connection between each pixel in the array and a logic circuit of at least one of the one or more sensor layers.

[0041] 19). A head-mounted display according to 14), wherein at least one of the one or more sensor layers comprises a logic circuit configured to extract one or more features from the captured one or more images.

[0042] 20). A head-mounted display according to 14), wherein at least one of the one or more sensor layers comprises a convolutional neural network (CNN) based on an array of memristors for storing trained network weights. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1A is a diagram of a head mounted display (HMD) according to one or more implementations.

[0044] Figure 1B According to one or more embodiments Figure 1A A cross section of the front rigid body of the HMD.

[0045] Figure 2 According to one or more embodiments, it may be Figure 1A Cross-sectional view of a stacked sensor system having multiple stacked sensor layers in a portion of an HMD.

[0046] Figure 3 According to one or more embodiments, it may be Figure 2 Detailed view of multiple stacked sensor layers of a portion of a stacked sensor system in FIG.

[0047] Figure 4 According to one or more embodiments, it may be Figure 2 An example sensor architecture consisting of coupled sensor layers as part of a stacked sensor system in FIG.

[0048] Figure 5 According to one or more embodiments, it may be Figure 2 An example of a memristor array-based neural network as part of a stacked sensor system in FIG.

[0049] Figure 6 is an example of a host-sensor closed-loop system according to one or more embodiments.

[0050] Figure 7 is a block diagram of an HMD system in which a console operates according to one or more implementations.

[0051] The accompanying drawings depict embodiments of the present disclosure for purposes of illustration only. Those skilled in the art will readily recognize from the following description that alternative embodiments of the structures and methods shown herein may be employed without departing from the principles or benefits of the present disclosure described herein. DETAILED DESCRIPTION

[0052] Embodiments of the present disclosure may include an artificial reality system or may be implemented in combination with an artificial reality system. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user, which may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured (e.g., real-world) content. Artificial reality content may include video, audio, tactile feedback, or some combination thereof, and any of the content may be presented in a single channel or in multiple channels (such as stereoscopic video that produces a three-dimensional effect for the viewer). In addition, in some embodiments, artificial reality may also be associated with, for example, applications, products, accessories, services, or some combination thereof for creating content in artificial reality and / or otherwise used for artificial reality (e.g., performing activities in artificial reality). An artificial reality system that provides artificial reality content can be implemented on a variety of platforms, including a head-mounted display (HMD) connected to a host computer system, a stand-alone HMD, a near-eye display (NED), a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more viewers.

[0053] A stacked sensor system for determining various characteristics of an environment is proposed herein, which can be integrated into an artificial reality system. The stacked sensor system includes a plurality of stacked sensor layers. Each sensor layer in the plurality of stacked sensor layers can represent a signal processing layer for performing a specific signal processing function. Analog sensor data related to the intensity of light reflected from the environment can be captured by a photodetector layer located on top of the stacked sensor system. The captured analog sensor data can be converted from the analog domain to the digital domain, for example, via an analog-to-digital conversion (ADC) layer located below the photodetector layer. The digital sensor data can then be provided to at least one signal processing layer of the stacked sensor system located below the ADC layer. The at least one signal processing layer will process the digital sensor data to determine one or more characteristics of the environment.

[0054] In some embodiments, multiple stacked sensor systems are integrated into an HMD. These stacked sensor systems (e.g., sensor devices) can capture data describing various features of the environment, including depth information of a local area surrounding part or all of the HMD. The HMD displays content to a user wearing the HMD. The HMD can be part of an artificial reality system. The HMD also includes an electronic display and an optical assembly. The electronic display is configured to emit image light. The optical assembly is configured to direct the image light to an eye box of the HMD corresponding to the position of the user's eyes. The image light may include depth information of a local area determined by at least one of the multiple stacked sensor systems.

[0055] In some other embodiments, multiple stacked sensor systems can be integrated into a glasses-like platform representing a NED. The NED can be part of an artificial reality system. The NED presents media to a user. Examples of media presented by the NED include one or more images, videos, audio, or a combination thereof. The NED also includes an electronic display and an optical component. The electronic display is configured to emit image light. The optical component is configured to guide the image light to an eye box of the NED corresponding to the position of the user's eyes. The image light may include depth information of a local area determined by at least one of the multiple stacked sensor systems.

[0056] Figure 1A 1 is a diagram of an HMD 100 according to one or more embodiments. The HMD 100 may be part of an artificial reality system. In embodiments describing an AR system and / or an MR system, portions of a front side 102 of the HMD 100 are at least partially transparent in the visible band (~380nm to 750nm), and portions of the HMD 100 between the front side 102 of the HMD 100 and the user's eyes are at least partially transparent (e.g., a partially transparent electronic display). The HMD 100 includes a front rigid body 105, a ribbon 110, and a reference point 115.

[0057] The front rigid body 105 includes one or more electronic display elements (in Figure 1A ), one or more integrated eye tracking systems (not shown in Figure 1A ), an inertial measurement unit (IMU) 120, one or more position sensors 125, and a reference point 115. Figure 1AIn the illustrated embodiment, the position sensor 125 is located within the IMU 120, and both the IMU 120 and the position sensor 125 are not visible to the user of the HMD 100. The IMU 120 is an electronic device that generates IMU data based on measurement signals received from one or more of the position sensors 125. The position sensor 125 generates one or more measurement signals in response to the movement of the HMD 100. Examples of the position sensor 125 include: one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor that detects movement, a type of sensor used to perform error correction on the IMU 120, or some combination thereof. The position sensor 125 can be located outside the IMU 120, inside the IMU 120, or some combination thereof.

[0058] HMD 100 includes a distributed network of sensor devices (cameras) 130, which may be embedded in the front rigid body 105. It should be noted that although Figure 1A Not shown, at least one sensor device 130 may also be embedded in the strip 110. Each sensor device 130 may be implemented as a relatively small sized camera. A distributed network of sensor devices 130 may replace multiple large conventional cameras. In some embodiments, each sensor device 130 of the distributed network embedded in the HMD 100 is implemented as a microchip camera with a predetermined limited resolution, for example, each sensor device 130 may include an array of 100 x 100 pixels or an array of 200 x 200 pixels. In some embodiments, each sensor device 130 in the distributed network has a field of view that does not overlap with the field of view of any other sensor device 130 integrated into the HMD 100. This is in contrast to the overlapping field of view of a large conventional camera, which may cause a large amount of overlapping data to be captured from the surrounding area. The HMD 100 may also include an imaging aperture (in Figure 1A (not shown). The sensor device 130 may capture light reflected from the surrounding area through an imaging aperture.

[0059] It should be noted that it would be impractical for each sensor device 130 in a distributed network of sensor devices 130 to have its own direct link (bus) to a central processing unit (CPU) or controller 135 embedded in the HMD 100. Instead, each individual sensor device 130 can be connected in a scalable manner via a shared bus (in Figure 1AThe controller 135 is coupled to the sensor device 130 (not shown) to provide an expandable network of sensor devices 130 embedded in the HMD 100. The expandable network of sensor devices 130 can be considered a redundant system. Taken together, the sensor devices 130 cover a much larger field of view (e.g., 360 degrees) than the field of view typically deployed by a large conventional camera (e.g., 180 degrees). The wider field of view obtained by the sensor device 130 provides increased robustness.

[0060] It should be noted that it is not necessary to always keep all sensor devices 130 embedded in the HMD 100 active (i.e., turned on). In some embodiments, the controller 135 is configured to dynamically activate a first subset of the sensor devices 130 and deactivate a second subset of the sensor devices 130, for example, based on a particular situation. In one or more embodiments, the controller 135 may deactivate a portion of the sensor devices 130 depending on the specific simulation running on the HMD 100. For example, after locating a preferred portion of the environment for scanning, certain sensor devices 130 may remain activated, while other sensor devices 130 may be deactivated in order to save power consumed by the distributed network of sensor devices 130.

[0061] A sensor device 130 or a group of sensor devices 130 may, for example, track one or more moving objects and specific features about the one or more moving objects during a period of time. For example, based on instructions from the controller 135, the features about the moving object obtained during the period of time may then be transmitted to another sensor device 130 or another group of sensor devices 130 for continued tracking during a subsequent period of time. For example, the HMD 100 may use features extracted in the scene as "ground markers" for user positioning and head posture tracking in a three-dimensional world. Features associated with the user's head may be extracted by, for example, one sensor device 130 at one moment. At the next moment, the user's head may move, and another sensor device 130 may be activated to locate the same feature for performing head tracking. The controller 135 may be configured to predict which new sensor device 130 may potentially capture the same feature of a moving object (e.g., a user's head). In one or more embodiments, the controller 135 may utilize IMU data obtained by the IMU 120 to perform a rough prediction. In this case, information about the tracking features may be transmitted from one sensor device 130 to another sensor device 130, for example, based on the rough prediction. Depending on the specific tasks being performed at a specific moment, the number of active sensor devices 130 may be dynamically adjusted (e.g., based on instructions from the controller 135). In addition, one sensor device 130 may perform extraction of specific features of the environment and provide the extracted feature data to the controller 135 for further processing and transmission to another sensor device 130. Thus, each sensor device 130 in the distributed network of sensor devices 130 may process a limited amount of data. In contrast, conventional sensor devices integrated into HMD systems typically perform continuous processing of large amounts of data, which consumes more power.

[0062] In some embodiments, each sensor device 130 integrated into the HMD 100 can be configured for a specific type of processing. For example, at least one sensor device 130 can be customized to track various features of the environment, such as determining sharp corners, hand tracking, etc. In addition, each sensor device 130 can be customized to detect one or more specific landmark features while ignoring other features. In some embodiments, each sensor device 130 can perform early processing that provides information related to a specific feature (e.g., the coordinates of the feature and a description of the feature). To support early processing, certain processing circuitry can be incorporated into the sensor device 130, such as in conjunction with Figures 2 to 5130 . The sensor device 130 may then transmit the data obtained based on the early processing, for example, to the controller 135 , thereby reducing the amount of data communicated between the sensor device 130 and the controller 135 . In this way, the frame rate of the sensor device 130 is increased while retaining the bandwidth requirements between the sensor device 130 and the controller 135 . In addition, because the partial processing is performed at the sensor device 130 , the power consumption and processing delay of the controller 135 may be reduced, and the computational burden of the controller 135 may be reduced and distributed to one or more sensor devices 130 . Another advantage of the partial and early processing of instructions at the sensor device 130 includes a reduced use of the internal memory (in the internal memory of the controller 135 ) of the sensor device 130 . Figure 1A In addition, the power consumption at the controller 135 can be reduced because fewer memory accesses result in lower power consumption.

[0063] In an embodiment, the sensor device 130 may include an array of 100 x 100 pixels or an array of 200 x 200 pixels coupled to a processing circuit that is customized to extract, for example, up to 10 features of the environment surrounding part or all of the HMD 100. In another embodiment, the processing circuit of the sensor device 130 may be customized to operate as a neural network that is trained to track, for example, up to 20 joint positions of a user's hand (which may be required to perform accurate hand tracking). In yet another embodiment, at least one sensor device 130 may be employed for facial tracking, where movements of the user's mouth and face may be captured. In this case, the at least one sensor device 130 may face downward to facilitate tracking of the user's facial features.

[0064] It should be noted that each sensor device 130 integrated into the HMD 100 can provide a level of signal-to-noise ratio (SNR) that is higher than a threshold level defined for that sensor device 130. Because the sensor device 130 is customized for a specific task, the sensitivity of the customized sensor device 130 can be improved compared to a conventional camera. It should also be noted that the distributed network of sensor devices 130 is a redundant system, and the sensor device 130 of the distributed network that produces a preferred SNR level can be selected (e.g., by the controller 135). In this way, the tracking accuracy and robustness of the distributed network of sensor devices 130 can be greatly improved. Each sensor device 130 can also be configured to operate in an expanded wavelength range (e.g., in the infrared and / or visible spectrum).

[0065] In some embodiments, the sensor device 130 includes a photodetector layer having an array of silicon-based photodiodes. In alternative embodiments, the photodetector layer of the sensor device 130 can be implemented using non-silicon-based materials and technologies, which can provide improved sensitivity and wavelength range. In one embodiment, the photodetector layer of the sensor device 130 is based on an organic photonic film (OPF) photodetector material suitable for capturing photodetector materials with wavelengths greater than 1000nm. In another embodiment, the photodetector layer of the sensor device 130 is based on a quantum dot (QD) photodetector material. The QD-based sensor device 130 can, for example, be suitable for integration into AR systems and applications related to outdoor environments with low visibility (e.g., at night). Most of the available ambient light is then located in the long-wavelength non-visible range, for example, between about 1 μm and 2.5 μm, that is, in the short-wave infrared range. The photodetection layer of the sensor device 130 implemented based on the optimized QD film can detect both visible light and short-wave infrared light, while the silicon-based film may only be sensitive to wavelengths of light around about 1.1 μm.

[0066] In some embodiments, a controller 135 of the sensor devices 130 embedded in the front rigid body 105 and coupled to the distributed sensor network is configured to combine captured information from the sensor devices 130. The controller 135 can be configured to appropriately integrate data associated with different features collected by different sensor devices 130. In some embodiments, the controller 135 determines depth information of one or more objects in a local area surrounding part or all of the HMD 100 based on data captured by one or more of the sensor devices 130.

[0067] Figure 1B According to one or more embodiments Figure 1A 100. The front rigid body 105 includes a sensor device 130, a controller 135 coupled to the sensor device 130, an electronic display 155, and an optical assembly 160. The electronic display 155 and the optical assembly 160 together provide image light to an eye box 165. The eye box 165 is the area in the space occupied by the user's eyes 170. For purposes of illustration, Figure 1B The cross section 150 is shown associated with a single eye 170, while another optical assembly 160 separate from the optical assembly 160 provides altered image light to the user's other eye.

[0068] The electronic display 155 emits image light toward the optical assembly 160. In various embodiments, the electronic display 155 may include a single electronic display or multiple electronic displays (e.g., a display for each eye of the user). Examples of electronic displays 155 include: a liquid crystal display (LCD), an organic light emitting diode (OLED) display, an inorganic light emitting diode (ILED) display, an active matrix organic light emitting diode (AMOLED) display, a transparent organic light emitting diode (TOLED) display, some other display, a projector, or some combination thereof. The electronic display 155 also includes an aperture, a Fresnel lens, a convex lens, a concave lens, a diffractive element, a waveguide, a filter, a polarizer, a diffuser, a fiber taper, a reflective surface, a polarized reflective surface, or any other suitable optical element that affects the image light emitted from the electronic display 155. In some embodiments, the electronic display 155 may have one or more coatings such as an anti-reflective coating.

[0069] The optical assembly 160 receives image light emitted from the electronic display 155 and directs the image light to the eye box 165 of the user's eye 170. The optical assembly 160 also amplifies the received image light, corrects optical aberrations associated with the image light, and presents the corrected image light to the user of the HMD 100. In some embodiments, the optical assembly 160 includes a collimating element (lens) for collimating a beam of image light emitted from the electronic display 155. At least one optical element of the optical assembly 160 may be an aperture, a Fresnel lens, a refractive lens, a reflective surface, a diffractive element, a waveguide, a filter, or any other suitable optical element that affects the image light emitted from the electronic display 155. In addition, the optical assembly 160 may include a combination of different optical elements. In some embodiments, one or more of the optical elements in the optical assembly 160 may have one or more coatings such as an anti-reflective coating, a dichroic coating, and the like. The amplification of the image light by the optical assembly 160 allows the elements of the electronic display 155 to be physically smaller, lighter, and consume less power than larger displays. In addition, magnification can increase the field of view of the displayed media. For example, the field of view of the displayed media is such that the displayed media is presented using substantially all (e.g., 110 degrees diagonally) and in some cases the entire field of view of the user. In some embodiments, optical assembly 160 is designed such that its effective focal length is greater than the spacing to electronic display 155, which magnifies the image light projected by electronic display 155. In addition, in some embodiments, the amount of magnification can be adjusted by adding or removing optical elements.

[0070] In some embodiments, the front rigid body 105 further includes an eye tracking system (in Figure 1B165), the eye tracking system determines eye tracking information for the user's eye 170. The determined eye tracking information may include information about the position (including orientation) of the user's eye 170 in the eye box 165, that is, information about the angle of the eye gaze. In one embodiment, the eye tracking system illuminates the user's eye 170 with structured light. The eye tracking system may determine the position of the user's eye 170 based on deformations in the structured light pattern reflected from the surface of the user's eye and captured by a camera of the eye tracking system. In another embodiment, the eye tracking system determines the position of the user's eye 170 based on the amplitude of image light captured at multiple times.

[0071] In some embodiments, the front rigid body 105 further includes a zoom module (in Figure 1B 15). The varifocal module can adjust the focus of one or more images displayed on the electronic display 155 based on the eye tracking information obtained from the eye tracking system. In one embodiment, the varifocal module adjusts the focus of the displayed image and alleviates the vergence accommodation conflict by adjusting the focal distance of the optical assembly 160 based on the determined eye tracking information. In other embodiments, the varifocal module adjusts the focus of the displayed image by foveated rendering the one or more images based on the determined eye tracking information.

[0072] Figure 2 is a cross-sectional view of a sensor assembly 200 having multiple stacked sensor layers according to one or more embodiments. The sensor assembly 200 may be Figure 1A 100. In some embodiments, the sensor assembly 200 includes multiple silicon layers stacked on top of each other. In alternative embodiments, at least one of the multiple stacked sensor layers in the sensor assembly 200 is implemented based on a non-silicon photodetection material. The top sensor layer in the sensor assembly 200 can be customized for photodetection and can be referred to as a photodetector layer 205. The photodetector layer 205 can include a two-dimensional array of pixels 210. Each pixel 210 of the photodetector layer 205 can be, for example, bonded via copper (in Figure 2 The processing circuit 215 (not shown) is directly coupled to the processing layer 220, which is located below the photodetector layer 205 within the sensor assembly 200.

[0073] As in Figure 2The stacking of multiple sensor layers (wafers) shown in Figure 2 allows copper bonding between the photodetector layer 205 and the processing layer 220 at each pixel resolution. By placing the two wafers face to face, the copper pad connection from one wafer in the sensor assembly 200 to the other wafer in the sensor assembly 200 can be generated at each pixel level, that is, the electrical signal 225 corresponding to a single pixel can be sent from the photodetector layer 205 to the processing circuit 215 of the processing layer 220. In one or more embodiments, the interconnection between the processing layer 220 and at least one additional layer of the multiple stacked structures of the sensor assembly 200 can be achieved using, for example, "through silicon via" (TSV) technology. Due to the geometric size of the TSV (e.g., about 10 μm), the interconnection between the processing layer 220 and the at least one additional layer of the sensor assembly 200 is not at the pixel level, but can still be very dense.

[0074] In some embodiments, by adopting wafer scaling, a small-sized sensor component 200 can be effectively implemented. For example, the wafer of the photodetector layer 205 can be implemented using, for example, a 45nm process technology, while the wafer of the processing layer 220 can be implemented using a more advanced process technology (for example, a 28nm process technology). Because the transistors of the 28nm process technology occupy a very small area, a large number of transistors can be assembled into the small area of ​​the processing layer 220. In an exemplary embodiment, the sensor component 200 can be implemented as a 1mm x 1mm x 1mm cube with a power consumption of approximately 10mW. In contrast, a conventional sensor (camera) includes a processing circuit and a photodetector pixel array implemented on a single silicon layer, and the total sensor area is determined to be the sum of the areas of all functional blocks. Without such a large area, the sensor component 200 can be realized in a small area. Figure 2 Due to the benefits of vertical stacking in the embodiment shown in , conventional sensors occupy a much larger area than sensor assembly 200 .

[0075] Figure 3 is a detailed view of a sensor assembly 300 including a plurality of stacked sensor layers according to one or more embodiments. The sensor assembly 300 may be Figure 1A Embodiments of the sensor device 130 and Figure 2. In some embodiments, the photodetector layer 305 can be positioned on top of the sensor assembly 300 and can include an array of pixels 310, such as a two-dimensional array of photodiodes. Because the processing circuitry of the sensor assembly 300 can be integrated into other layers below the photodetector layer 305, the photodetector layer 305 can be customized for photodetection only. Therefore, the area of ​​the photodetector layer 305 can be relatively small, and the photodetector layer 305 can consume a limited amount of power. In some embodiments, an ADC layer 315 customized to convert an analog signal (e.g., the intensity of light captured by the photodetector layer 305) into a digital signal can be placed immediately below the photodetector layer 305. The ADC layer 315 can be configured to convert (e.g., through its processing circuitry or ADC logic circuitry, in Figure 3 The ADC layer 315 may also include a memory (not shown in detail) corresponding to, for example, a digital value of an image frame data. Figure 3 ), which is used to store the digital value obtained after the conversion.

[0076] In some embodiments, a feature extraction layer 320 having processing circuitry customized for feature extraction can be placed immediately below the ADC layer 315. The feature extraction layer 320 can also include a memory for storing, for example, digital sensor data generated by the ADC layer 315. The feature extraction layer 320 can be configured to extract one or more features from the digital sensor data obtained by the ADC layer 315. Because the feature extraction layer 320 is customized to extract specific features, the feature extraction layer 320 can be efficiently designed to occupy a small area size and consume a limited amount of power. Figure 4 More details are provided regarding the feature extraction layer 320 .

[0077] In some embodiments, a convolutional neural network (CNN) layer 325 can be placed immediately below the feature extraction layer 320. The neural network logic of the CNN layer 325 can be trained and optimized for specific input data, for example, data having information about a specific feature or set of features obtained by the feature extraction layer 320. Because the input data is fully anticipated, the neural network logic of the CNN layer 325 can be efficiently implemented and customized for specific types of feature extraction data, thereby achieving reduced processing latency and reduced power consumption.

[0078] In some embodiments, the CNN layer 325 is designed to perform image classification and recognition applications. Training of the neural network logic circuit of the CNN layer 325 can be performed offline, and the network weights in the neural network logic circuit of the CNN layer 325 can be trained prior to utilizing the CNN layer 325 for image classification and recognition. In one or more embodiments, the CNN layer 325 is implemented to perform inference, i.e., applying the trained network weights to an input image to determine an output, e.g., an image classifier. Compared to designing a general-purpose CNN architecture, the CNN layer 325 can be implemented as a customized and dedicated neural network, and can be designed for preferred levels of power consumption, area size, and efficiency (computational speed). In combination Figure 5 Details regarding a specific implementation of CNN layer 325 are provided.

[0079] In some embodiments, each layer 305, 315, 320, 325 in the sensor assembly 300 that is customized for a specific processing task can be implemented using silicon-based technology. Alternatively, at least one of the layers 305, 315, 320, 325 can be implemented based on non-silicon photodetection materials, for example, OPF photodetection materials and / or QD photodetection materials. In some embodiments, instead of a silicon-based photodetector layer 305 including an array of photodiode-based pixels 310, a non-silicon photodetector layer 330 can be placed on top of the sensor assembly 300. In one embodiment, the non-silicon photodetector layer 330 is implemented as a photodetector layer of QD photodetection material and can be referred to as a QD photodetector layer. In another embodiment, the non-silicon photodetector layer 330 is implemented as a photodetector layer of OPF photodetection material and can be referred to as an OPF photodetector layer. In yet another embodiment, more than one photodetector layer may be used in the sensor assembly 300 for photodetection, for example, at least one silicon-based photodetector layer 305 and at least one non-silicon-based photodetector layer 330 .

[0080] In some embodiments, direct copper bonding may be used for inter-layer coupling between the photodetector layer 305 and the ADC layer 315. Figure 3 As shown in FIG. 3 , copper pad connections 335 may be used as an interface between the pixels 310 in the photodetector layer 305 and the processing circuitry (ADC logic circuitry) of the ADC layer 315. For example, where the photodetector layer 305 is implemented as a 20M pixel camera, up to about 20 mega-pixel copper pad connections 335 may be implemented between the photodetector layer 305 and the processing circuitry of the ADC layer 315. It should be noted that the pitch of the photodetector layer 305 may be relatively small, for example, between about 1 μm and 2 μm.

[0081] In some embodiments, as discussed, interconnections between sensor layers located below the photodetector layer 305 in the sensor assembly 300 can be implemented using, for example, TSV technology. Figure 3 As shown in FIG. 3 , a TSV interface 340 may interconnect the ADC logic / memory of the ADC layer 315 with the feature extraction logic of the feature extraction layer 320. The TSV interface 340 may provide digital sensor data obtained by the ADC layer 315 to the feature extraction layer 320 for extracting one or more specific features, wherein the digital sensor data may relate to image data captured by the photodetector layer 305. Similarly, another TSV interface 345 may be used to interconnect the feature extraction logic of the feature extraction layer 320 with the neural network logic of the CNN layer 325. The TSV interface 345 may provide the feature extraction data obtained by the feature extraction logic of the feature extraction layer 320 as input to the neural network logic of the CNN layer 325 for, for example, image classification and / or recognition.

[0082] In some embodiments, the optical assembly 350 can be positioned on top of the silicon-based photodetector layer 305 (or the non-silicon-based photodetector layer 330). The optical assembly 350 can be configured to direct at least a portion of light reflected from one or more objects in a local area surrounding the sensor assembly 300 to the pixels 310 of the silicon-based photodetector layer 305 (or the sensor elements of the non-silicon-based photodetector layer 330). In some embodiments, the optical assembly 350 can be positioned on top of the silicon-based photodetector layer 305 (or the non-silicon-based photodetector layer 330). Figure 3 The optical assembly 350 may be implemented as a glass wafer and represent an individual lens element of the optical assembly 350. In one or more embodiments, a polymer-based material may be molded on the top and / or bottom surface of the glass wafer to serve as a reflective surface for the individual lens elements of the optical assembly 350. This technique may be referred to as wafer-level optics. In addition, the spacers (in Figure 3 ) may be included between a pair of adjacent glass wafers (layers) in the optical assembly 350, thereby adjusting the space between the adjacent glass layers.

[0083] In some embodiments, all glass wafers of the optical assembly 350 and all silicon wafers of the layers 305, 315, 320, 325 can be fabricated and stacked together before cutting each individual sensor lens unit from the wafer stack, thereby obtaining one instantiation of the sensor assembly 300. Once fabrication is complete, each cube obtained from the wafer stack becomes a complete, fully functional camera, e.g. Figure 3It should be understood that the sensor assembly 300 does not require any plastic housing for the layers 305, 315, 320, 325, which facilitates the implementation of the sensor assembly 300 as a cube with a predetermined volume size that is smaller than the threshold volume.

[0084] In some embodiments, when a non-silicon-based photodetector layer 330 (e.g., a QD photodetector layer or an OPF photodetector layer) is part of the sensor assembly 300, the non-silicon-based photodetector layer 330 can be directly coupled to the ADC layer 315. The electrical connections between the sensor elements (pixels) in the non-silicon-based photodetector layer 330 and the ADC layer 315 can be made as copper pads. In this case, the non-silicon-based photodetector layer 330 can be deposited on the ADC layer 315 after all other layers 315, 320, 325 are stacked. After the non-silicon-based photodetector layer 330 is deposited on the ADC layer 315, the optical assembly 350 is applied on top of the non-silicon-based photodetector layer 330.

[0085] Figure 4 4 is an example sensor architecture 400 of coupled sensor layers in a stacked sensor assembly according to one or more implementations. The sensor architecture 400 may include a sensor circuit 405 coupled to a feature extraction circuit 410, for example, via a TSV interface. The sensor architecture 400 may be implemented as Figure 3 In some embodiments, the sensor circuit 405 may be a part of at least two sensor layers of the sensor assembly 300 in the embodiment. In some embodiments, the sensor circuit 405 may be a part of the photodetector layer 305 and the ADC layer 315, and the feature extraction circuit 410 may be a part of the feature extraction layer 320.

[0086] The sensor circuit 405 may acquire and pre-process the sensor data and then provide the acquired sensor data to the feature extraction circuit 410, for example, via a TSV interface. The sensor data may correspond to an image captured by a two-dimensional pixel array 415 (e.g., an M x N array of digital pixels), where M and N are integers of the same or different values. It should be noted that the two-dimensional pixel array 415 may be Figure 3 A portion of the photodetector layer 305 of the sensor assembly 300. In addition, the two-dimensional pixel array 415 may include pixel-by-pixel AND ADC logic circuits (in Figure 4 The ADC logic circuit may be Figure 3430. The pixel data 420 from the multiplexer 425 may include digital sensor data related to the captured image. The pixel data 420 may be stored in a line buffer 430 and provided to the feature extraction circuit 410, for example, via a TSV interface for extracting one or more specific features. It should be noted that the full frame read from the multiplexer 425 may be output via a high-speed mobile industry processor interface (MIPI) 435 for generating a raw stream output 440.

[0087] Feature extraction circuitry 410 may determine one or more features from the captured image represented by pixel data 420. Figure 4 In an exemplary embodiment of the invention, the feature extraction circuit 410 includes a point / feature / key point (KP) / mapped event block 445, a convolution engine 450, a centroid estimation block 455, and a threshold detection block 460. The convolution engine 450 can process the pixel data 420 cached in the online buffer 430 by, for example, applying a 3x3 convolution using various filter coefficients (kernels) 465 (such as filter coefficients 465 for Gaussian filtering, first order derivatives, second order derivatives, etc.). The filtered data from the convolution engine 450 can be further fed into the threshold detection block 460, where specific key features or events in the captured image can be detected based on, for example, filtering and thresholding. The location of the key features / events determined by the threshold detection block 460 can be written to the mapping block 445. The mapping of one or more key features / events can be uploaded to, for example, a host computer (on a Figure 4 Another mapping of one or more key features / events can also be written from the host to the mapping block 445. In an embodiment, based on the structured light principle, the sensor architecture 400 can be used to measure the depth information in the scene. In this case, the laser point centroid can be extracted from the centroid estimation block 455 and the centroid can be written to the mapping block 445.

[0088] It should be understood that in Figure 4 The sensor architecture 400 shown in FIG. 4 represents an exemplary implementation. Other implementations of the sensor circuit 405 and / or the feature extraction circuit 410 may include different and / or additional processing blocks.

[0089] Figure 5 An example of a neural network 500 based on a memristor array according to one or more embodiments is shown. The neural network 500 may be a Figure 3 325 of the sensor assembly 300. In some embodiments, the neural network 500 represents a CNN that is trained and used for certain processing based on a machine learning algorithm, such as for image classification and / or recognition. It should be noted that Figure 4The feature extraction circuit 410 and the neural network 500 may coexist in a smart sensor system (eg Figure 3 sensor assembly 300), because more traditional computer vision algorithms (e.g., for depth extraction) can be implemented on the feature extraction circuit 410.

[0090] In some embodiments, the neural network 500 can be optimized for neuromorphic computing with a memristor crossbar matrix suitable for performing vector-matrix multiplication. Learning in the neural network 500 is represented according to a set of parameters including a conductance value G=G at the crossbar matrix of the neural network 500. n,m (n=1, 2, ... N; m=1, 2, ... M) and resistance value R S (For example, M resistors with values ​​of r S Operational amplifier (op-amp) 502 and its associated resistor r S They are used as the output driver of each column of memristor elements and the weighting coefficient of each column respectively.

[0091] In some embodiments, instead of retrieving parameters from, for example, a dynamic random access memory (DRAM), the parameters in the form of conductance and resistance values ​​are directly available at the crossbar matrix points of the neural network 500 and can be used directly during calculations (e.g., during vector-matrix multiplication). Figure 5 The neural network 500 of the memristor crossbar switch matrix shown in FIG. 5 can have dual functions, that is, the neural network 500 can be used as a memory storage device and a computing device. Therefore, the neural network 500 can be referred to as "memory-implemented computing", which has a preferred power consumption level, area size, and computing efficiency. The neural network 500 can effectively replace the combination of a memory storage device and a CPU, which makes the neural network 500 suitable for being effectively implemented as Figure 3 A portion of a CNN layer 325 of the sensor assembly 300 .

[0092] The initial weight G of the neural network 500 n,m Writing can be done via input 505 (which can represent a matrix input) having values ​​organized in, for example, N rows and M columns. In one or more embodiments, matrix input 505 can correspond to a kernel for a convolution operation. In some embodiments, input 510 can correspond to, for example, a matrix of pixels captured by photodetector layer 305 and represented by Figure 3 The ADC layer 315 and the feature extraction layer 320 of the sensor assembly 300 process the digital pixel values ​​of the image. The input 510 may include N voltage values ​​V organized as, for example 1 I 、V 2I 、V 3 I , ……V N I The digital voltage value V of the vector I During an inference operation, vector input 510 may be applied to matrix input 505. As a result of the multiplication between vector input 510 and matrix input 505 (i.e., as a result of a vector-matrix multiplication), output 515 may be obtained. Figure 5 As shown in FIG. 1 , output 515 represents a digital voltage value V O A vector of M voltage values ​​V 1 O 、V 2 O 、V 3 O , ……V M I , where V O =V I GR S The output vector V O Can be used to infer functions such as object classification.

[0093] In some embodiments, neural network 500 can be connected to the CNN layer 325 (i.e., neural network 500) via a TSV interface 345 between feature extraction layer 320 and CNN layer 325 (i.e., neural network 500). Figure 3 The photodetector layer 305, ADC layer 315, and feature extraction layer 320 of the sensor assembly 300 are effectively connected. In addition, the neural network 500 implemented based on the memristor crossbar switch matrix can avoid parallel-to-serial and serial-to-parallel conversion of data, which simplifies the implementation and improves the processing speed. Alternatively, the neural network 500 can be used for image segmentation and semantic applications, which can utilize Figure 5 The proposed method is implemented by a crossbar matrix of memristors with different sets of learned coefficients.

[0094] Figure 6 is an example of a host-sensor closed-loop system 600 according to one or more embodiments. The host-sensor closed-loop system 600 includes a sensor system 605 and a host system 610. The sensor system 605 may be Figure 1A Embodiments of the sensor device 130 in Figure 2 Embodiments of the sensor assembly 200 and / or Figure 3 The host system 610 may be an embodiment of the sensor assembly 300 in Figure 1AIn some embodiments, the sensor system 605 can obtain one or more key features of at least a portion of the environment, such as a captured image. The sensor system 605 can initially provide the one or more key features to the host system 610 as full-resolution key frames 615 at a rate of, for example, 10 frames per second. Based on the processed one or more full-resolution key frames 615, the host system 610 can be configured to predict one or more key points representing the future position of the one or more key features, such as in the next one or more image frames. The host system 610 can then provide a key point map 620 to the sensor system 605 at a rate of, for example, 10 frames per second.

[0095] After receiving the key point map 620, the sensor system 605 can activate, for example, a portion of pixels corresponding to the nearby area of ​​the predicted one or more features. The sensor system 605 can then capture and process only those light intensities partially related to the activated portion of pixels. By activating only that portion of pixels and processing only a portion of intensity values ​​captured by the activated portion of pixels, the power consumed by the sensor system 605 can be reduced. The sensor system 605 can derive one or more updated positions of the one or more key features. The sensor system 605 can then send the one or more updated positions of the one or more key features to the host system 610 as an updated key point map 625 at an increasing rate of, for example, 100 frames per second, because the updated key point map 625 includes less data than the full-resolution key frame 615. The host system 610 can then process the updated key point map 625 with a reduced amount of data compared to the full-resolution key frame 615, which provides power consumption savings at the host system 610 while also reducing computational delays at the host system 610. In this way, sensor system 605 and host system 610 form a host-sensor closed-loop system 600 with predictive sparse capture. Host-sensor closed-loop system 600 provides power savings at both sensor system 605 and host system 610 with increased communication rates between sensor system 605 and host system 610.

[0096] System environment

[0097] Figure 7 is a block diagram of one embodiment of an HMD system 700 in which a console 710 operates. The HMD system 700 may operate in an artificial reality system. Figure 7 The illustrated HMD system 700 includes an HMD 705 and an input / output (I / O) interface 715 coupled to a console 710. Figure 7An example HMD system 700 is shown that includes one HMD 705 and I / O interface 715, however in other embodiments, any number of these components may be included in the HMD system 700. For example, there may be multiple HMDs 705, each with an associated I / O interface 715, where each HMD 705 and I / O interface 715 communicates with the console 710. In alternative configurations, different and / or additional components may be included in the HMD system 700. Furthermore, in some embodiments, the HMD system 700 may include a plurality of HMDs 705, each with an associated I / O interface 715. Figure 7 The functionality described in one or more of the components shown may be achieved by combining with Figure 7 The different ways of describing are distributed among these components. For example, some or all of the functionality of console 710 is provided by HMD 705.

[0098] HMD 705 is a head-mounted display that presents content including virtual views and / or augmented views of a physical, real-world environment to a user through computer-generated elements (e.g., two-dimensional (2D) images, three-dimensional (3D) images, 2D or 3D video, sound, etc.). In some embodiments, the presented content includes audio presented via an external device (e.g., speakers and / or headphones) that receives audio information from HMD 705, console 710, or both and presents audio data based on the audio information. HMD 705 may include one or more rigid bodies that may be rigidly or non-rigidly coupled to each other. Rigid coupling between rigid bodies causes the coupled rigid bodies to act as a single rigid entity. In contrast, non-rigid coupling between rigid bodies allows the rigid bodies to move relative to each other. An embodiment of HMD 705 may be a combination of the above. Figure 1A An HMD 100 is described.

[0099] HMD 705 includes one or more sensor components 720, electronic display 725, optical component 730, one or more position sensors 735, IMU 740, optical eye tracking system 745, and optical zoom module 750. Some embodiments of HMD 705 have the following features: Figure 7 In addition, in other embodiments, by combining Figure 7 The functionality provided by the various components described may be allocated differently among the components of HMD 705 .

[0100] Each sensor assembly 720 may include multiple stacked sensor layers. A first sensor layer located on top of the multiple stacked sensor layers may include an array of pixels configured to capture one or more images of at least a portion of light reflected from one or more objects in a local area surrounding part or all of the HMD 705. At least one other sensor layer located below the first (top) sensor layer in the multiple stacked sensor layers may be configured to process data related to the captured one or more images. The HMD 705 or console 710 may dynamically activate a first subset of the sensor assemblies 720 and deactivate a second subset of the sensor assemblies 720, for example, based on an application running on the HMD 705. Therefore, at each moment, only a portion of the sensor assemblies 720 will be activated. In some embodiments, information about one or more tracking features of one or more moving objects may be transmitted from one sensor assembly 720 to another sensor assembly 720, so that the other sensor assembly 720 can continue to track the one or more features of the one or more moving objects.

[0101] In some embodiments, each sensor assembly 720 can be coupled to a host, i.e., a processor (controller) of the HMD 705 or the console 710. The sensor assembly 720 can be configured to send first data of a first resolution to the host using a first frame rate, the first data being associated with an image captured by the sensor assembly 720 at a first moment in time. The host can be configured to send information about one or more features obtained based on the first data received from the sensor assembly 720 using the first frame rate. The sensor assembly 720 can be further configured to send second data of a second resolution lower than the first resolution to the host using a second frame rate higher than the first frame rate, the second data being associated with another image captured by the sensor assembly at a second moment in time.

[0102] Each sensor assembly 720 may include an interface connection between each pixel in the array of the top sensor layer and logic circuitry of at least one of the one or more sensor layers located below the top sensor layer. At least one of the one or more sensor layers located below the top sensor layer of the sensor assembly 720 may include logic circuitry configured to extract one or more features from the captured one or more images. At least one of the one or more sensor layers located below the top sensor layer of the sensor assembly 720 may further include a CNN based on a memristor array for storing trained network weights.

[0103] At least one sensor assembly 720 can capture data describing depth information of a local area. The at least one sensor assembly 720 can use the data (e.g., based on the captured portion of the structured light pattern) to calculate the depth information. Alternatively, the at least one sensor assembly 720 can send this information to another device, such as the console 710, which can use the data from the sensor assembly 720 to determine the depth information. Each of these sensor assemblies 720 can be Figure 1A The sensor device 130 in Figure 2 The sensor assembly 200 in Figure 3 The sensor assembly 300, and / or Figure 6 An implementation of the sensor system 605 in.

[0104] The electronic display 725 displays a two-dimensional image or a three-dimensional image to the user based on the data received from the console 710. In various embodiments, the electronic display 725 includes a single electronic display or multiple electronic displays (e.g., a display for each eye of the user). Examples of electronic display 725 include: an LCD, an OLED display, an ILED display, an AMOLED display, a TOLED display, some other display, or some combination thereof. The electronic display 725 may be Figure 1B An embodiment of the electronic display 155 in.

[0105] Optical assembly 730 amplifies image light received from electronic display 725, corrects optical errors associated with the image light, and presents the corrected image light to a user of HMD 705. Optical assembly 730 includes a plurality of optical elements. Example optical elements included in optical assembly 730 include: an aperture, a Fresnel lens, a convex lens, a concave lens, a filter, a reflective surface, or any other suitable optical element that affects the image light. In addition, optical assembly 730 may include a combination of different optical elements. In some embodiments, one or more of the optical elements in optical assembly 730 may have one or more coatings such as a partially reflective coating or an anti-reflective coating.

[0106] The magnification and focusing of the image light by the optical assembly 730 allows the electronic display 725 to be physically smaller, lighter, and consume less power than larger displays. In addition, the magnification can increase the field of view of the content presented by the electronic display 725. For example, the field of view of the displayed content is such that the displayed content is presented using substantially all (e.g., approximately 110 degrees diagonally) and in some cases all of the field of view. In addition, in some embodiments, the amount of magnification can be adjusted by adding or removing optical elements.

[0107] In some embodiments, the optical assembly 730 can be designed to correct one or more types of optical errors. Examples of optical errors include barrel or pincushion distortion, longitudinal chromatic aberration, or lateral chromatic aberration. Other types of optical errors can further include spherical aberration, chromatic aberration or error due to lens field curvature, astigmatism, or any other type of optical error. In some embodiments, the content provided to the electronic display 725 for display is pre-distorted, and the optical assembly 730 corrects the distortion when receiving image light generated based on the content of the electronic display 725. In some embodiments, the optical assembly 730 is configured to direct image light emitted from the electronic display 725 to an eye box of the HMD 705 corresponding to the user's eye position. The image light may include depth information of a local area determined by at least one of the plurality of sensor assemblies 720 based in part on the processed data. The optical assembly 730 may be Figure 1B An embodiment of the optical component 160 in FIG.

[0108] The IMU 740 is an electronic device that generates data indicating the position of the HMD 705 based on measurement signals received from one or more of the position sensors 735 and depth information received from the at least one sensor assembly 720. The position sensor 735 generates one or more measurement signals in response to the movement of the HMD 705. Examples of the position sensor 735 include: one or more accelerometers, one or more gyroscopes, one or more magnetometers, another suitable type of sensor that detects movement, a type of sensor for error correction of the IMU 740, or some combination thereof. The position sensor 735 can be located outside the IMU 740, inside the IMU 740, or some combination thereof.

[0109] Based on the one or more measurement signals from the one or more position sensors 735, the IMU 740 generates data indicating an estimated current position of the HMD 705 relative to an initial position of the HMD 705. For example, the position sensors 735 include a plurality of accelerometers to measure translational motion (forward / backward, up / down, left / right) and a plurality of gyroscopes to measure rotational motion (e.g., pitch, yaw, roll). In some embodiments, the position sensors 735 may represent Figure 1A705. In some embodiments, the IMU 740 rapidly samples the measurement signals and calculates an estimated current position of the HMD 705 based on the sampled data. For example, the IMU 740 integrates the measurement signals received from the accelerometer over time to estimate a velocity vector and integrates the velocity vector over time to determine an estimated current position of a reference point on the HMD 705. Alternatively, the IMU 740 provides the sampled measurement signals to the console 710, which interprets the data to reduce errors. A reference point is a point that can be used to describe the position of the HMD 705. A reference point can generally be defined as a point or position in space that is related to the orientation and position of the HMD 705.

[0110] The IMU 740 receives one or more parameters from the console 710. The one or more parameters are used to maintain tracking of the HMD 705. Based on the received parameters, the IMU 740 can adjust one or more IMU parameters (e.g., sampling rate). In some embodiments, certain parameters cause the IMU 740 to update the initial position of the reference point so that it corresponds to the next position of the reference point. Updating the initial position of the reference point to the next calibration position of the reference point helps to reduce the accumulated error associated with the current position estimated by the IMU 740. The accumulated error (also called drift error) causes the estimated position of the reference point to "drift" away from the actual position of the reference point over time. In some embodiments of the HMD 705, the IMU 740 can be a dedicated hardware component. In other embodiments, the IMU 740 can be a software component implemented on one or more processors. In some embodiments, the IMU 740 can represent Figure 1A IMU 130.

[0111] In some embodiments, an eye tracking system 745 is integrated into the HMD 705. The eye tracking system 745 determines eye tracking information associated with the eyes of a user wearing the HMD 705. The eye tracking information determined by the eye tracking system 745 may include information about the orientation of the user's eyes, i.e., information about the angle at which the eyes are looking. In some embodiments, the eye tracking system 745 is integrated into the optical assembly 730. Embodiments of the eye tracking system 745 may include an illumination source and an imaging device (camera).

[0112] In some embodiments, a varifocal module 750 is further integrated into the HMD 705. The varifocal module 750 can be coupled to the eye tracking system 745 to obtain eye tracking information determined by the eye tracking system 745. The varifocal module 750 can be configured to adjust the focus of one or more images displayed on the electronic display 725 based on the determined eye tracking information obtained from the eye tracking system 745. In this way, the varifocal module 750 can mitigate the visual convergence accommodation conflict regarding the image light. The varifocal module 750 can be connected (e.g., mechanically or electrically) to at least one of the electronic display 725 and at least one optical element of the optical assembly 730. The varifocal module 750 can then be configured to adjust the focus of the one or more images displayed on the electronic display 725 by adjusting the position of at least one of the electronic display 725 and the at least one optical element of the optical assembly 730 based on the determined eye tracking information obtained from the eye tracking system 745. By adjusting the position, the varifocal module 750 changes the focus of the image light output from the electronic display 725 toward the user's eyes. The varifocal module 750 may also be configured to adjust the resolution of the image displayed on the electronic display 725 by concave rendering the displayed image based at least in part on the determined eye tracking information obtained from the eye tracking system 745. In this case, the varifocal module 750 provides an appropriate image signal for the electronic display 725. The varifocal module 750 provides an image signal with a maximum pixel density for the electronic display 725 only in the concave region where the user's eye is focused, and provides an image signal with a lower pixel density in other regions of the electronic display 725. In one embodiment, the varifocal module 750 may utilize the depth information obtained by the at least one sensor assembly 720, for example, to generate content for presentation on the electronic display 725.

[0113] The I / O interface 715 is a device that allows a user to send an action request to the console 710 and receive a response from the console. An action request is a request to perform a specific action. For example, an action request may be an instruction to start or end capturing image or video data, or an instruction to perform a specific action within an application. The I / O interface 715 may include one or more input devices. Example input devices include: a keyboard, a mouse, a game controller, or any other suitable device for receiving action requests and communicating these action requests to the console 710. The action request received by the I / O interface 715 is communicated to the console 710, which performs an action corresponding to the action request. In some embodiments, the I / O interface 715 includes an IMU 740 that captures IMU data, which indicates an estimated position of the I / O interface 715 relative to the initial position of the I / O interface 715. In some embodiments, the I / O interface 715 may provide tactile feedback to the user according to the instructions received from the console 710. For example, tactile feedback is provided when an action request is received, or the console 710 communicates instructions to the I / O interface 715 so that the I / O interface 715 generates tactile feedback when the console 710 performs an action.

[0114] The console 710 provides content to the HMD 705 for processing based on information received from one or more of the at least one sensor assembly 720, the HMD 705, and the I / O interface 715. Figure 7 In the example shown in FIG. 7 , the console 710 includes an application storage device 755, a tracking module 760, and an engine 765. Some embodiments of the console 710 have a combination of Figure 7 Similarly, the different modules or components described herein can be combined with Figure 7 The different ways described distribute the functionality described further below among the components of console 710.

[0115] The application storage device 755 stores one or more applications executed by the console 710. An application is a set of instructions that, when executed by a processor, generates content for presentation to a user. The content generated by the application may be responsive to input received from the user via the I / O interface 715 or movement of the HMD 705. Examples of applications include gaming applications, conferencing applications, video playback applications, or other suitable applications.

[0116] The tracking module 760 calibrates the HMD system 700 using one or more calibration parameters and can adjust one or more calibration parameters to reduce errors in determining the position of the HMD 705 or the I / O interface 715. For example, the tracking module 760 communicates the calibration parameters to the at least one sensor assembly 720 to adjust the focus of the at least one sensor assembly 720 to more accurately determine the position of the structured light elements captured by the at least one sensor assembly 720. The calibration performed by the tracking module 760 is also responsible for information received from the IMU 740 in the HMD 705 and / or the IMU 740 included in the I / O interface 715. In addition, if tracking of the HMD 705 is lost (e.g., the at least one sensor assembly 720 loses sight of at least a threshold number of structured light elements), the tracking module 760 can recalibrate part or all of the HMD system 700.

[0117] The tracking module 760 uses information from the at least one sensor assembly 720, the one or more position sensors 735, the IMU 740, or some combination thereof to track the movement of the HMD 705 or the I / O interface 715. For example, the tracking module 760 determines the location of a reference point of the HMD 705 in a map of the local area based on the information from the HMD 705. The tracking module 760 may also use data from the IMU 740 indicating the location of the HMD 705 or use data from the IMU 740 included in the I / O interface 715 indicating the location of the I / O interface 715 to determine the location of the reference point of the HMD 705 or the reference point of the I / O interface 715, respectively. Furthermore, in some embodiments, the tracking module 760 may use a portion of the data from the IMU 740 indicating the location of the HMD 705 and a representation of the local area from the at least one sensor assembly 720 to predict a future location of the HMD 705. The tracking module 760 provides the estimated or predicted future position of the HMD 705 or the I / O interface 715 to the engine 765 .

[0118] The engine 765 generates a 3D map of a local area surrounding part or all of the HMD 705 based on information received from the HMD 705. In some embodiments, the engine 765 determines depth information for the 3D map of the local area based on information related to the technique used in calculating depth received from the at least one sensor assembly 720. The engine 765 can calculate the depth information using one or more techniques in calculating depth from structured light. In various embodiments, the engine 765 uses the depth information to, for example, update a model of the local area and generate content based in part on the updated model.

[0119] The engine 765 also executes applications within the HMD system 700 and receives position information, acceleration information, velocity information, predicted future position, or some combination thereof of the HMD 705 from the tracking module 760. Based on the received information, the engine 765 determines the content provided to the HMD 705 for presentation to the user. For example, if the received information indicates that the user has looked to the left, the engine 765 generates content for the HMD 705 that reflects the user's movement in a virtual environment or in an environment that enhances the local area with additional content. In addition, the engine 765 performs an action within an application executed on the console 710 in response to an action request received from the I / O interface 715 and provides feedback that the action is performed to the user. The feedback provided may be visual feedback or auditory feedback via the HMD 705 or tactile feedback via the I / O interface 715.

[0120] In some embodiments, based on eye tracking information received from the eye tracking system 745 (e.g., the orientation of the user's eyes), the engine 765 determines the resolution of content provided to the HMD 705 for presentation to the user on the electronic display 725. The engine 765 provides content with maximum pixel resolution to the HMD 705 in the recessed area of ​​the electronic display 725 where the user is looking, while the engine 765 provides lower pixel resolution in other areas of the electronic display 725, thereby achieving less power consumption at the HMD 705 and saving computing cycles of the console 710 without compromising the user's visual experience. In some embodiments, the engine 765 can further use the eye tracking information to adjust where objects are displayed on the electronic display 725, thereby preventing a vergence-accommodation conflict.

[0121] Additional configuration information

[0122] The above description of the embodiments of the present disclosure is presented for illustrative purposes only and is not intended to be exhaustive or to limit the present disclosure to the exact form disclosed. It should be understood by those skilled in the relevant art that there may be many modifications and variations based on the above disclosure.

[0123] Some parts of this description describe the embodiments of the present disclosure from the perspective of algorithms and symbolic representations of information operations. These algorithmic descriptions and representations are usually used by technicians in the field of data processing to effectively convey their work essence to other technicians in the field. Although these operations are described functionally, computationally or logically, they should be understood to be implemented by computer programs or equivalent circuits, microcodes, etc. In addition, it is sometimes convenient to refer to the arrangement of these operations as modules without losing their generality. The described operations and their associated modules can be embodied as software, firmware, hardware or any combination thereof.

[0124] Any steps, operations or processes described herein may be performed or implemented using one or more hardware or software modules alone or in combination with other devices. In one embodiment, the software module may be implemented using a computer program product comprising a computer readable medium containing computer program code that may be executed by a computer processor that is used to perform any or all of the steps, operations or processes described.

[0125] Embodiments of the present disclosure may also relate to a device for performing the operations herein. The device may be specifically constructed for the desired purpose, and / or the device may include a general-purpose computing device selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-temporary tangible computer-readable storage medium that can be coupled to a computer system bus or any type of medium suitable for storing electronic instructions. In addition, any computing system mentioned in this specification may include a single processor or may be an architecture that uses a multi-processor design to increase computing power.

[0126] Embodiments of the present disclosure may also relate to products produced by the computing processes described herein. Such products may include information obtained by the computing processes, wherein the information is stored on a non-transitory tangible computer-readable storage medium and may include any embodiment of the computer program product or other data combination described herein.

[0127] Finally, the language used in this specification is selected in principle for the purpose of readability and illustrativeness, and the language used is not selected to delimit or limit the subject matter of the present invention. Therefore, the scope of the present disclosure is not intended to be limited by this detailed description, but is defined by any claims issued by this application based on the detailed description. Therefore, the disclosure of the embodiments is intended to illustrate, but not to limit the scope of the present disclosure set forth in the appended claims.

Claims

1. A head mounted display system, comprising: a first camera positioned to capture an image of a first field of view; a second camera positioned to capture an image of a second field of view that does not overlap with the first field of view; as well as a controller communicatively coupled to the first camera and the second camera, the controller being configured to: Based on specific circumstances, dynamically activating the second camera to capture an image of the second field of view; as well as deactivating the first camera; Wherein the particular situation comprises a physical object moving from the first field of view to the second field of view, and wherein the first camera and the second camera are configured to track the movement of the physical object.

2. The system according to claim 1, wherein: The specific scenario includes performing a simulation.

3. The system according to claim 1, wherein: The first camera and the second camera are positioned on a head mounted display.

4. The system according to claim 1, wherein: The first camera is configured to: capturing an image of a physical object in the first field of view; performing feature extraction on the captured image to identify features of the physical object; and Data regarding the characteristic is transmitted to the second camera.

5. The system according to claim 4, wherein: The second camera is configured to track the feature of the physical object based on the data.

6. The system according to claim 1, wherein: The controller is configured to: predicting that a particular camera in a set of cameras including the first camera and the second camera will be able to capture features of a physical object at a future point in time; and In response to predicting that the particular camera will be able to capture the feature of the physical object at the future point in time, the particular camera is activated.

7. The system according to claim 6, wherein: The controller is configured to, in response to predicting that the particular camera will be able to capture the feature of the physical object at the future point in time: At least one camera in the set of cameras is deactivated, the at least one camera being different from the particular camera.

8. A method of operating a head mounted display, comprising: dynamically activating, by the controller, the first camera to capture an image of the first field of view; Based on specific circumstances, deactivating, by the controller, the first camera; as well as activating, by the controller, a second camera; the second camera being positioned to capture an image of a second field of view that does not overlap with the first field of view; Wherein the particular situation comprises a physical object moving from the first field of view to the second field of view, and wherein the first camera and the second camera are configured to track the movement of the physical object.

9. The method according to claim 8, wherein: The specific scenario includes performing a simulation.

10. The method according to claim 8, wherein: The first camera and the second camera are positioned on a head mounted display.

11. The method according to claim 8, wherein: The first camera is configured to: capturing an image of a physical object in the first field of view; performing feature extraction on the captured image to identify features of the physical object; and Data regarding the characteristic is transmitted to the second camera.

12. The method according to claim 11, wherein: The second camera is configured to track the feature of the physical object based on the data.

13. The method according to claim 8, wherein: The controller is configured to: predicting that a particular camera in a set of cameras including the first camera and the second camera will be able to capture features of a physical object at a future point in time; and In response to predicting that the particular camera will be able to capture the feature of the physical object at the future point in time, the particular camera is activated.

14. The method according to claim 13, wherein: The controller is configured to, in response to predicting that the particular camera will be able to capture the feature of the physical object at the future point in time: At least one camera in the set of cameras is deactivated, the at least one camera being different from the particular camera.

15. A non-transitory computer readable medium comprising program code executable by a processor to cause the processor to: dynamically activating a first camera to capture an image of a first field of view; Based on specific circumstances, deactivating the first camera; as well as activating a second camera; the second camera being positioned to capture an image of a second field of view that does not overlap with the first field of view; The particular scenario includes a physical object moving from the first field of view to the second field of view, and wherein the first camera and the second camera are configured to track the movement of the physical object.

16. The non-transitory computer readable medium of claim 15, wherein: The specific scenario includes performing a simulation.

17. The non-transitory computer-readable medium of claim 15, further comprising program code executable by the processor to cause the processor to: predicting that a particular camera in a set of cameras including the first camera and the second camera will be able to capture features of a physical object at a future point in time; and In response to predicting that the particular camera will be able to capture the feature of the physical object at the future point in time, the particular camera is activated.

Citation Information

Patent Citations

  • Head-mounted display tracking system

    CN111194423A

  • Position tracking system for head-mounted displays that includes sensor integrated circuits

    CN111602082A

  • Intelligent sensor

    CN115299039A

  • Golfer's Eye View

    US20170142329A1

  • Method for processing media content and technical equipment for the same

    US20180205933A1