Energy-saving adaptive three-dimensional sensing

By using a distributed laser beam in an adaptive 3D sensing system, 3D data capture is only performed in designated areas, which solves the problem of high energy consumption of existing 3D sensing methods and achieves efficient and energy-saving 3D sensing effects.

CN120019637APending Publication Date: 2025-05-16SNAP INC

Patent Information

Application Number
CN202380072221.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-13
Filing Date
2023-10-05
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing 3D sensing methods consume high energy when achieving accurate three-dimensional sensing and have strict requirements on heat, which affects the sustainability of the equipment and user experience.

Method used

Adaptive 3D sensing system is adopted, which uses a distributed laser beam to capture 3D data only in designated areas of real-world scenes, reducing the lighting and processing requirements for the entire scene, thereby reducing energy consumption.

Benefits of technology

It achieves a significant reduction in energy consumption while maintaining accurate 3D sensing, improves the energy saving and user experience of the equipment, and increases eye safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019637A_ABST
    Figure CN120019637A_ABST
Patent Text Reader

Abstract

The invention discloses an energy-saving adaptive three-dimensional sensing system. An adaptive 3D sensing system includes one or more camera devices and one or more projectors. An adaptive 3D sensing system captures an image of a real-world scene using one or more camera devices and computes a depth estimate and a depth estimate confidence value for pixels of the image. The adaptive 3D sensing system calculates an attention mask based on the one or more depth estimation confidence values, and commands the one or more projectors to send a distributed laser beam into one or more regions of the real-world scene based on the attention mask. An adaptive 3D sensing system captures 3D sensed image data of one or more regions of a real-world scene and generates 3D sensed data of the real-world scene based on the 3D sensed image data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of priority to Greek patent application serial number 20220100840 filed on October 12, 2022, and U.S. patent application serial number 18 / 299,923 filed on April 13, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0003] The present disclosure relates generally to user interfaces, and more particularly to user interfaces used in augmented reality and virtual reality. Background Art

[0004] A head mounted device may be implemented with a transparent or translucent display through which a user of the head mounted device can view the surrounding environment. Such a device enables a user to view the surrounding environment through the transparent or translucent display, and also to see objects (e.g., virtual objects, such as renderings of 2D or 3D graphics models, images, videos, text, etc.) generated for display that appear as part of the surrounding environment and / or superimposed on the surrounding environment. This is commonly referred to as "augmented reality" or "AR". A head mounted device may also completely block the user's field of view and display a virtual environment through which the user can move or be moved. This is commonly referred to as "virtual reality" or "VR". In a hybrid form, a camera is used to capture a view of the surrounding environment, and then the view is displayed to the user together with the augmentation on a display that blocks the user's eyes. As used herein, unless the context indicates otherwise, the term extended reality (XR) refers to augmented reality, virtual reality, and any mixture of these technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] To easily identify the discussion of any particular element or act, the highest digit or digits in a reference number refer to the figure number in which the element is first introduced.

[0006] Figure 1 is a perspective view of a head mounted device according to some examples.

[0007] Figure 2 According to some examples Figure 1 Additional views of the head-mounted device.

[0008] Figure 3 is a diagrammatic representation of a machine according to some examples within which a set of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein.

[0009] Figure 4is an illustration of an adaptive 3D sensing system according to some examples of the present disclosure.

[0010] Figure 5 is an illustration of another adaptive 3D sensing system according to some examples of the present disclosure.

[0011] Figure 6 are illustrations of SLM-based projectors according to some examples of the present disclosure.

[0012] Figure 7 is an illustration of a MEMS+DOE based projector according to some examples of the present disclosure.

[0013] Fig. 8A is a process flow chart of an adaptive 3D sensing method according to some examples of the present disclosure.

[0014] Figure 8B is a diagram of an operating environment for an adaptive 3D sensing system according to some examples of the present disclosure.

[0015] Figure 8C is an illustration of stages of an adaptive 3D sensing method according to some examples of the present disclosure.

[0016] Fig.8D is an illustration of stages of an adaptive 3D sensing method according to some examples of the present disclosure.

[0017] Fig. 8E is an illustration of a projector component of an adaptive 3D sensing system according to some examples of the present disclosure.

[0018] Fig.8F is an illustration of another projector assembly of an adaptive 3D sensing system according to some examples of the present disclosure.

[0019] Fig. 9 is a block diagram illustrating a software architecture within which the present disclosure may be implemented, according to some examples.

[0020] Fig.10 is a block diagram showing a networked system including details of a head-mounted XR system according to some examples.

[0021] Fig.11 is a diagrammatic representation of a networking environment in which the present disclosure may be deployed, according to some examples. DETAILED DESCRIPTION

[0022] Users of mobile devices (such as smart phones and XR glasses) use their mobile devices for a variety of purposes where accurate three-dimensional (3D) sensing of real-world scenes is desired. Example uses include XR applications, where 3D models of real-world scenes are used to place virtual objects in the XR experience provided to the user (such as digital avatars, etc.), and hand tracking services are provided for the user interface of the XR application. Due to the physical limitations of mobile devices, the capacity of the total energy of the battery is very limited. In addition, wearables have strict requirements on heat, and energy-saving devices generally generate much less heat. Therefore, energy-saving 3D sensing methods are desired.

[0023] Some 3D sensing methods use a large amount of energy to drive a light projector to illuminate the entire real-world scene object, to power a camera to collect image data, and to provide computing resources for computer vision processing of the image data captured by the camera. In a full-mode 3D sensing system, a light source that consumes a large amount of energy is used to illuminate the entire real-world scene. In addition, as the entire real-world scene is illuminated, the image data collected for the entire real-world scene is processed using computing resources at a high resolution that requires additional energy. In a line scanning 3D method, a laser is used to sequentially scan the entire real-world scene, and a rolling shutter camera is used to capture the image data. Although an improvement is provided in the operating distance, this method still scans the entire real-world scene and therefore uses a comparable amount of energy as the full-mode method. Point scanning 3D sensing methods use a single point laser light source to scan the entire real-world scene. Although an improved distance is provided compared to full-mode 3D sensing, systems using point scanning methods are slower and still require a large amount of energy when the system point scans the entire real-world scene. Additionally, eye safety is an issue in full-field point scanning and line scanning because the laser energy is concentrated not only in space but also in time.

[0024] In some examples of the present disclosure, an adaptive 3D sensing system simultaneously captures images using one or more cameras of an area of ​​a 3D real-world scene. The adaptive 3D sensing system calculates a depth estimate and a confidence value for each pixel of the image. The adaptive 3D sensing system calculates an attention mask for an area of ​​the image that satisfies one or more conditions, such as, but not limited to: (1) the depth confidence is below a specified threshold for pixels in the area; (2) a virtual object of an XR user interface will be rendered in an area of ​​the real-world scene corresponding to the area; and / or (3) an area of ​​the real-world scene corresponds to an area of ​​the image that has not yet been mapped to a 3D model of the real-world scene. The attention mask includes mask data that is used to send a distributed laser beam into an area of ​​the real-world scene corresponding to an area of ​​the image that satisfies one or more conditions. The adaptive 3D sensing system commands one or more projectors to project or send a distributed laser beam into an area of ​​the real-world scene based on the attention mask.

[0025] In some examples, an adaptive 3D sensing system uses less energy than a system that scans an entire real-world scene by sending a distributed laser beam to specified areas of the real-world scene and capturing 3D data for those areas of the real-world scene rather than capturing 3D data in other areas.

[0026] In some examples, energy reduction is achieved by computing partial depth of the real-world scene.

[0027] In some examples, the adaptive 3D sensing system increases eye safety because the distributed laser beam can have a lower power and can therefore stay at a certain position for a longer time than in a full-view scanning system (e.g., a line scanning system), resulting in the laser energy being dispersed in time. For example, a line scanning system can scan 100 to 1000 lines per frame, and given the same frame rate, the time that a distributed laser beam in an adaptive 3D sensing system stays at a certain position may be 100 to 1000 times longer than a distributed laser beam (which is a line) of a line scanning system. The maximum permissible exposure (MPE) is the maximum permissible laser energy that does not cause damage to the human eye. As the exposure time becomes longer, the MPE increases. For the same amount of energy, a system with a longer duration is safer.

[0028] In some examples, the projector includes a diffractive optical element (DOE), a laser, and a movable micro-electromechanical system (MEMS) mirror. The MEMS mirror is operable to deflect a laser beam through the DOE at a smaller dispersion angle and send the distributed laser beam to one or more designated areas of a real-world scene selected for 3D sensing based on an attention mask. In some examples, during exposure of a depth frame, the MEMS mirror does not rotate, and the laser pattern is deflected to a position considered to be of interest. In some examples, during exposure of a depth frame, the MEMS mirror rotates to deflect the laser pattern to several positions while capturing the depth frame; here, in some examples, the camera device captures several frames, each of which corresponds to a pattern position, thereby improving the signal-to-noise ratio compared to capturing one frame when the ambient lighting noise is increased.

[0029] In some examples, a phase spatial light modulator (SLM) is used to send a distributed laser beam of a laser onto one or more designated areas of a real-world scene for 3D sensing selected based on an attention mask.

[0030] In some examples, one or more cameras are used to capture image data including valid 3D image sensing data in the illuminated area.

[0031] In some examples, one or more depth sensors are used to capture 3D data including 3D depth sensor data.

[0032] Other technical features will be readily apparent to those skilled in the art from the following drawings, descriptions and claims.

[0033] Figure 1 is a head-mounted XR system according to some examples (e.g., Figure 1 100). The glasses 100 may include a frame 102, which is made of any suitable material, such as plastic or metal, including any suitable shape memory alloy. In one or more examples, the frame 102 includes a first optical element holder or a left optical element holder 104 (e.g., a display or lens holder) and a second optical element holder or a right optical element holder 106 connected by a bridge 112. A first optical element or a left optical element 108 and a second optical element or a right optical element 110 may be disposed in the left optical element holder 104 and the right optical element holder 106, respectively. The right optical element 110 and the left optical element 108 may be lenses, displays, display components, or a combination of the foregoing. Any suitable display component may be disposed in the glasses 100.

[0034] The frame 102 additionally includes a left arm or temple piece 122 and a right arm or temple piece 124. In some examples, the frame 102 can be formed from a single piece of material to have a unitary or unitary construction.

[0035] The glasses 100 may include a computing system such as a computer 120, which may be of any suitable type so as to be carried by the frame 102, and in one or more examples, the computing system may be of a suitable size and shape to be partially disposed in one of the temple pieces 122 or temple pieces 124. The computer 120 may include multiple processors, memories, and various communication components that share a common power supply. As discussed below, the various components of the computer 120 may include low-power circuits, high-speed circuits, and display processors. Various other examples may include these elements in different configurations or integrated together in different ways. Additional details of various aspects of the computer 120 may be implemented as shown in the data processor 1002 discussed below.

[0036] The computer 120 additionally includes a battery 118 or other suitable portable power source. In some examples, the battery 118 is disposed in the left temple piece 122 and is electrically coupled to the computer 120 disposed in the right temple piece 124. The glasses 100 may include a connector or port (not shown) suitable for charging the battery 118, a wireless receiver, transmitter or transceiver (not shown), or a combination of such devices.

[0037] The glasses 100 include a first or left camera 114 and a second or right camera 116. Although two cameras are depicted, other examples contemplate the use of a single or additional (i.e., more than two) cameras. In one or more examples, the glasses 100 include any number of input sensors or other input / output devices in addition to the left camera 114 and the right camera 116. Such sensors or input / output devices may additionally include biometric sensors, position sensors, motion sensors, etc.

[0038] In some examples, left camera 114 and right camera 116 provide video frame data for use by glasses 100 to extract 3D information from an ambient scene of a real-world scene.

[0039] The glasses 100 may also include a touchpad 126 mounted to or integrated with one or both of the left temple piece 122 and the right temple piece 124. The touchpad 126 is typically arranged vertically, approximately parallel to the temple of the user in some examples. As used herein, typically vertically aligned refers to the touchpad being more vertical than horizontal, although potentially more vertical than that. Additional user input may be provided by one or more buttons 128, which in the example shown are disposed on the outer upper edges of the left optical element holder 104 and the right optical element holder 106. The one or more touchpads 126 and buttons 128 provide a means by which the glasses 100 can receive input from a user of the glasses 100.

[0040] In some examples, the glasses 100 have a projector 130 mounted in a forward-facing position on the frame 102 of the glasses 100. The adaptive 3D sensing system of the glasses 100 can use the projector to project a focused light beam, thereby enabling the adaptive 3D sensing system to perform adaptive 3D sensing.

[0041] Figure 2 The glasses 100 are shown from the user's perspective. Figure 1 Several elements shown in FIG. have been omitted. Figure 1 As described, Figure 2 The illustrated eyeglasses 100 include left and right optical elements 108, 110 secured within left and right optical element holders 104, 106, respectively.

[0042] The glasses 100 include a forward optical assembly 202 including a right projector 204 and a right near-eye display 206 , and a forward optical assembly 210 including a left projector 212 and a left near-eye display 216 .

[0043] In some examples, the near-eye display is a waveguide. The waveguide includes a reflective structure or a diffractive structure (e.g., a grating and / or an optical element such as a mirror, a lens, or a prism). Light 208 emitted by projector 204 encounters the diffractive structure of the waveguide of near-eye display 206, which directs the light toward the right eye of the user to provide an image superimposed with a view of the real-world scene environment seen by the user on or in the right optical element 110. Similarly, light 214 emitted by projector 212 encounters the diffractive structure of the waveguide of near-eye display 216, which directs the light toward the left eye of the user to provide an image superimposed with a view of the real-world scene environment seen by the user on or in the left optical element 108. The combination of the GPU, the forward optical assembly 202, the left optical element 108, and the right optical element 110 provides an optical engine of the glasses 100. The glasses 100 use the optical engine to generate an overlay of an environmental view of the user's real-world scene, including displaying a user interface to the user of the glasses 100.

[0044] However, it will be appreciated that other display technologies or configurations may be utilized within the optical engine to display images to a user within the user's field of view. For example, instead of providing a projector 204 and a waveguide, an LCD, LED or other display panel or surface may be provided.

[0045] In use, information, content, and various user interfaces may be presented to a user of the glasses 100 on the near-eye display. As described in more detail herein, the user may then use the touch pad 126 and / or buttons 128, associated devices (e.g., Fig.10 The user can interact with the glasses 100 by using voice input or touch input on the mobile computing system 1026 shown, and / or hand movements, positions, and locations detected by the glasses 100.

[0046] In some examples, glasses 100 include a standalone AR system that provides an AR experience to a user of glasses 100. In some examples, glasses 100 are a component of an AR system that includes one or more other devices that provide additional computing resources and / or additional user input and output resources. Other devices may include smartphones, general-purpose computers, etc.

[0047] Figure 3 is a diagrammatic representation of a machine 300 within which instructions 310 (e.g., software, programs, applications, applet, app, other executable code) may be executed that cause the machine 300 to perform any one or more of the methodologies discussed herein. The machine 300 may be used as an AR system (e.g., Figure 1The computer 120 of the glasses 100). For example, the instructions 310 can cause the machine 300 to perform any one or more of the methods described herein. The instructions 310 transform the general, unprogrammed machine 300 into a specific machine 300 that is programmed to perform the functions described and shown in the described manner. The machine 300 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 300 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 300, in combination with other components of the AR system, may be used as, but not limited to: a server, a client, a computer, a personal computer (PC), a tablet computer, a laptop computer, a notebook, a set-top box (STB), a PDA, an entertainment media system, a cellular phone, a smart phone, a mobile device, a head-mounted device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, web appliances, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 310 specifying actions to be taken by the machine 300. In addition, although a single machine 300 is shown, the term "machine" may also be understood to include a collection of machines that individually or collectively execute instructions 310 to perform any one or more of the methods discussed herein.

[0048] Machine 300 may include processor 302, memory 304, and I / O device interface 306, which may be configured to communicate with each other via bus 344. In an example, processor 302 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 308 and processor 312 that execute instructions 310. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously. Although Figure 3 Multiple processors 302 are shown, but machine 300 may include a single processor with a single core, a single processor with multiple cores (eg, a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0049] The memory 304 includes a main memory 314, a static memory 316, and a storage unit 318, which are all accessible by the processor 302 via the bus 344. The main memory 304, the static memory 316, and the storage unit 318 store instructions 310 that implement any one or more of the methods or functions described herein. During execution of the instructions 310 by the machine 300, the instructions 310 may also reside, in whole or in part, within the main memory 314, within the static memory 316, within the non-transitory machine-readable medium 320, within the storage unit 318, within one or more of the processors 302 (e.g., within a processor's cache memory), or within any suitable combination thereof.

[0050] The I / O device interface 306 couples the machine 300 to I / O devices 346. One or more of the I / O devices 346 may be components of the machine 300 or may be separate devices. The I / O device interface 306 may include various interfaces to the I / O devices 346 that are used by the machine 300 to receive input, provide output, generate output, transmit information, exchange information, capture measurements, etc. The specific I / O device interface 306 included in a particular machine will depend on the type of machine. It will be appreciated that the I / O device interface 306, the I / O devices 346 may include various interfaces to the I / O devices 346. Figure 3 306. In various examples, the I / O device interface 306 may include an output component interface 328 and an input component interface 332. The output component interface 328 may include an interface to a visual component (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), an acoustic component (e.g., a speaker), a tactile component (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. The input component interface 332 may include an interface to an alphanumeric input component (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input component), a pointing-based input component (e.g., a mouse, a touch pad, a trackball, a joystick, a motion sensor, or other pointing instrument), a tactile input component (e.g., a physical button, a touch screen that provides a touch position and / or a touch force or touch gesture, or other tactile input component), an audio input component (e.g., a microphone), etc.

[0051] In another example, the I / O device interface 306 may include: a biometric component interface 334, a motion component interface 336, an environmental component interface 338, or a positioning component interface 340, as well as various other component interfaces. For example, the biometric component interface 334 may include an interface to the following components: a component for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biological signals (e.g., blood pressure, heart rate, body temperature, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition), etc. The motion component interface 336 may include an interface to the following components: an inertial measurement unit (IMU), an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environmental component interface 338 may include, for example, an interface to an illumination sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor that detects concentrations of hazardous gases for safety or measures pollutants in the atmosphere), or other components that can provide indications, measurements, or signals associated with the surrounding real-world scene. The positioning component interface 340 includes an interface to a position sensor component (e.g., a global positioning system (GPS) receiver component and / or an inertial measurement unit (IMU), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure (from which altitude can be derived)), an orientation sensor component (e.g., a magnetometer), etc.

[0052] A variety of technologies may be used to achieve communication. I / O device interface 306 also includes a communication component interface 342 that is operable to couple machine 300 to network 322 or device 324 via coupling 330 and coupling 326, respectively. For example, communication component interface 342 may include an interface to a network interface component or another suitable device that interfaces with network 322. In other examples, communication component interface 342 may include an interface to a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, Components (e.g. Low power consumption), Device 324 may be another machine or any of a variety of peripherals (eg, a peripheral coupled via USB).

[0053] In addition, the communication component interface 342 may include an interface to a component operable to detect an identifier. For example, the communication component interface 342 may include an interface to a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes, such as universal product codes (UPC) barcodes; multi-dimensional barcodes, such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). In addition, various information may be obtained via the communication component interface 342, such as a location obtained via Internet Protocol (IP) geolocation, ... Location derived from signal triangulation, location derived via detection of NFC beacon signals that can indicate a specific location, etc.

[0054] Various memories (e.g., memory 304, main memory 314, static memory 316, and / or memory of processor 302) and / or storage unit 318 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. These instructions (e.g., instructions 310) when executed by processor 302 cause various operations to implement the disclosed examples.

[0055] Instructions 310 may be transmitted or received over network 322 via a network interface device (e.g., a network interface component included in communication component interface 342) using a transmission medium and using any of a number of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 310 may be transmitted or received to device 324 via coupling 326 (e.g., a peer-to-peer coupling) using a transmission medium.

[0056] Figure 44 is an illustration of an adaptive 3D sensing system according to some examples of the present disclosure. The XR system uses an adaptive 3D sensing system 410 to generate 3D sensing data 416 and transmits the 3D sensing data 416 to an XR application 408, which provides an XR user interface 424 to a user 422 interacting with the XR application 408. The adaptive 3D sensing system 410 includes one or more cameras, such as camera 402 and camera 404. The adaptive 3D sensing system 410 uses one or more cameras to capture image data of a real-world scene 418, such as image data 412 and image data 414. The image data is transmitted to an adaptive 3D sensing component 406 of the adaptive 3D sensing system 410. The adaptive 3D sensing component 406 receives the image data and generates projector command data 426 based on the image data (as shown in FIG. 4 ). Fig. 8A 4. The adaptive 3D sensing component 406 transmits the projector command data 426 to the projector 420. The projector 420 receives the projector command data 426 and generates a distributed laser beam 432 based on the projector command data 426. The projector 420 uses the projector 420 to send the distributed laser beam 432 to an area of ​​the real world scene 418 (as shown in FIG. 4). Figure 6 and Figure 7 4 and 5), and the adaptive 3D sensing system 410 uses one or more cameras to capture 3D sensing image data of the real world scene 418, such as 3D sensing image data 430 and 3D sensing image data 428. The adaptive 3D sensing component 406 receives the 3D sensing image data and generates 3D sensing data 416 based on the 3D sensing image data 428, and transmits the 3D sensing data 416 to the XR application 408. The XR application receives the 3D sensing data 416 and uses the 3D sensing data 416 to generate or update the XR user interface 424.

[0057] In some examples, including head mounted devices such as ( Figure 1 The XR system of the glasses 100 uses the adaptive 3D sensing system 410 to generate 3D sensing data of a real-world scene viewed by a user of the XR system and display an XR user interface to the user.

[0058] In some examples, an XR system including a smartphone, tablet, or other mobile device uses adaptive 3D sensing system 410 to generate 3D sensing data of a real-world scene viewed by a user of the XR system and display an XR user interface to the user.

[0059] In some examples, an XR system including a computing system (such as a personal computer) uses adaptive 3D sensing system 410 to generate 3D sensing data of a real-world scene viewed by a user of the XR system and display an XR user interface to the user.

[0060] In some examples, adaptive 3D sensing system 410 includes two or more cameras and two or more projectors to capture a large real-world scene 418 .

[0061] In some examples, a projector of the adaptive 3D sensing system is mounted adjacent to one or more cameras of the adaptive 3D sensing system.

[0062] In some examples, a projector of the adaptive 3D sensing system is mounted at a location spaced apart relative to one or more cameras of the adaptive 3D sensing system.

[0063] Figure 5 5 is an illustration of an adaptive 3D sensing system according to some examples of the present disclosure. The XR system uses an adaptive 3D sensing system 510 to generate 3D sensing data 516 and transmits the 3D sensing data 516 to an XR application 508, which provides an XR user interface 524 to a user 522 interacting with the XR application 508. The adaptive 3D sensing system 510 includes one or more cameras, such as camera 502 and camera 504. The adaptive 3D sensing system 510 uses one or more cameras to capture image data of a real-world scene 518, such as 3D sensing image data 512 and 3D sensing image data 514. The image data is transmitted to an adaptive 3D sensing component 506 of the adaptive 3D sensing system 510. The adaptive 3D sensing component 506 receives the image data and generates projector command data 526 based on the image data (as shown in FIG. 5 ). Fig. 8A 526). The adaptive 3D sensing component 506 transmits the projector command data 526 to the projector 520. The projector 520 receives the projector command data 526 and generates a distributed laser beam 530 based on the projector command data 526. The adaptive 3D sensing system 510 uses the projector 520 to send the distributed laser beam 530 to an area of ​​the real world scene 518 (as shown in FIG. Figure 6 and Figure 75 (described more fully in the accompanying drawings), and the adaptive 3D sensing system 510 uses one or more depth sensors 532 to capture 3D depth sensor data 528 of the real world scene 518. The adaptive 3D sensing component 506 receives the 3D depth sensor data 528 and generates 3D sensing data 516 based on the 3D depth sensor data 528, and transmits the 3D sensing data 516 to the XR application 508. The XR application receives the 3D sensing data 516 and uses the 3D sensing data 516 to generate or update the XR user interface 524.

[0064] In some examples, depth sensor 532 includes a Laser Detection and Ranging - Time of Flight (LIDAR-ToF) depth sensor.

[0065] In some examples, depth sensor 532 includes a continuous wave-time of flight (CW-ToF) depth sensor.

[0066] In some examples, depth sensor 532 includes a single photon avalanche diode (SPAD)-ToF depth sensor for estimating the depth of real-world scene 518 .

[0067] In some examples, depth sensor 532 includes a frequency modulated continuous wave (FMCW)-ToF depth sensor for estimating the depth of real-world scene 518 .

[0068] In some examples, depth sensor 532 is mounted adjacent to one or more projectors 520. In some examples, depth sensor 532 is mounted in spaced relationship to one or more projectors.

[0069] In some examples, including head mounted devices such as ( Figure 1 The XR system of the glasses 100 uses the adaptive 3D sensing system 510 to generate 3D sensing data of a real-world scene viewed by a user of the XR system and display an XR user interface to the user.

[0070] In some examples, an XR system comprising a smartphone, tablet, or other mobile device uses adaptive 3D sensing system 510 to generate 3D sensing data of a real-world scene viewed by a user of the XR system and display an XR user interface to the user.

[0071] In some examples, an XR system comprising a computing system such as a personal computer uses adaptive 3D sensing system 510 to generate 3D sensing data of a real-world scene viewed by a user of the XR system and display an XR user interface to the user.

[0072] In some examples, adaptive 3D sensing system 510 includes two or more depth sensors and two or more projectors to capture large real-world scenes 418 .

[0073] Figure 6 6 is a diagram of an SLM-based projector according to some examples of the present disclosure. The adaptive 3D sensing system uses an SLM-based projector 600 to distribute a laser beam 608 and selectively sends the laser beam 608 as a distributed laser beam (e.g., distributed laser beam 622) into a real-world scene 610. The SLM-based projector 600 includes an SLM 606, a laser 602 that projects the laser beam 608 onto the SLM 606, and a projector controller 604. The SLM is operable to freely distribute the laser beam so that different random dot patterns are projected onto regions of the real-world scene, such as region 612 and region 614.

[0074] The projector controller 604 is operable to receive the projector command data 616, and to generate a phase mask signal 618 based on the projector command data 616, which causes the SLM 606 to generate a distributed laser beam, such as the distributed laser beam 622, using the laser beam 608. The SLM 606 sends the distributed laser beam 622 to one or more regions of the real-world scene 610, such as the region 612 and the region 614, based on the projector command data 616. The projector controller 604 is also operable to generate a laser control signal 620 based on the projector command data 616, and to control the laser 602 using the laser control signal 620.

[0075] In some embodiments, SLM 606 diffracts laser beam 608 and distributes laser beam 608 into a pattern having two or more spots or beams simultaneously based on phase mask signal 618. In some embodiments, SLM 606 distributes laser beam 608 into an arbitrary pattern based on phase mask signal 618.

[0076] In some embodiments, the SLM 606 is included in an assembly with one or more optical elements 626 (eg, but not limited to, a lens, etc.).

[0077] In some examples, the projector controller 604 is operable to generate laser synchronization data 624 for the laser 602, which is used by the adaptive 3D sensing system to synchronize the capture of 3D sensing image data using a camera or the capture of 3D depth sensor data using a depth sensor. For example, the laser synchronization data 624 may include timing data for powering on and off the laser 602 for determining when an image is captured by one or more cameras or for ToF calculations. The laser synchronization data 624 may also include phase data of the laser beam 608 generated by the laser 602 for ToF calculations.

[0078] Figure 7 7 is an illustration of a MEMS+DOE-based projector according to some examples of the present disclosure. An adaptive 3D sensing system uses a MEMS+DOE-based projector 700 to generate a distributed laser beam, such as a distributed laser beam 718, and selectively focuses the distributed laser beam on a real-world scene 716. The MEMS+DOE-based projector 700 includes a DOE 708 with a low dispersion angle, a laser 706, and a movable MEMS mirror 704, which can be operated by a projector controller 702 to move the MEMS mirror 704 to deflect the laser beam 710 through the DOE 708. The projector controller 702 is operable to receive projector command data 720 and generate a MEMS mirror control signal 722 based on the projector command data 720. The projector controller 702 transmits the MEMS mirror control signal 722 to cause the MEMS mirror 704 to deflect the laser beam 710 through the DOE 708. When laser beam 710 passes through DOE 708, laser beam 710 is diffracted to generate distributed laser beams, such as distributed laser beam 718 and distributed laser beam 726. In addition, by positioning MEMS mirror 704, the projector sends the distributed laser beam to a specified area of ​​real world scene 716, such as area 712 and area 714.

[0079] In some embodiments, a distributed laser beam generated by passing the laser beam through a DOE includes a randomized pattern produced by the DOE. The DOE is constructed in such a way that a grating of the DOE produces a randomized pattern in the distributed laser beam.

[0080] In some examples, the projector controller 702 is operable to generate laser synchronization data 724 for the laser 706, which is used by the adaptive 3D sensing system to synchronize the capture of 3D sensing image data using a camera or the capture of 3D depth sensor data using a depth sensor. For example, the laser synchronization data 724 may include timing data for powering on and off the laser 706 for determining when an image is captured by one or more cameras or for ToF calculations. The laser synchronization data 724 may also include phase data of the laser beam 710 generated by the laser 706 for ToF calculations.

[0081] Fig. 8A is a process flow chart of an adaptive 3D sensing method 800, Figure 8B is a diagram of the operating environment of the adaptive 3D sensing system, Figure 8C and Fig.8D is a diagram of the stages of an adaptive 3D sensing method 800, Fig. 8E and Fig.8F is an illustration of a projector component of an adaptive 3D sensing system according to some examples of the present disclosure.The adaptive 3D sensing system uses an adaptive 3D sensing method 800 to perform adaptive 3D sensing of an area of ​​a real-world scene.

[0082] In operation 802 , an adaptive 3D sensing system simultaneously uses one or more cameras, such as camera 818 and camera 820 , to capture image data of a real-world scene 822 .

[0083] In operation 804, the adaptive 3D sensing system calculates an attention mask 832 for an area 824 of the real-world scene 822 based on the image data, the area 824 to be subjected to adaptive 3D sensing by the adaptive 3D sensing system. For example, the adaptive 3D sensing system determines the area 824 of the real-world scene 822 to be subjected to adaptive 3D sensing based on an area of ​​the image data corresponding to the area 824 that satisfies one or more conditions, such as, but not limited to: (1) the depth confidence is below a specified threshold level for pixels in the image area of ​​the image data corresponding to the area 824; (2) a virtual object of the XR user interface is to be rendered in the area 824 of the real-world scene; and / or (3) the area 824 of the real-world scene 822 has not yet been mapped into a 3D model of the real-world scene 822 that is used by the XR application to provide the XR user interface to the user. The attention mask 832 includes mask data (e.g., mask data 866) for distributing a laser beam having a distributed laser beam pattern (e.g., distributed laser beam pattern 868 and distributed laser beam pattern 870) to one or more regions (e.g., region 830) of the real-world scene corresponding to image regions that satisfy one or more conditions. Thus, the attention mask 832 determines regions in the real-world scene 822 that will be subjected to adaptive 3D sensing by the adaptive 3D sensing system.

[0084] In some examples, to determine that the depth confidence is below a specified threshold level for pixels in the image area corresponding to area 824, the adaptive 3D sensing system calculates a depth estimate and confidence for each pixel of the image of the real-world scene 822 based on the image data. For example, the adaptive 3D sensing system extracts first feature data from first image data received from a first camera, and extracts second feature data from second image data received from a second camera. The adaptive 3D sensing system extracts a depth estimate for each pixel in the image data based on the first image data and the second image data using a stereo image processing method. In some examples, binocular disparity is calculated for a pixel p in the first image data, which establishes a correspondence between pixel p and pixel q in the second image data. The disparity is estimated so that the intensity values ​​at one or more pixels around p in the first image data are structurally similar to the intensity values ​​at pixels around q in the second image data. The adaptive 3D sensing system determines the confidence value for the pixel of the image data based on whether such a correspondence can be found uniquely at each pixel, that is, whether there is another pixel q' whose surrounding pixels also look similar to p. The higher the uniqueness, the higher the confidence value. When the estimated confidence value of a pixel is lower than a specified threshold confidence value, this means that the area of ​​the real-world scene 822 corresponding to the pixel should be subjected to adaptive 3D sensing by the adaptive 3D sensing system because the depth data of the area of ​​the real-world scene 822 is unreliable. Therefore, the adaptive 3D sensing system generates an attention mask 832 based on the estimated depth confidence value, for example, by generating the following attention mask: the attention mask directs the distributed laser beam to those areas of the real-world scene where the depth estimation confidence value is lower than the specified threshold. The adaptive 3D sensing system generates projector command data based on the attention mask 832 and transmits the projector command data to the projector 816 to command the projector 816 to project the distributed laser beam 826 to selectively sense the areas 824 of the real-world scene 822 where the depth estimation confidence value is lower than the specified confidence value.

[0085] In some examples, the adaptive 3D sensing system uses a single camera. The projector 130 generates a known structured light pattern that deforms when it falls on various 3D surfaces. The single camera is positioned so that the optical axis of the camera is offset from the optical axis of the projector. The adaptive 3D sensing captures image data and estimates depth based on the image data, the known structured light pattern, and the offset between the optical axis of the camera and the optical axis of the projector.

[0086] In some examples, the adaptive 3D sensing system generates the attention mask 832 based on the real-world scene 822 position of the virtual object to be rendered in the area 824 of the real-world scene 822. For example, the adaptive 3D sensing system receives the coordinates of the area 824 of the real-world scene 822 from the XR application, where the virtual object of the XR user interface is to be rendered in the area 824. The adaptive 3D sensing system generates the attention mask 832 based on the coordinates of the area in which the virtual object is to be rendered. The adaptive 3D sensing system generates projector command data based on the attention mask 832, and transmits the projector command data to the projector 816 to command the projector 816 to project the distributed laser beam 826, thereby selectively sensing the area 824 of the real-world scene 822 in which the virtual object is to be rendered.

[0087] In some examples, the adaptive 3D sensing system generates an attention mask 832 based on an area of ​​the real-world scene 822 that has not been mapped to a 3D model of the real-world scene 822, which is maintained by the XR system that provides the XR user interface to the user. The 3D model allows the glasses 100 to visually place virtual objects relative to physical objects within the user's field of view. For example, as the user of the XR system moves through the real-world scene, the XR system continuously generates a 3D model of the real-world scene. The tracking component of the XR system estimates the pose of a head-mounted device (e.g., glasses 100) worn by the user. The tracking component uses image data from one or more cameras (e.g., left camera 114 and right camera 116), and associated position data provided by one or more positioning components of the XR system to track the position and determine the pose of the glasses 100 relative to a reference frame (e.g., a real-world scene environment). The tracking component continuously collects and uses updated sensor data describing the movement of the glasses 100 to generate and update the 3D model. The XR system detects that a user of the XR system has entered a new area of ​​the real-world scene based on current position data and previous position data included in the 3D model. In response to determining that the user has entered the new area of ​​the real-world scene, the XR system generates new 3D sensing data mapping the new area of ​​the real-world scene using adaptive 3D sensing, and adds the new area of ​​the real-world scene to the 3D model based on the new 3D sensing data.

[0088] In operation 806, the adaptive 3D sensing system commands the projector 816 to send the distributed laser beam 826 to one or more designated areas 824 of the real-world scene 822 based on the attention mask 832. For example, the adaptive 3D sensing system generates projector command data for the projector 816 based on the attention mask. The adaptive 3D sensing system transmits the projector command data to the projector 816. The projector 816 receives the projector command data from the adaptive 3D sensing system and sends the distributed laser beam 826 to one or more designated areas 824 of the real-world scene 822 that is undergoing adaptive 3D sensing based on the areas designated in the attention mask 832.

[0089] In some examples, the projector 816 includes a DOE 840, a laser 844, and a movable MEMS mirror 842 operable to deflect a laser beam generated by the laser 844 through the DOE 840. The DOE generates a distributed laser beam 850 having a distributed laser beam pattern 846 using the laser beam, and sends the distributed laser beam 850 into one or more regions 830 of the real-world scene 822 based on the attention mask 832.

[0090] In some examples, projector 816 includes a laser 836 and an SLM 834. SLM 834 is operable to generate a distributed laser beam pattern 838 from a laser beam generated by laser 836 and send the distributed laser beam pattern 838 to one or more designated areas 830 of real-world scene 822 based on attention mask 832.

[0091] In some examples, in operation 808, the adaptive 3D sensing system uses camera 818 and camera 820 to capture 3D sensing image data of real-world scene 822 in real time when projector 816 sends a distributed laser beam into one or more areas 824. In operation 810, the adaptive 3D sensing system generates 3D sensing data of one or more areas 824 of real-world scene 822 based on the 3D sensing image data captured by the one or more cameras.

[0092] In some examples, in operation 808, the adaptive 3D sensing system uses one or more depth sensors 828 to capture 3D depth sensor data of the real-world scene 822 in real time as the projector 816 sends the distributed laser beam into the one or more areas 824. In operation 810, the adaptive 3D sensing system generates 3D sensing data of the area 824 of the real-world scene 822 based on the 3D depth sensor data captured by the one or more depth sensors 532.

[0093] In some examples, the real world scene 872 will include two or more regions, such as region 862 and region 864, including region 830 to be subjected to adaptive 3D sensing. The adaptive 3D sensing system calculates an attention mask 874 having mask data (e.g., mask data 852 and mask data 848 corresponding to region 862 and region 864). The adaptive 3D sensing system sends a distributed laser beam having a distributed laser beam pattern into region 830 based on the attention mask 874. In the case where the adaptive 3D sensing system uses an SLM-based projector, the pattern of 830 includes a distributed laser beam pattern, such as distributed laser beam pattern 860 and distributed laser beam pattern 858. In the case where the adaptive 3D sensing system uses a MEMS+DOE-based projector, the pattern of 830 includes a distributed laser beam pattern, such as distributed laser beam pattern 856 and distributed laser beam pattern 854.

[0094] In some examples, during the exposure of the depth frame, the MEMS mirror 842 does not rotate and deflects the laser pattern to one location considered to be of interest. In some examples, during the exposure of the depth frame, the MEMS mirror 842 rotates to deflect the laser pattern to several locations while capturing the depth frame; here, in some examples, the camera captures several frames, each of which corresponds to a pattern location, thereby improving the signal-to-noise ratio compared to capturing one frame when the ambient lighting noise is elevated.

[0095] In operation 812 , the adaptive 3D sensing system transmits the 3D sensing data to the XR application.

[0096] In operation 814 , the XR application uses the 3D sensing data to provide or update an XR user interface provided by the XR application to the user.

[0097] Fig. 9 900 is a block diagram illustrating a software architecture 904 that may be installed on any one or more of the devices described herein. The software architecture 904 is supported by hardware, such as a machine 902 including a processor 920, a memory 926, and an I / O component interface 938. In this example, the software architecture 904 may be conceptualized as a stack of layers in which each layer provides specific functionality. The software architecture 904 includes layers such as an operating system 912, a library 908, a framework 910, and an application 906. In operation, the application 906 invokes an API call 950 through the software stack and receives a message 952 in response to the API call 950.

[0098] The operating system 912 manages hardware resources and provides public services. The operating system 912 includes, for example, a kernel 914, services 916, and drivers 922. The kernel 914 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 914 provides memory management, processor management (e.g., scheduling), component management, networking and security settings, and other functions. Services 916 can provide other public services to other software layers. Drivers 922 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 922 may include display drivers, camera drivers, or Low-power drivers, Flash drivers, Serial communications drivers (e.g., Universal Serial Bus (USB) drivers), Drivers, audio drivers, power management drivers, etc.

[0099] The library 908 provides a low-level common infrastructure used by the application 906. The library 908 may include a system library 918 (e.g., a C standard library) that provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the library 908 may include an API library 924, such as a media library (e.g., a library for supporting presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., an OpenGL framework for rendering two-dimensional (2D) and three-dimensional (3D) graphics content on a display, GLMotif for implementing a user interface), an image feature extraction library (e.g., OpenIMAJ), a database library (e.g., SQLite providing various relational database functions), a web library (e.g., WebKit providing web browsing functions), etc. The library 908 may also include various other libraries 928 to provide many other APIs to the application 906 .

[0100] The framework 910 provides a high-level common infrastructure used by the applications 906. For example, the framework 910 provides various graphical user interface (GUI) functions, advanced resource management, and advanced location services. The framework 910 can provide a wide range of other APIs that can be used by the applications 906, some of which can be specific to a particular operating system or platform.

[0101] In an example, applications 906 may include a home application 936, a contacts application 930, a browser application 932, a book reader application 934, a location application 942, a media application 944, a messaging application 946, a game application 948, and a variety of other applications such as third-party applications 940. Application 906 is a program that executes the functions defined in the program. Various programming languages ​​can be used to create one or more of the applications 906 constructed in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C language or assembly language). In a specific example, third-party applications 940 (e.g., provided by an entity other than the vendor of a particular platform using ANDROID TM or IOS TM Software Development Kit (SDK) can be used to develop applications on platforms such as IOS TM ANDROID TM , Mobile software running on the mobile operating system of the Phone or other mobile operating systems. In this example, third-party applications 940 can call API calls 950 provided by the operating system 912 to facilitate the functions described herein.

[0102] Fig.10 1 is a block diagram showing a networked system 1000 including details of the glasses 100 according to some examples. The networked system 1000 includes the glasses 100, a mobile computing system 1026, and a server system 1032. The mobile computing system 1026 may be a smart phone, a tablet computer, a tablet phone, a laptop computer, an access point, or any other such device capable of connecting to the glasses 100 using a low power wireless connection 1036 and / or a high speed wireless connection 1034. The mobile computing system 1026 is connected to the server system 1032 via a network 1030. The network 1030 may include any combination of wired and wireless connections. The server system 1032 may be one or more computing devices that are part of a service or network computing system. The mobile computing system 1026 and any elements of the server system 1032 and the network 1030 may use Fig. 9 and Figure 3 The details of the software architecture 904 or machine 300 described respectively are implemented.

[0103] The glasses 100 include a data processor 1002, a display 1010, one or more cameras 1008, and additional input / output elements 1016. The input / output elements 1016 may include a microphone, an audio speaker, a biometric sensor, additional sensors, or additional display elements integrated with the data processor 1002. Fig. 9 and Figure 3 Examples of input / output elements 1016 are further discussed. For example, input / output elements 1016 may include any I / O device interface 306, including output component interface 328, motion component interface 336, etc. Figure 2 Examples of display 1010 are discussed in . In the particular examples described herein, display 1010 includes displays for the left and right eyes of a user.

[0104] The data processor 1002 includes an image processor 1006 (eg, a video processor), a GPU and display driver 1038, a tracking component 1040, an interface 1012, low power circuits 1004, and high speed circuits 1020. The components of the data processor 1002 are interconnected by a bus 1042.

[0105] The interface 1012 refers to any source of user commands provided to the data processor 1002. In one or more examples, the interface 1012 is a physical button that, when pressed, sends a user input signal from the interface 1012 to the low-power processor 1014. The low-power processor 1014 can process pressing such a button and then immediately releasing it as a request to capture a single image, and vice versa. The low-power processor 1014 can process pressing such a button within a first time period as a request to capture video data when the button is pressed and stop video capture when the button is released, wherein the video captured when the button is pressed is stored as a single video file. Alternatively, pressing the button for a long period of time can capture a still image. In some examples, the interface 1012 can be any mechanical switch or physical interface capable of accepting user input, wherein the user input is associated with requesting data from the camera 1008. In other examples, the interface 1012 can have a software component, or can be associated with commands received wirelessly from other sources (e.g., from the mobile computing system 1026).

[0106] Image processor 1006 includes circuitry for receiving signals from camera 1008 and processing those signals from camera 1008 into a format suitable for storage in memory 1024 or for transmission to mobile computing system 1026. In one or more examples, image processor 1006 (e.g., a video processor) includes a microprocessor integrated circuit (IC) customized for processing sensor data from camera 1008, and volatile memory used by the microprocessor in operation.

[0107] The low power circuit 1004 includes a low power processor 1014 and a low power wireless circuit 1018. These elements of the low power circuit 1004 can be implemented as separate elements, or can be implemented on a single IC as part of a single system on a chip. The low power processor 1014 includes logic for managing the other elements of the glasses 100. As described above, for example, the low power processor 1014 can accept user input signals from the interface 1012. The low power processor 1014 can also be configured to receive input signals or command communications from the mobile computing system 1026 via the low power wireless connection 1036. The low power wireless circuit 1018 includes circuit elements for implementing a low power wireless communication system. Bluetooth TM Smart, also known as Bluetooth h TM Low power consumption is a standard implementation of a low power wireless communication system that may be used to implement the low power wireless circuit 1018. In other examples, other low power communication systems may be used.

[0108] High-speed circuitry 1020 includes a high-speed processor 1022, a memory 1024, and a high-speed wireless circuit 1028. High-speed processor 1022 may be any processor capable of managing high-speed communications and operations of any general-purpose computing system used by data processor 1002. High-speed processor 1022 includes processing resources used to manage high-speed data transmission over high-speed wireless connection 1034 using high-speed wireless circuitry 1028. In some examples, high-speed processor 1022 executes an operating system such as a LINUX operating system or a program such as a UNIX operating system. Fig. 9 The high-speed processor 1022, which executes the software architecture of the data processor 1002, is used to manage data transmission with the high-speed wireless circuit 1028, in addition to any other duties. In some examples, the high-speed wireless circuit 1028 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, which is also referred to as Wi-Fi herein. In other examples, the high-speed wireless circuit 1028 can implement other high-speed communication standards.

[0109] The memory 1024 includes any storage device capable of storing camera data generated by the camera 1008 and the image processor 1006. Although the memory 1024 is shown as being integrated with the high-speed circuit 1020, in other examples, the memory 1024 may be a separate, independent element of the data processor 1002. In some such examples, electrical wiring may provide a connection from the image processor 1006 or the low-power processor 1014 to the memory 1024 through a chip including the high-speed processor 1022. In other examples, the high-speed processor 1022 may manage addressing of the memory 1024 so that the low-power processor 1014 initiates the high-speed processor 1022 whenever a read or write operation involving the memory 1024 is needed.

[0110] The tracking component 1040 estimates the pose of the glasses 100. For example, the tracking component 1040 uses image data and associated inertial data and GPS data from the camera device 1008 and the positioning component interface 340 to track the position and determine the pose of the glasses 100 relative to a reference frame (e.g., a real-world scene environment). The tracking component 1040 continuously collects and uses updated sensor data describing the movement of the glasses 100 to determine an updated three-dimensional pose of the glasses 100, which indicates changes in relative position and orientation relative to physical objects in the real-world scene environment. The tracking component 1040 allows the glasses 100 to visually place virtual objects relative to physical objects within the user's field of view via the display 1010.

[0111] When the glasses 100 operate in a conventional augmented reality mode, the GPU and display driver 1038 may use the pose of the glasses 100 to generate frames of virtual content or other content to be presented on the display 1010. In this mode, the GPU and display driver 1038 generate updated frames of virtual content based on the updated three-dimensional pose of the glasses 100, which reflects changes in the position and orientation of the user relative to physical objects in the user's real-world scene environment.

[0112] One or more functions or operations described herein may also be performed in an application resident on the glasses 100 or on the mobile computing system 1026 or on a remote server. For example, one or more functions or operations described herein may be performed by one of the applications 906 (e.g., the messaging application 946).

[0113] Fig.111 is a block diagram illustrating an example interactive system 1100 that facilitates interaction over a network (e.g., exchanging text messages, conducting text, audio, and video calls, or playing games). The interactive system 1100 includes multiple computing systems 1102, each of which hosts multiple applications, including interactive clients 1104 and other applications 1106. Each interactive client 1104 is communicatively coupled to other instances of the interactive client 1104 (e.g., hosted on respective other computing systems 1102), an interactive server system 1110, and a third-party server 1112 via one or more communication networks including a network 1108 (e.g., the Internet). The interactive client 1104 can also communicate with the locally hosted application 1106 using an application program interface (API).

[0114] Each computing system 1102 may include one or more user devices, such as a mobile device 1114 , a head-mounted XR system 1116 , and a computer client device 1118 , which are communicatively connected to exchange data and messages.

[0115] The interactive clients 1104 interact with other interactive clients 1104 and with the interactive server system 1110 via the network 1108. The data exchanged between the interactive clients 1104 (e.g., interaction 1120) and between the interactive clients 1104 and the interactive server system 1110 includes functions (e.g., commands for invoking functions) and payload data (e.g., text, audio, video or other multimedia data).

[0116] The interactive server system 1110 provides server-side functions to the interactive clients 1104 via the network 1108. Although certain functions of the interactive system 1100 are described herein as being performed by the interactive clients 1104 or by the interactive server system 1110, it may be a design choice whether certain functions are located within the interactive clients 1104 or within the interactive server system 1110. For example, it may be technically preferred to initially deploy certain technologies and functions within the interactive server system 1110, but later migrate the technologies and functions to the interactive clients 1104 where the computing system 1102 has sufficient processing power.

[0117] The interactive server system 1110 supports various services and operations provided to the interactive clients 1104. Such operations include sending data to the interactive clients 1104, receiving data from the interactive clients 1104, and processing data generated by the interactive clients 1104. The data may include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. The data exchange within the interactive system 1100 is invoked and controlled by functions available through the user interface (UI) of the interactive client 1104.

[0118] Turning now specifically to the interaction server system 1110, an application program interface (API) server 1122 is coupled to and provides a programming interface to the interaction server 1124, making the functionality of the interaction server 1124 accessible to the interaction clients 1104, other applications 1106, and third-party servers 1112. The interaction server 1124 is communicatively coupled to a database server 1126, thereby facilitating access to a database 1128 that stores data associated with interactions processed by the interaction server 1124. Similarly, a web server 1130 is coupled to the interaction server 1124 and provides a web-based interface to the interaction server 1124. To this end, the web server 1130 handles incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0119] The application program interface (API) server 1122 receives and sends interaction data (e.g., commands and message payloads) between the interaction server 1124 and the computing system 1102 (and, for example, the interaction client 1104 and other applications 1106) and the third-party server 1112. Specifically, the application program interface (API) server 1122 provides a set of interfaces (e.g., routines and protocols) that the interaction client 1104 and other applications 1106 can call or query to call the functions of the interaction server 1124. The application program interface (API) server 1122 exposes various functions supported by the interaction server 1124, including: account registration; login functionality; sending interaction data from a particular interaction client 1104 to another interaction client 1104 via the interaction server 1124; transferring media files (e.g., images or videos) from the interaction client 1104 to the interaction server 1124; setting up a collection of media data (e.g., a story); retrieving a friend list of a user of the computing system 1102; retrieving messages and content; adding and removing entities (e.g., friends) from an entity graph (e.g., a social graph); locating friends in a social graph; and opening application events (e.g., related to the interaction client 1104).

[0120] Returning to the interactive client 1104, the features and functions of the external resource (e.g., the linked application 1106 or applet) are available to the user via the interface of the interactive client 1104. In this context, "external" means that the application 1106 or applet is outside the interactive client 1104. External resources are usually provided by a third party, but can also be provided by the creator or provider of the interactive client 1104. The interactive client 1104 receives a user selection of an option to launch or access the features of such an external resource. The external resource can be an application 1106 installed on the computing system 1102 (e.g., a "local app"), or a small-scale version of an application hosted on the computing system 1102 or located at the remote end of the computing system 1102 (e.g., on a third-party server 1112) (e.g., a "applet"). The small-scale version of the application includes a subset of the features and functions of the application (e.g., the full-scale, local version of the application) and is implemented using a markup language document. In some examples, a small-scale version of an application (e.g., a "mini program") is a web-based markup language version of the application and is embedded in the interactive client 1104. In addition to using markup language documents (e.g., .*ml files), mini programs can also include scripting languages ​​(e.g., .*js files or .json files) and style sheets (e.g., .*ss files).

[0121] In response to receiving a user selection of an option to launch or access a feature of an external resource, the interactive client 1104 determines whether the selected external resource is a web-based external resource or a locally installed application 1106. In some cases, the application 1106 installed locally on the computing system 1102 can be independent of the interactive client 1104 and launched separately from the interactive client 1104, such as by selecting an icon corresponding to the application 1106 on a home screen of the computing system 1102. A small-scale version of such an application can be launched or accessed via the interactive client 1104, and in some examples, no portion of the small-scale application can be accessed outside of the interactive client 1104, or a limited portion of the small-scale application can be accessed outside of the interactive client 1104. The small-scale application can be launched by the interactive client 1104 receiving, for example, a markup language document associated with the small-scale application from the third-party server 1112 and processing such a document.

[0122] In response to determining that the external resource is a locally installed application 1106, the interactive client 1104 instructs the computing system 1102 to launch the external resource by executing locally stored code corresponding to the external resource. In response to determining that the external resource is a web-based resource, the interactive client 1104 communicates with the third-party server 1112 (for example) to obtain a markup language document corresponding to the selected external resource. The interactive client 1104 then processes the obtained markup language document to present the web-based external resource within the user interface of the interactive client 1104.

[0123] Interactive client 1104 can notify the user of computing system 1102 or other users (e.g., "friends") related to such user of one or more external resources of activities taking place. For example, interactive client 1104 can provide participants of a conversation (e.g., a chat session) in interactive client 1104 with notifications related to external resources currently or recently used by one or more members in a user group. One or more users can be invited to join an active external resource, or an external resource that was recently used but is currently inactive (in the friend group) can be started. External resources can provide participants in a conversation who each use a corresponding interactive client 1104 with the ability to share items, conditions, states, or locations in an external resource with one or more members in a user group in a chat session. The shared items can be interactive chat cards, and members of a chat can interact with the interactive chat cards, such as for starting a corresponding external resource, viewing specific information in an external resource, or bringing members of a chat to a specific location or state in an external resource. Within a given external resource, a response message can be sent to a user on interactive client 1104. Based on the current context of the external resource, the external resource can selectively include different media items in the response.

[0124] The interactive client 1104 can present a list of available external resources (e.g., applications 1106 or applets) to the user to launch or access a given external resource. The list can be presented in the form of a context-sensitive menu. For example, icons representing different applications (or applets) of the application 1106 (or applets) can change based on how the user launches the menu (e.g., from a conversational interface or a non-conversational interface).

[0125] Additional examples include:

[0126] Example 1: A system comprises: one or more cameras; one or more projectors; one or more processors, the one or more processors being operably connected to the one or more cameras and the one or more projectors; and a memory, the memory being operably connected to the one or more processors, the memory storing instructions, which, when executed by the one or more processors, cause the system to perform operations, the operations comprising: capturing image data of a real-world scene using the one or more cameras; calculating an attention mask based on the image data; commanding the one or more projectors to send distributed laser beams into an area of ​​the real-world scene based on the attention mask; when one or more projectors send the distributed laser beams into the area of ​​the real-world scene, capturing three-dimensional (3D) sensing image data of the area of ​​the real-world scene using the one or more cameras; and generating 3D sensing data of the real-world scene based on the 3D sensing image data.

[0127] Example 2: The subject matter of Example 1 includes: wherein the one or more projectors are spatial light modulator (SLM) based projectors.

[0128] Example 3: The subject matter according to any one of Examples 1 to 2 includes: wherein the one or more projectors are micro-electromechanical systems (MEMS) + diffractive optical elements (DOE) based projectors.

[0129] Example 4: The subject matter of Example 3 includes: wherein the one or more MEMS+DOE based projectors generate the distributed laser beam to include a randomized pattern.

[0130] Example 5: The subject matter of any one of Examples 1 to 4 includes determining the area of ​​the real-world scene based on a depth estimation confidence value.

[0131] Example 6: The subject matter of any one of Examples 1 to 5 includes determining the area of ​​the real-world scene based on a position of a virtual object of an extended reality (XR) user interface of an XR application in the real-world scene.

[0132] Example 7: The subject matter of any one of Examples 1 to 6 includes: generating a 3D model of the real-world scene based on the image data of the real-world scene; determining that the area of ​​the real-world scene is not mapped into the 3D model based on the position data of the system; and specifying the area of ​​the real-world scene in response to determining that the area of ​​the real-world scene is not mapped into the 3D model.

[0133] Example 8: The subject matter of any one of Examples 1 to 7 includes: wherein the system also includes a head-mounted device.

[0134] Example 9: The subject matter of any one of Examples 1 to 8 includes: wherein the system further includes a smart phone.

[0135] "Carrier signal" refers to any intangible medium that can store, encode or carry instructions for execution by a machine and includes digital or analog communications signals or other intangible media to facilitate communication of such instructions. Instructions may be sent or received over a network using a transmission medium via a network interface device.

[0136] "Client device" refers to any machine that connects to a communication network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smart phone, a tablet computer, an ultrabook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, or any other communication device that a user may use to access a network.

[0137] "Communications network" means one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, The coupling may be a network, another type of network, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other types of cellular or wireless couplings. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world wide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other data transmission technologies defined by various standard setting organizations, other long distance protocols, or other data transmission technologies.

[0138] "Machine-readable media" refers to both machine storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "machine-readable medium," and "device-readable medium" refer to the same thing and may be used interchangeably in this disclosure.

[0139] "Machine storage media" refers to a single or multiple storage devices and / or media (e.g., centralized or distributed databases, and / or associated caches and servers) that store executable instructions, routines, and / or data. The term includes, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and / or device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage media," "device storage media," and "computer storage media" refer to the same thing and may be used interchangeably in this disclosure. The terms "machine storage media," "computer storage media," and "device storage media" expressly exclude carrier waves, modulated data signals, and other such media, some of which are covered by the term "signal media."

[0140] A "processor" refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values ​​according to control signals (e.g., "commands," "opcodes," "machine codes," etc.), and produces associated output signals that are applied to operate a machine. For example, a processor may be a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor may also be a multi-core processor, having two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously.

[0141] "Signal medium" refers to any intangible medium that is capable of storing, encoding or carrying machine-executable instructions, and "signal medium" includes digital or analog communication signals or other intangible media to facilitate the communication of software or data. The term "signal medium" may be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" refers to a signal in which one or more of the characteristics of the signal is set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" refer to the same thing and may be used interchangeably in this disclosure.

[0142] Changes and modifications may be made to the disclosed examples without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.

Claims

1. A system comprising: one or more camera devices; one or more projectors; one or more processors operably connected to the one or more cameras and the one or more projectors; as well as a memory operatively connected to the one or more processors, the memory storing instructions that, when executed by the one or more processors, cause the system to perform operations comprising: capturing image data of a real-world scene using the one or more cameras; calculating an attention mask based on the image data; commanding the one or more projectors to send distributed laser beams into areas of the real-world scene based on the attention mask; When one or more projectors transmit the distributed laser beam into the area of ​​the real-world scene, capturing three-dimensional (3D) sensing image data of the area of ​​the real-world scene using the one or more cameras; and 3D sensing data of the real world scene is generated based on the 3D sensing image data.

2. The system according to claim 1, wherein: The one or more projectors are spatial light modulator (SLM) based projectors.

3. The system according to claim 1, wherein: The one or more projectors are micro-electromechanical systems (MEMS) + diffractive optical elements (DOE) based projectors.

4. The system according to claim 3, wherein: The one or more MEMS+DOE based projectors generate the distributed laser beam to include a randomized pattern.

5. The system according to claim 1, wherein: The operations also include: The area of ​​the real-world scene is determined based on a depth estimate confidence value.

6. The system according to claim 1, wherein: The operations also include: The area of ​​the real-world scene is determined based on a position of a virtual object of an extended reality (XR) user interface of an XR application in the real-world scene.

7. The system according to claim 1, wherein: The operations also include: generating a 3D model of the real-world scene based on the image data of the real-world scene; determining, based on position data of the system, that the area of ​​the real-world scene is not mapped in the 3D model; and The area of ​​the real-world scene is specified in response to determining that the area of ​​the real-world scene is not mapped in the 3D model.

8. The system according to claim 1, wherein: The system also includes a head mounted device.

9. The system according to claim 1, wherein: The system also includes a smart phone.

10. A computer-implemented method comprising: capturing image data of a real-world scene using one or more cameras of the system; calculating an attention mask based on the image data; commanding one or more projectors of the system to send distributed laser beams into areas of the real-world scene based on the attention mask; When one or more projectors transmit the distributed laser beam into the area of ​​the real-world scene, capturing three-dimensional (3D) sensing image data of the area of ​​the real-world scene using the one or more camera devices, and 3D sensing data of the real world scene is generated based on the 3D sensing image data.

11. The computer-implemented method for a system according to claim 10, wherein: The one or more projectors are spatial light modulator (SLM) based projectors.

12. The computer-implemented method of claim 10, wherein: The one or more projectors are micro-electromechanical systems (MEMS) + diffractive optical elements (DOE) based projectors.

13. The computer-implemented method of claim 12, wherein: The one or more MEMS+DOE based projectors generate the distributed laser beam to include a randomized pattern.

14. The computer-implemented method of claim 10, further comprising: The area of ​​the real-world scene is determined based on a depth estimate confidence value.

15. The computer-implemented method of claim 10, further comprising: The area of ​​the real-world scene is determined based on a position of a virtual object of an extended reality (XR) user interface of an XR application in the real-world scene.

16. The computer-implemented method of claim 10, further comprising: generating a 3D model of the real-world scene based on the image data of the real-world scene; determining, based on position data of the system, that the area of ​​the real-world scene is not mapped in the 3D model; as well as The area of ​​the real-world scene is specified in response to determining that the area of ​​the real-world scene is not mapped in the 3D model.

17. The computer-implemented method of claim 10, wherein: The system also includes a head mounted device.

18. The computer-implemented method of claim 10, wherein: The system also includes a smart phone.

19. A non-transitory computer storage medium, the computer readable storage medium comprising instructions, the instructions, when executed by a computer, causing the computer to perform operations comprising: capturing image data of a real-world scene using one or more cameras; calculating an attention mask based on the image data; commanding one or more projectors to send distributed laser beams into areas of the real-world scene based on the attention mask; capturing three-dimensional (3D) sensing image data of the area of ​​the real-world scene using the one or more camera devices when the one or more projectors transmit the distributed laser beam into the area of ​​the real-world scene; as well as 3D sensing data of the real world scene is generated based on the 3D sensing image data.

20. The non-transitory computer storage medium of claim 19, wherein: The operations also include: The area of ​​the real-world scene is determined based on a depth estimate confidence value.

Citation Information

Patent Citations

  • Static configuration of accelerator card security modes

    US20220100840A1

Cited By

  • Energy-efficient adaptive 3D sensing

    US12498579B2