Proximity-based protocol for implementing multi-user augmented reality (XR) experience

Through a proximity-based protocol, user equipment shares image features with host equipment for authentication and positioning, solving the complexity of XR room management in the existing technology, realizing a multi-user XR room experience under a unified framework, and improving user experience and security.

CN120476432APending Publication Date: 2025-08-12QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380090723.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-12
Filing Date
2023-12-08
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, users need to install customized applications for each XR room and create separate user accounts, resulting in poor user experience and limiting the number of XR rooms that can be joined, and lacking a unified framework to manage sharing and navigation of multi-user XR rooms.

Method used

Adopting a proximity-based protocol, the user equipment communicates with the user equipment through a static host device, scans and generates a 3D scene map, and the user equipment shares image features with the host device for authentication and positioning, providing a unified framework to manage the joining and mastering of multi-user XR rooms.

Benefits of technology

It realizes a seamless multi-user XR room experience, reduces the hassle of users installing customized applications, simplifies user account management, improves security and control, and supports applications in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476432A_ABST
    Figure CN120476432A_ABST
Patent Text Reader

Abstract

Systems and techniques for implementing a multi-user augmented reality (XR) experience are described herein. In one exemplary example, a user device may receive, from a host device, a message associated with a user including a prompt to join an XR room hosted by the host device. The user device may connect to the host device based on the message. The user device may obtain images of a three-dimensional (3D) scene of a physical environment, and may send the images to the host device. The user device may receive composite content from the host device. A virtual representation of the user may be located relative to the XR room based on features in the images matching features of the 3D scene of the physical environment. The user equipment may render the synthetic content of the XR room based on a pose of the apparatus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to processing virtual content of a virtual environment or a partially virtual environment. For example, aspects of the present disclosure include systems and techniques for providing a proximity-based protocol for enabling multi-user extended reality (XR) experiences. Background Art

[0002] Extended reality (XR) (e.g., virtual reality, augmented reality, mixed reality) systems can provide users with virtual experiences by immersing them in a completely virtual environment (composed of virtual content), and / or can provide users with augmented or mixed reality experiences by combining the real world or physical environment with the virtual environment.

[0003] Users may desire to enhance their physical three-dimensional (3D) space with synthetic content. These users may want to share their augmented 3D scenes with others (e.g., each of which may be referred to as an XR room). Sharing of XR rooms may have a variety of applications, such as for classrooms, restaurants, businesses, factories, homes (e.g., living rooms and kitchens), museums, etc. These XR rooms may be multifunctional by providing users with interactive information, entertainment, and / or synthetic light experiences. Therefore, in the case of such ubiquitous XR rooms, a unified framework may be needed to help users join, host, and navigate through XR rooms. Summary of the Invention

[0004] The following presents a simplified summary of one or more aspects disclosed herein. Therefore, the following summary should neither be considered an exhaustive overview of all contemplated aspects nor be considered to identify key or critical elements related to all contemplated aspects or to delineate the scope associated with any particular aspect. Therefore, the sole purpose of the following summary is to present certain concepts related to one or more aspects of the mechanisms disclosed herein in a simplified form prior to the detailed description presented below.

[0005] Systems and techniques for providing a multi-user extended reality (XR) experience are described. According to at least one illustrative example, a method for providing an XR experience is provided. The method includes: receiving, by a user device associated with a user, a message from a host device including a prompt to join an XR room hosted by the host device; connecting to the host device based on the message; obtaining, by the user device, images of a three-dimensional (3D) scene of a physical environment; sending, by the user device, the images to the host device; receiving, by the user device, synthetic content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on features in the images matching features of the 3D scene of the physical environment; and rendering, by the user device, the synthetic content of the XR room based on a pose of the user device.

[0006] In another illustrative example, an apparatus associated with a user for providing an extended reality (XR) experience is provided. The apparatus includes at least one memory and at least one processor, the at least one processor being coupled to the at least one memory and configured to: receive a message associated with the user from a host device including a prompt to join an XR room hosted by the host device; connect to the host device based on the message; obtain images of a three-dimensional (3D) scene of a physical environment; send the images to the host device; receive synthetic content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on features in the images matching features of the 3D scene of the physical environment; and render the synthetic content of the XR room based on a pose of the apparatus.

[0007] In another illustrative example, a non-transitory computer-readable medium of a user device associated with a user is provided. The non-transitory computer-readable medium has instructions that, when executed by at least one processor, cause the at least one processor to: receive a message associated with the user from a host device including a prompt to join an XR room hosted by the host device; connect to the host device based on the message; obtain images of a three-dimensional (3D) scene of a physical environment; send the images to the host device; receive composite content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on features in the images matching features of the 3D scene of the physical environment; and render the composite content of the XR room based on a pose of the device.

[0008] In another illustrative example, an apparatus associated with a user for providing an extended reality (XR) experience is provided. The apparatus includes: means for receiving a message including a prompt to join an XR room hosted by the host device; means for connecting to the host device based on the message; means for obtaining images of a three-dimensional (3D) scene of a physical environment; means for sending the images to the host device; means for receiving synthesized content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on features in the images matching features of the 3D scene of the physical environment; and means for rendering the synthesized content of the XR room based on a pose of the apparatus.

[0009] In some aspects, one or more of the devices described herein are, are part of, and / or include an XR device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device (e.g., a mobile phone or other mobile device), a wearable device, a wireless communication device, a camera, a personal computer, a laptop computer, a vehicle or a computing device or component of a vehicle, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above-mentioned device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).

[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. This subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0011] The foregoing and other features and aspects will become more apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Illustrative examples of the present application are described in detail below with reference to the following drawings:

[0013] Figure 1 is a diagram illustrating an example of an extended reality (XR) system according to aspects of the present disclosure;

[0014] Figure 2 is a diagram illustrating an example of a three-dimensional (3D) collaborative virtual environment according to aspects of the present disclosure;

[0015] Figure 3 is a block diagram illustrating an example architecture of an XR system according to some examples;

[0016] Figure 4 is a block diagram illustrating the architecture of a simultaneous localization and mapping (SLAM) device according to some examples;

[0017] Figure 5 is a flow chart illustrating an example of a process for 3D reconstruction and generation of a feature database according to some examples of the present disclosure;

[0018] Figure 6 is a flow chart illustrating an example of a process for authenticating and locating a user using a head mounted device (HMD) according to some examples of the present disclosure;

[0019] Figure 7 is a diagram illustrating an example in which a user located within proximity of an XR room is able to join the XR room according to some examples of the present disclosure;

[0020] Figure 8 is a diagram illustrating an example of an enterprise with corresponding XR rooms according to some examples of the present disclosure;

[0021] Figure 9 is a diagram illustrating an example of users collaboratively hosting an XR room according to some examples of the present disclosure;

[0022] Figure 10 is a diagram illustrating an example of a user interacting with synthetic content in an XR room according to some examples of the present disclosure;

[0023] Figure 11 is a diagram illustrating an example of an XR classroom according to some examples of the present disclosure;

[0024] Figure 12 is a diagram illustrating an example of an XR sub-room within an XR classroom according to some examples of the present disclosure;

[0025] Figure 13 is a flowchart illustrating an example of a process for implementing a multi-user XR experience according to some examples of the present disclosure; and

[0026] Figure 14 is a diagram illustrating an example of a computing system according to aspects of the present disclosure. DETAILED DESCRIPTION

[0027] Certain aspects of the present disclosure are provided below. Some of these aspects can be applied independently, and some of them can be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for explanation purposes to provide a thorough understanding of various aspects of the application. However, it will be apparent that various aspects can be implemented without these specific details. Each drawing and description is not intended to be restrictive.

[0028] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0029] An extended reality (XR) system or device can provide an XR experience to a user by presenting virtual content to the user (e.g., for a fully immersive experience) and / or can combine a view of the real world or physical environment with a display of a virtual environment (composed of virtual content). The real world environment may include real-world objects (also called physical objects) such as people, vehicles, buildings, tables, chairs, and / or other real-world objects or physical objects. As used herein, the terms XR system and XR device are used interchangeably. Examples of XR systems or devices include head-mounted displays (HMDs), smart glasses (e.g., AR glasses, MR glasses, etc.), and the like.

[0030] XR systems may include VR systems that facilitate interaction with virtual reality (VR) environments, AR systems that facilitate interaction with augmented reality (AR) environments, MR systems that facilitate interaction with mixed reality (MR) environments, and / or other XR systems. For example, VR provides a fully immersive experience in a three-dimensional (3D) computer-generated VR environment or video that depicts a virtual version of a real-world environment. VR content may, in some cases, include VR video, which may be captured and rendered at very high quality, potentially providing a truly immersive virtual reality experience. Virtual reality applications may include games, training, education, sports videos, online shopping, and the like. VR content may be rendered and displayed using a VR system or device that completely covers the user's eyes during the VR experience, such as a VR HMD or other VR head-mounted device.

[0031] AR is a technology that provides virtual or computer-generated content (referred to as AR content) over a user's view of a physical, real-world scene or environment. AR content can include any virtual content, such as video, images, graphical content, location data (e.g., Global Positioning System (GPS) data or other location data), sound, any combination thereof, and / or other augmented content. AR systems are designed to enhance (or augment) rather than replace a person's current perception of reality. For example, a user may see a real, stationary or moving physical object through an AR device display, but the user's visual perception of the physical object may be enhanced or improved by a virtual image of the object (e.g., a real-world car replaced by a virtual image of a DeLorean), by AR content added to the physical object (e.g., virtual wings added to a living animal), by AR content displayed relative to the physical object (e.g., informational virtual content displayed near a sign on a building, a virtual coffee cup virtually anchored to a real-world table in one or more images (e.g., placed on top of the real-world table), etc.), and / or by displaying other types of AR content. Various types of AR systems can be used for gaming, entertainment, and / or other applications.

[0032] MR technology can combine aspects of VR and AR to provide users with an immersive experience. For example, in an MR environment, real-world and computer-generated objects can interact (e.g., real people can interact with virtual people as if the virtual people were real people).

[0033] An XR environment can be interactive in a way that appears real or physical. As a user experiencing an XR environment (e.g., an immersive VR environment) moves in the real world, the rendered virtual content (e.g., images rendered in the virtual environment within the VR experience) also changes, giving the user the perception that they are moving within the XR environment. For example, the user can turn left or right, look up or down, and / or move forward or backward, thereby changing the user's viewpoint of the XR environment. The XR content presented to the user can change accordingly, making the user's experience in the XR environment as seamless as in the real world.

[0034] In some cases, the XR system may match the relative pose and movement of objects and devices in the physical world. For example, the XR system may use tracking information to calculate the relative pose of devices, objects, and / or features of the real-world environment in order to match the relative positioning and movement of devices, objects, and / or the real-world environment. In some cases, the XR system may use the pose and movement of one or more devices, objects, and / or the real-world environment to render content in a convincing manner relative to the real-world environment. Relative pose information may be used to match virtual content to the user's perceived motion and the spatiotemporal state of devices, objects, and the real-world environment. In some cases, the XR system may track parts of the user (e.g., the user's hands and / or fingertips) to allow the user to interact with items of virtual content.

[0035] As mentioned previously, users may desire to enhance the physical three-dimensional (3D) space around them with synthetic content. These users may want to share their augmented 3D scenes (e.g., XR rooms) with others. Sharing XR rooms can have a variety of applications, such as in classrooms, restaurants, businesses, factories, homes (e.g., living rooms and kitchens), museums, and more. These XR rooms can be multifunctional by providing users with interactive information, entertainment, and / or synthetic light experiences. In the case of such XR rooms, a unified framework can help enable users to navigate through XR rooms in their daily lives. Therefore, there should be a single protocol for users to host and join these XR rooms. Without a unified framework, each XR room may need to host a unique software application (app) that each user would need to manually install on their device (e.g., HMD), and user profiles may need to be created for each different app (e.g., for each different XR room). These apps can be annoying to users, significantly reducing their chances of entering an XR room. These apps may also effectively limit the number of XR rooms a single user can join.

[0036] Described herein are systems, apparatuses, electronic devices, methods (also referred to as processes), and computer-readable media (collectively referred to herein as "systems and techniques") for providing a proximity-based protocol for enabling multi-user XR experiences. Specifically, the systems and techniques provide a framework for hosting and joining multi-user XR rooms. In one or more aspects, a static host device (e.g., a wireless router) may be employed that communicates with a user's device (e.g., the user's AR / VR HMD) using a wireless protocol. The protocol may allow a user to smoothly transition between different XR rooms based on the user's proximity to the host device.

[0037] A user device (e.g., an HMD or other XR device associated with a user) located within proximity to a host device may connect to the host device as a manager. When operating as a manager, the user device may scan the entire 3D scene of its environment (e.g., a physical room) to capture images of the 3D scene, and the images captured by the user device may be shared with the host device. The images may be processed within the host device and / or offline (e.g., in a cloud such as a cloud server) to generate a 3D map of the scene (e.g., containing features such as points of interest, and the geometry of the scene). The user device may add synthetic objects and / or light to the scene from the obtained scene geometry when operating as a manager.

[0038] Other user devices (e.g., HMDs, etc.) within proximity of the host device (e.g., within communication range) may receive the broadcast message sent from the host device, and the host device may prompt these other user devices to join the XR room. In some cases, the host device may authenticate the user device before the user can join the XR room. The user device may transmit an image (including image features) of the physical environment surrounding the user device obtained to the host device. The image features in the image may be matched with scene features at the host device to locate the user (or user device) relative to the XR room. Positioning of the user device may be performed periodically to avoid positioning drift (e.g., drift of the position of a particular user within the XR environment of the XR room). The host device may communicate the geometry of the synthesized content to the user device. Each of these user devices may then use the head / eye pose of the user associated with the user device to render the synthesized content to the visual perspective camera view of the user device.

[0039] The disclosed systems and techniques have numerous advantages. One advantage is that the systems and techniques provide a unified framework for users to join and host multi-user XR rooms. This universal framework can be adopted as a standard for any XR device without requiring significant overhead. The unified framework relieves users of the cumbersome task of installing custom apps for each XR room. The unified framework also relieves users from having to create separate user accounts for each app in each XR room.

[0040] Another advantage is that the systems and techniques can make it easier for users to create XR rooms and share them with other users, which can enable a wide variety of applications in a variety of different scenarios, including but not limited to classrooms, restaurants, bars, businesses, factories, personalized homes, and museums. An additional advantage is that the systems and techniques allow multiple users to seamlessly participate in creating an XR room. Additionally, another advantage is that the systems and techniques employ the use of local host devices, which can ensure better security (e.g., by leveraging multimodal authentication of the user) and give the user greater control when setting up the XR room.

[0041] Various aspects of the application will be described with respect to the accompanying drawings.

[0042] Figure 1 An example of an extended reality system 100 is illustrated. As shown, the extended reality system 100 includes a device 105, a network 120, and a communication link 125. In some cases, the device 105 may be an extended reality (XR) device, which generally implements aspects of extended reality, including virtual reality (VR), augmented reality (AR), mixed reality (MR), etc. A system that includes the device 105, the network 120, or other elements of the extended reality system 100 may be referred to as an extended reality system.

[0043] Device 105 may overlay virtual objects (e.g., synthetic content) with real-world objects in view 130. For example, view 130 may generally relate to visual input to user 110 via device 105, a display generated by device 105, a virtual object configuration generated by device 105, and the like. For example, view 130-A may relate to visible real-world objects (also referred to as physical objects) at some initial time, along with visible virtual objects overlaid on or coexisting with the real-world objects. View 130-B may relate to visible real-world objects at some later time, along with visible virtual objects overlaid on or coexisting with the real-world objects. As discussed herein, shifting from view 130-A to view 130-B at 135 due to head motion 115 may result in positioning differences in real-world objects (e.g., and thus, overlaid virtual objects). In another example, view 130-A may relate to a completely virtual environment or scene at the initial time, and view 130-B may relate to a virtual environment or scene at a later time.

[0044] Generally speaking, device 105 may generate, display, project, etc., virtual objects and / or a virtual environment to be viewed by user 110 (e.g., where a portion of the virtual objects and / or the virtual environment may be displayed based on a predicted head pose of user 110 according to the techniques described herein). In some examples, device 105 may include a transparent surface (e.g., optical glass) such that virtual objects may be displayed on the transparent surface so as to overlay the virtual objects on real-world objects viewed through the transparent surface. Additionally or alternatively, device 105 may project virtual objects onto the real-world environment. In some cases, device 105 may include a camera and may display both real-world objects (e.g., as frames or images captured by the camera) and virtual objects overlaid on the displayed real-world objects. In various examples, device 105 may include aspects of a virtual reality headset, a head-mounted display (HMD), smart glasses, a live feed video camera, a GPU, one or more sensors (e.g., such as one or more IMUs, image sensors, microphones, etc.), one or more output devices (e.g., such as speakers, a display, smart glasses, etc.), and the like.

[0045] In some cases, head movement 115 may include rotation of user 110's head, translational movement of the head, etc. Device 105 may update view 130 of user 110 based on head movement 115. For example, before head movement 115, device 105 may display view 130-A for user 110. In some cases, after head movement 115, device 105 may display view 130-B to user 110. As view 130-A shifts to view 130-B, the augmented reality system (e.g., device 105) may render or update virtual objects and / or other portions of the virtual environment for display.

[0046] In some cases, extended reality system 100 may provide various types of virtual experiences, such as a three-dimensional (3D) collaborative virtual environment, to a group of users (eg, including user 110 ). Figure 2 2 is a diagram illustrating an example of a 3D collaborative virtual environment 200 in which various users interact with each other in a virtual session via virtual representations (or avatars) of the users in the virtual environment 200. The virtual representations include a virtual representation 202 of a first user, a virtual representation 204 of a second user, a virtual representation 206 of a third user, a virtual representation 208 of a fourth user, and a virtual representation 210 of a fifth user. Other contextual information for the virtual environment 200 is also shown, including a virtual calendar 212, a virtual webpage 214, and a virtual video conferencing interface 216. When interacting with the virtual representations of other users, the users can experience the virtual environment visually, aurally, haptically, or in other ways from each user's perspective. For example, the virtual environment 200 is shown from the perspective of the first user (represented by virtual representation 202).

[0047] As previously noted, it is important for XR systems to efficiently generate high-quality virtual representations (or avatars) with low latency. It may also be important for XR systems to render audio in an efficient manner to enhance the XR experience. For example, in Figure 2 In the example of a 3D collaborative virtual environment 200 of FIG. 1 , a first user's XR system (e.g., XR system 100) displays virtual representations 204-210 of other users participating in a virtual session. The users' virtual representations 204-210 and the background of the virtual environment 200 should be displayed in a realistic manner (e.g., as if the users were meeting in a real-world environment), such as by animating the heads, bodies, arms, and hands of the other users' virtual representations 204-210 as the users move in the real world. Audio captured by the other users' XR systems may need to be rendered spatially or may be rendered in mono for output to the first user's XR system. The latency of rendering and animating the virtual representations 204-210 should be minimized so that the first user's user experience is as if the user is interacting with the other users in the real-world environment.

[0048] Figure 3 is a diagram illustrating the architecture of an example system 300 according to some aspects of the present disclosure. The system 300 can be an XR system (e.g., a system that runs (or executes) XR applications and / or implements XR operations), a vehicle, a robotic system, or other type of system. The system 300 can perform tracking and positioning, mapping of an environment in the physical world (e.g., a scene), and / or positioning and rendering of virtual content on a display 309 (e.g., positioning and rendering of virtual content on a screen, visible plane / area, and / or other displays as part of an XR experience). For example, the system 300 can generate a map of the environment in the physical world (e.g., a three-dimensional (3D) map), track the pose (e.g., position and location) of the system 300 relative to the environment (e.g., relative to a 3D map of the environment), and / or determine positioning and / or anchor points in specific locations on the map of the environment. In one example, the system 300 may position and / or anchor virtual content in a specific location on a map of the environment, and may render the virtual content on a display 309 so that the virtual content appears to be located at a location in the environment that corresponds to the specific location on the map of the scene at which the virtual content is positioned and / or anchored. The display 309 may include a monitor, glass, screen, lens, projector, and / or other display mechanism. For example, in the context of an XR system, the display 309 may allow a user to see a real-world environment, and also allow XR content to be overlaid on, overlapped with, blended with, or otherwise displayed on the real-world environment.

[0049] In this illustrative example, system 300 may include one or more image sensors 302, accelerometers 304, gyroscopes 306, storage devices 307, computing components 310, pose engines 320, image processing engines 324, and rendering engines 326. Figure 3 The components 302-326 shown are non-limiting examples provided for purposes of illustration and explanation, and other examples may include those related to Figure 3 For example, in some cases, the system 300 may include one or more other sensors (e.g., one or more inertial measurement units (IMUs), radars, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, audio sensors, etc.), one or more display devices, one or more other processing engines, one or more other hardware components, and / or Figure 3 Although various components of system 300 (such as image sensor 302) may be referred to herein in the singular, it should be understood that system 300 may include multiples of any component discussed herein (e.g., multiple image sensors 302).

[0050] System 300 includes or communicates (wired or wirelessly) with an input device 308. Input device 308 can include any suitable input device, such as a touch screen, a pen or other pointer device, a keyboard, a mouse, buttons or keys, a microphone for receiving voice commands, a gesture input device for receiving gesture commands, a video game controller, a steering wheel, a joystick, a set of buttons, a trackball, a remote control, any other input device discussed herein, or any combination thereof. In some cases, image sensor 302 can capture images that can be processed for interpreting gesture commands.

[0051] In some implementations, one or more image sensors 302, accelerometers 304, gyroscopes 306, storage 307, computing component 310, pose engine 320, image processing engine 324, and rendering engine 326 can be part of the same computing device. For example, in some cases, one or more image sensors 302, accelerometers 304, gyroscopes 306, storage 307, computing component 310, pose engine 320, image processing engine 324, and rendering engine 326 can be integrated into a device or system such as an HMD, XR glasses (e.g., AR glasses), a vehicle or a system of a vehicle, a smartphone, a laptop, a tablet, a gaming system, and / or any other computing device. However, in some implementations, one or more image sensors 302, accelerometers 304, gyroscopes 306, storage 307, computing component 310, pose engine 320, image processing engine 324, and rendering engine 326 can be part of two or more separate computing devices. For example, in some cases, some of the components 302 - 326 may be part of or implemented by one computing device, and the remaining components may be part of or implemented by one or more other computing devices.

[0052] The storage device 307 can be any storage device for storing data. In addition, the storage device 307 can store data from any of the components of the system 300. For example, the storage device 307 can store data from the image sensor 302 (e.g., image or video data), data from the accelerometer 304 (e.g., measurements), data from the gyroscope 306 (e.g., measurements), data from the computing component 310 (e.g., processing parameters, preferences, virtual content, rendered content, scene maps, tracking and positioning data, object detection data, privacy data, XR application data, facial recognition data, occlusion data, etc.), data from the pose engine 320, data from the image processing engine 324, and / or data from the rendering engine 326 (e.g., output frames). In some examples, the storage device 307 may include a buffer for storing frames for processing by the computing component 310.

[0053] One or more computing components 310 may include a central processing unit (CPU) 312, a graphics processing unit (GPU) 314, a digital signal processor (DSP) 316, an image signal processor (ISP) 318, and / or other processors (e.g., a neural processing unit (NPU) implementing one or more trained neural networks). The computing component 310 may perform various operations such as image enhancement, computer vision, graphics rendering, tracking, localization, pose estimation, mapping, content anchoring, content rendering, image and / or video processing, sensor processing, recognition (e.g., text recognition, facial recognition, object recognition, feature recognition, tracking or pattern recognition, scene recognition, occlusion detection, etc.), trained machine learning operations, filtering, and / or any of the various operations described herein. In some examples, the computing component 310 may implement (e.g., control, operate, etc.) a pose engine 320, an image processing engine 324, and a rendering engine 326. In other examples, the computing component 310 may also implement one or more other processing engines.

[0054] Image sensor 302 can include any image and / or video sensor or capture device. In some examples, image sensor 302 can be part of a multi-camera assembly (such as a dual-camera assembly). Image sensor 302 can capture image and / or video content (e.g., raw image and / or video data), which can then be processed by computation component 310, pose engine 320, image processing engine 324, and / or rendering engine 326, as described herein.

[0055] In some examples, image sensor 302 can capture image data and can generate an image (also referred to as a frame) based on the image data and / or can provide the image data or frame to pose engine 320, image processing engine 324, and / or rendering engine 326 for processing. An image or frame can include a video frame from a video sequence or a still image. An image or frame can include an array of pixels representing a scene. For example, the image can be: a red-green-blue (RGB) image having red, green, and blue color components per pixel; a luminance, redness, blueness (YCbCr) image having a luminance component and two chrominance (color) components (redness and blueness) per pixel; or any other suitable type of color or monochrome image.

[0056] In some cases, image sensor 302 (and / or other cameras of system 300) can be configured to also capture depth information. For example, in some implementations, image sensor 302 (and / or other cameras) can include an RGB depth (RGB-D) camera. In some cases, system 300 can include one or more depth sensors (not shown) that are separate from image sensor 302 (and / or other cameras) and can capture depth information. For example, such depth sensors can obtain depth information independently of image sensor 302. In some examples, the depth sensor can be physically mounted in the same general location as image sensor 302 but can operate at a different frequency or frame rate than image sensor 302. In some examples, the depth sensor can take the form of a light source that can project a structured or textured light pattern (which can include one or more narrowband lights) onto one or more objects in a scene. Depth information can then be obtained by exploiting the geometric deformation of the projected pattern caused by the surface shape of the objects. In one example, depth information can be obtained from a stereo sensor, such as a combination of an infrared structured light projector and an infrared camera registered to a camera (e.g., an RGB camera).

[0057] System 300 may also include other sensors in its one or more sensors. The one or more sensors may include one or more accelerometers (e.g., accelerometer 304), one or more gyroscopes (e.g., gyroscope 306), and / or other sensors. The one or more sensors may provide velocity, orientation, and / or other positioning-related information to computing component 310. For example, accelerometer 304 may detect the acceleration of system 300 and may generate acceleration measurements based on the detected acceleration. In some cases, accelerometer 304 may provide one or more translation vectors (e.g., up / down, left / right, forward / backward), which may be used to determine the position or pose of system 300. Gyroscope 306 may detect and measure the orientation and angular velocity of system 300. For example, gyroscope 306 may be used to measure the pitch, roll, and yaw of system 300. In some cases, gyroscope 306 may provide one or more rotation vectors (e.g., pitch, yaw, roll). In some examples, the image sensor 302 and / or the pose engine 320 can use measurements obtained by the accelerometer 304 (e.g., one or more translation vectors) and / or measurements obtained by the gyroscope 306 (e.g., one or more rotation vectors) to calculate the pose of the system 300. As previously mentioned, in other examples, the system 300 can also include other sensors, such as an inertial measurement unit (IMU), a magnetometer, a gaze and / or eye tracking sensor, a machine vision sensor, an intelligent scene sensor, a voice recognition sensor, an impact sensor, a vibration sensor, a positioning sensor, a tilt sensor, and the like.

[0058] As described above, in some cases, the one or more sensors may include at least one IMU. An IMU is an electronic device that uses a combination of one or more accelerometers, one or more gyroscopes, and / or one or more magnetometers to measure specific forces, angular velocities, and / or orientations of system 300. In some examples, the one or more sensors may output measured information associated with the capture of images captured by image sensor 302 (and / or other cameras of system 300) and / or depth information obtained using one or more depth sensors of system 300.

[0059] The output of one or more sensors (e.g., accelerometer 304, gyroscope 306, one or more IMUs, and / or other sensors) can be used by pose engine 320 to determine the pose (also known as head pose) of system 300 and / or the pose of image sensor 302 (or other cameras of system 300). In some cases, the pose of system 300 and the pose of image sensor 302 (or other cameras) can be the same. The pose of image sensor 302 refers to the position and orientation of image sensor 302 relative to a reference frame (e.g., with respect to an object). In some implementations, the camera pose can be determined for six degrees of freedom (6DoF), which refers to three translational components (e.g., which can be given by X (horizontal), Y (vertical), and Z (depth) coordinates relative to a reference frame (such as the image plane)) and three angular components (e.g., roll, pitch, and yaw relative to the same reference frame). In some implementations, the camera pose can be determined for 3 degrees of freedom (3DoF), which refers to three angular components (e.g., roll, pitch, and yaw).

[0060] In some cases, a device tracker (not shown) can use measurements from one or more sensors and image data from image sensor 302 to track the pose (e.g., 6DoF pose) of system 300. For example, the device tracker can fuse visual data from the image data (e.g., using a visual tracking solution) with inertial data from the measurements to determine the position and motion of system 300 relative to the physical world (e.g., a scene) and a map of the physical world. As described below, in some examples, when tracking the pose of system 300, the device tracker can generate a three-dimensional (3D) map of the scene (e.g., the real world) and / or generate updates to the 3D map of the scene. 3D map updates can include, for example, but not limited to, new or updated features and / or features or landmark points associated with the scene and / or the 3D map of the scene, positioning updates that identify or update the position of system 300 within the scene and the 3D map of the scene, and the like. The 3D map can provide a digital representation of the scene in the real / physical world. In some examples, the 3D map can anchor location-based objects and / or content to real-world coordinates and / or objects. The system 300 may use a mapped scene (eg, a scene in the physical world represented by a 3D map and / or associated with the 3D map) to merge physical and virtual worlds and / or merge virtual content or objects with the physical environment.

[0061] In some aspects, the computing component 310 can use a visual tracking solution to determine and / or track the pose (also referred to as the camera pose) of the image sensor 302 and / or the system 300 as a whole based on images captured by the image sensor 302 (and / or other cameras of the system 300). For example, in some examples, the computing component 310 can perform tracking using computer vision-based tracking, model-based tracking, and / or simultaneous localization and mapping (SLAM) techniques. For example, the computing component 310 can perform SLAM or can be integrated with a SLAM system (in Figure 3 not shown) such as Figure 4SLAM system 400 communicates (wired or wirelessly). SLAM refers to a class of technologies that creates a map of an environment (e.g., a map of the environment modeled by system 300) while tracking the pose of a camera (e.g., image sensor 302) and / or system 300 relative to the map. The map can be referred to as a SLAM map and can be three-dimensional (3D). SLAM technology can be performed using color or grayscale image data captured by image sensor 302 (and / or other cameras of system 300) and can be used to generate an estimate of 6DoF pose measurements of image sensor 302 and / or system 300. Such SLAM technology configured to perform 6DoF tracking can be referred to as 6DoF SLAM. In some cases, the output of one or more sensors (e.g., accelerometer 304, gyroscope 306, one or more IMUs and / or other sensors) can be used to estimate, correct and / or otherwise adjust the estimated pose.

[0062] In some cases, 6DoF SLAM (e.g., 6DoF tracking) can associate features (e.g., key points) observed from certain input images from the image sensor 302 (and / or other cameras or sensors) to a SLAM map. For example, 6DoF SLAM can use feature point associations from the input image (or other sensor data, such as a radar sensor, a LIDAR sensor, etc.) to determine the pose (positioning and orientation) of the image sensor 302 and / or the system 300 for the input image. 6DoF map building can also be performed to update the SLAM map. In some cases, a SLAM map maintained using 6DoF SLAM can include 3D feature points (e.g., key points) triangulated from two or more images. For example, key frames can be selected from the input image or video stream to represent the observed scene. For each key frame, a corresponding 6DoF camera pose associated with the image can be determined. The pose of the image sensor 302 and / or system 300 may be determined by projecting features (eg, feature points or keypoints) from a 3D SLAM map into an image or video frame and updating the camera pose based on the verified 2D-3D correspondences.

[0063] In one illustrative example, the computing component 310 may extract feature points (e.g., key points) from certain input images (e.g., each input image, a subset of the input images, etc.) or from each key frame. As used herein, a feature point (also referred to as a registration point) is a unique or identifiable portion of an image, such as a portion of a hand, the edge of a table, and other examples. Features extracted from a captured image may represent different feature points along three-dimensional space (e.g., coordinates on the X, Y, and Z axes), and each feature point may have an associated feature location. Feature points in a key frame may match (be identical to or correspond to) or fail to match feature points of a previously captured input image or key frame. Feature detection may be used to detect feature points. Feature detection may include image processing operations for examining one or more pixels of an image to determine whether a feature exists at a particular pixel. Feature detection may be used to process the entire captured image or portions of an image. For each image or key frame, once a feature has been detected, a local image patch surrounding the feature may be extracted. Features may be extracted using any suitable technique, such as Scale Invariant Feature Transform (SIFT) (which localizes features and generates their descriptions), Learned Invariant Feature Transform (LIFT), Speeded Up Robust Features (SURF), Gradient Location Orientation Histogram (GLOH), Orientation Rapid Rotation Brief (ORB), Binary Robust Invariant Scalable Keypoints (BRISK), Fast Retinal Keypoints (FREAK), KAZE, Accelerated KAZE (AKAZE), Normalized Cross Correlation (NCC), descriptor matching, another suitable technique, or a combination thereof.

[0064] In some cases, the system 300 may also track the user's hands and / or fingers to allow the user to interact with and / or control virtual content in the virtual environment. For example, the system 300 may track the position and / or movement of the user's hands and / or fingertips to identify or interpret the user's interaction with the virtual environment. User interaction may include, for example, but is not limited to, moving virtual content items, resizing virtual content items, selecting input interface elements in a virtual user interface (e.g., a virtual representation of a mobile phone, a virtual keyboard, and / or other virtual interfaces), providing input through the virtual user interface, etc.

[0065] Figure 4 is a block diagram illustrating the architecture of a simultaneous localization and mapping (SLAM) system 400. In some examples, the SLAM system 400 may be Figure 3In some examples, the SLAM system 400 may be, may include, or may be a part of an XR device, an autonomous vehicle, a vehicle, a computing system of a vehicle, a wireless communication device, a mobile device or handset (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device (e.g., a web-connected watch), a personal computer, a laptop computer, a server computer, a portable video game console, a portable media player, a camera device, a manned or unmanned ground vehicle, a manned or unmanned air vehicle, a manned or unmanned water vehicle, a manned or unmanned underwater vehicle, a manned or unmanned vehicle, a robot, another device, or any combination thereof, or may be a part thereof.

[0066] Figure 4 The SLAM system 400 may include or be coupled to each of the one or more sensors 405. The one or more sensors 405 may include one or more cameras 410. Each of the one or more cameras 410 may include an image capture device, an image processing device (e.g., Figure 14 The processor 1410 of the image capture and processing system, another type of camera, or a combination thereof. Each of the one or more cameras 410 can respond to light from a specific spectrum. The spectrum can be a subset of the electromagnetic (EM) spectrum. For example, each of the one or more cameras 410 can be a visible light (VL) camera that responds to the VL spectrum, an infrared (IR) camera that responds to the IR spectrum, an ultraviolet (UV) camera that responds to the UV spectrum, a camera that responds to another spectrum of light from another part of the EM spectrum, or some combination thereof.

[0067] The one or more sensors 405 may include one or more other types of sensors in addition to the camera 410, such as one or more of each of the following: an accelerometer, a gyroscope, a magnetometer, an inertial measurement unit (IMU), an altimeter, a barometer, a thermometer, a radio detection and ranging (RADAR) sensor, a light detection and ranging (LIDAR) sensor, a sound navigation and ranging (SONAR) sensor, a sound detection and ranging (SODAR) sensor, a global navigation satellite system (GNSS) receiver, a global positioning system (GPS) receiver, a BeiDou navigation satellite system (BDS) receiver, a Galileo receiver, a GLONASS satellite navigation system (GLONASS) receiver, an Indian constellation navigation (NavIC) receiver, a Quasi-Zenith Satellite System (QZSS) receiver, a Wi-Fi positioning system (WPS) receiver, a cellular network positioning system receiver, Beacon positioning receivers, short-range wireless beacon positioning receivers, personal area network (PAN) positioning receivers, wide area network (WAN) positioning receivers, wireless local area network (WLAN) positioning receivers, other types of positioning receivers, other types of sensors discussed herein, or combinations thereof. In some examples, one or more sensors 405 may include Figure 3 Any combination of sensors of system 300.

[0068] Figure 4 The SLAM system 400 may include a visual-inertial odometry (VIO) tracker 415. The term "visual-inertial odometry" may also be referred to herein as visual odometry. The VIO tracker 415 may receive sensor data 465 from the one or more sensors 405. For example, the sensor data 465 may include one or more images captured by the one or more cameras 410. The sensor data 465 may include other types of sensor data from the one or more sensors 405, such as data from any of the types of sensors 405 listed herein. For example, the sensor data 465 may include IMU data from one or more inertial measurement units (IMUs) in the one or more sensors 405.

[0069] Upon receiving sensor data 465 from one or more sensors 405, the VIO tracker 415 may perform feature detection, extraction, and / or tracking using the feature tracking engine 420 of the VIO tracker 415. For example, where the sensor data 465 includes one or more images captured by one or more cameras 410 of the SLAM system 400, the VIO tracker 415 may identify, detect, and / or extract features in each image. Features may include visually distinct points in an image, such as portions of an image that depict edges and / or corners. The VIO tracker 415 may periodically and / or continuously receive sensor data 465 from the one or more sensors 405, for example, by continuously receiving more images from the one or more cameras 410 as the one or more cameras 410 capture a video, where the images are video frames of the video. The VIO tracker 415 may generate descriptors for the features. The feature descriptors may be generated, at least in part, by generating a description of the feature, as depicted in a local image patch extracted around the feature. In some examples, the feature descriptors may describe the feature as a collection of one or more feature vectors.

[0070] The VIO tracker 415, in some cases, together with the mapping engine 430 and / or the relocalization engine 455, can associate multiple features with a map of the environment based on such feature descriptors. The feature tracking engine 420 of the VIO tracker 415 can perform feature tracking by identifying features in each image that the VIO tracker 415 has previously identified in one or more previous images (in some cases, based on identifying features with matching feature descriptors in different images). The feature tracking engine 420 can track changes in one or more locations at which a feature is depicted in each of the different images. For example, the feature extraction engine can detect a particular corner of a depicted room on the left side of a first image captured by a first camera in the cameras 410. The feature extraction engine can detect the same depicted feature (e.g., the same particular corner of the same room) on the right side of a second image captured by the first camera. The feature tracking engine 420 can identify that the feature detected in the first image and the second image are two depictions of the same feature (e.g., the same particular corner of the same room), and that the feature appears in two different locations in the two images. The VIO tracker 415 may determine that the first camera has moved based on the appearance of the same feature on the left side of the first image and on the right side of the second image, for example, where the feature depicts a static portion of the environment (e.g., a particular corner of a room).

[0071] The VIO tracker 415 may include a sensor integration engine 425. The sensor integration engine 425 may use sensor data from other types of sensors 405 (in addition to the camera 410) to determine information that can be used by the feature tracking engine 420 when performing feature tracking. For example, the sensor integration engine 425 may receive IMU data from an IMU in one or more sensors 405 (e.g., which may be included as part of the sensor data 465). The sensor integration engine 425 may determine, based on the IMU data in the sensor data 465, that the SLAM system 400 has rotated 15 degrees in a clockwise direction from the time the first camera in the cameras 410 acquired or captured the first image to the time the second image was acquired or captured. Based on this determination, the sensor integration engine 425 may identify that a feature depicted at a first location in the first image is expected to appear at a second location in the second image, and that the second location is expected to be a predetermined distance (e.g., a predetermined number of pixels, inches, centimeters, millimeters, or another distance metric) to the left of the first location. The feature tracking engine 420 may take this expectation into account when tracking the feature between the first image and the second image.

[0072] Based on the feature tracking performed by the feature tracking engine 420 and / or the sensor integration performed by the sensor integration engine 425, the VIO tracker 415 may determine a 3D feature location 472 for a particular feature. The 3D feature location 472 may include one or more 3D feature locations and may also be referred to as a 3D feature point. The 3D feature location 472 may be a set of coordinates along three different axes perpendicular to each other, such as an X coordinate along an X axis (e.g., in a horizontal direction), a Y coordinate along a Y axis perpendicular to the X axis (e.g., in a vertical direction), and a Z coordinate along a Z axis perpendicular to both the X axis and the Y axis (e.g., in a depth direction). In some aspects, the VIO tracker 415 may also determine one or more keyframes 470 (hereinafter referred to as keyframes 470) corresponding to the particular feature. A keyframe (from one or more keyframes 470) corresponding to the particular feature may be an image in which the particular feature is clearly depicted. In some examples, a keyframe (from one or more keyframes 470) corresponding to the particular feature may be an image in which the particular feature is clearly depicted. In some examples, a keyframe corresponding to a particular feature can be an image that, when considered by the feature tracking engine 420 and / or the sensor integration engine 425 for determining the 3D feature location 472, reduces the uncertainty of the 3D feature location 472 of the particular feature. In some examples, a keyframe corresponding to a particular feature also includes data regarding the pose 485 of the SLAM system 400 and / or the camera 410 during the capture of the keyframe. In some examples, the VIO tracker 415 can transmit the 3D feature location 472 and / or keyframes 470 corresponding to one or more features to the mapping engine 430. In some examples, the VIO tracker 415 can receive a map slice 475 from the mapping engine 430. The VIO tracker 415 can perform feature extraction on the information within the map slice 475 for feature tracking using the feature tracking engine 420.

[0073] Based on the feature tracking performed by the feature tracking engine 420 and / or the sensor integration performed by the sensor integration engine 425, the VIO tracker 415 can determine the pose 485 of the SLAM system 400 and / or the camera 410 during the capture of each of the images in the sensor data 465. The pose 485 may include the position of the SLAM system 400 and / or the camera 410 in 3D space, such as a set of coordinates along three different axes perpendicular to each other (e.g., X coordinate, Y coordinate, and Z coordinate). The pose 485 may include the orientation of the SLAM system 400 and / or the camera 410 in 3D space, such as pitch, roll, yaw, or some combination thereof. In some examples, the VIO tracker 415 may transmit the pose 485 to the relocalization engine 455. In some examples, the VIO tracker 415 may receive the pose 485 from the relocalization engine 455.

[0074] The SLAM system 400 also includes a mapping engine 430. The mapping engine 430 can generate a 3D map of the environment based on the 3D feature locations 472 and / or keyframes 470 received from the VIO tracker 415. The mapping engine 430 can include a map densification engine 435, a keyframe remover 440, a bundle adjuster 445, and / or a loop closure detector 450. The map densification engine 435 can perform map densification, in some examples, increasing the number and / or density of 3D coordinates describing the map geometry. The keyframe remover 440 can remove keyframes and / or add keyframes in some cases. In some examples, the keyframe remover 440 can remove keyframes 470 corresponding to areas of the map to be updated and / or areas with low corresponding confidence values. In some examples, the bundle adjuster 445 can refine the 3D coordinates describing the scene geometry, parameters of relative motion, and / or optical properties of the image sensor used to generate the frame according to an optimality criterion involving corresponding image projections of all points. The loop closure detector 450 can identify when the SLAM system 400 has returned to a previously mapped area and can use such information to update the map slice and / or reduce uncertainty in certain 3D feature points or other points in the map geometry.

[0075] The mapping engine 430 may output map slices 475 to the VIO tracker 415. The map slices 475 may represent 3D portions or subsets of the map. The map slices 475 may include map slices 475 that represent new, previously unmapped areas of the map. The map slices 475 may include map slices 475 that represent updates (or modifications or revisions) to previously mapped areas of the map. The mapping engine 430 may output map information 480 to the relocalization engine 455. The map information 480 may include at least a portion of the map generated by the mapping engine 430. The map information 480 may include one or more 3D points that constitute the geometry of the map, such as one or more 3D feature locations 472. The map information 480 may include one or more keyframes 470 corresponding to certain features and certain 3D feature locations 472.

[0076] The SLAM system 400 also includes a relocalization engine 455. The relocalization engine 455 can perform relocalization, for example, when the VIO tracker 415 fails to identify more than a threshold number of features in an image and / or when the VIO tracker 415 loses track of the pose 485 of the SLAM system 400 within the map generated by the mapping engine 430. The relocalization engine 455 can perform relocalization by performing extraction and matching using the extraction and matching engine 460. For example, the extraction and matching engine 460 can extract features from an image captured by the camera 410 of the SLAM system 400 when the SLAM system 400 is in the current pose 485, and can match the extracted features with features depicted in different keyframes 470, identified by the 3D feature localization 472, and / or identified in the map information 480. By matching these extracted features with the features previously identified, the pose 485 of the repositioning engine 455 that can identify the SLAM system 400 is the pose 485 at which the feature previously identified is visible to the camera 410 of the SLAM system 400, and is therefore similar to one or more previous poses 485 at which the feature previously identified is visible to the camera 410. In some cases, the repositioning engine 455 can be performed repositioning based on the distance between the camera positioning at which the feature was initially captured based on wide baseline mapping or current camera positioning. The repositioning engine 455 can receive information (e.g., information about one or more recent poses of the SLAM system 400 and / or camera 410) of the pose 485 from the VIO tracker 415, and the repositioning engine 455 can determine its repositioning based on this information. Once the repositioning engine 455 repositions the SLAM system 400 and / or camera 410 and therefore determines pose 485, the repositioning engine 455 can output pose 485 to the VIO tracker 415.

[0077] As previously noted, a user may wish to augment the physical three-dimensional (3D) space around the user with synthetic content. A user may want to share an augmented 3D scene (e.g., an XR room) with other users, such as a classroom (e.g., Figure 11 and Figure 12 as shown), restaurants (e.g. Figure 8 and Figure 10 ), businesses, factories, homes (e.g. Figure 9 These XR rooms can be multifunctional by providing users with interactive information, entertainment, and / or synthetic light experiences. In the case of such XR rooms, the unified framework described herein can help users navigate through the XR room.

[0078] Systems and techniques provide a proximity-based protocol for enabling multi-user XR experiences. In particular, the protocol provides a unified framework for users to host and join multi-user XR rooms. In one or more aspects, a static host device (e.g., Figure 7 The static host device 730, which may be a wireless router, may communicate with the user's device (such as the user's AR / VR HMD (e.g., Figure 7 The user HMD 720 and the user HMD 740)) communicate.

[0079] During operation, a host device (e.g. Figure 7 A user device (such as an HMD or other XR device (e.g., Figure 7 A user HMD 720) can connect to the host device as an administrator. The user device can scan the entire 3D scene of its environment (e.g., Figure 7 The user device may capture images of an XR room 710, which may be a physical room. The images captured by the user device may be sent to or otherwise shared with the host device. The images may be processed within the host device itself and / or offline (e.g., in a cloud such as a cloud server) to generate a 3D map of the scene (e.g., containing features such as points of interest, and the geometry of the scene). The user device, when operating as an administrator, may place synthetic objects (e.g., Figure 7 Synthetic content 750, in the form of lamps and / or lights) is added to the scene from the obtained scene geometry.

[0080] One or more other user devices (such as other HMDs or other XR devices (e.g., user HMD 740)) located near the host device (e.g., within communication range) may receive the broadcast message sent from the host device. The message may prompt the one or more other user devices to join the XR room (e.g., XR room 710). The host device may authenticate the one or more other user devices before the user can join the XR room. The user devices may communicate image features of their physical location that they have obtained (e.g., which may have different views and perspectives of the physical location) to the host device. For example, in aspects where the user devices are HMDs, one or more of the user devices (e.g., HMDs) may provide (e.g., transmit) one or more corresponding images from the HMD's forward-facing camera to the host device. The host device may match the image features from the one or more other user devices with scene features in the host device, and the host device may then locate the one or more other user devices relative to the XR room. For example, to locate the user devices, the host device may detect features (e.g., points of interest) in the image and match the features with a database of scene features in the host device. Positioning of the one or more other user devices may be performed periodically by the host device to avoid any drift. The host device may communicate the synthetic content geometry to the one or more other user devices. Each of the one or more other user devices may render the synthetic content to the HMD's visual perspective camera view using the head / eye pose of the user associated with the HMD.

[0081] In one or more aspects, a host device (e.g., Figure 7 The host device 730) can be a static device, similar to a wireless router. The host device can be assigned to a specific 3D scene (e.g., a specific physical room). Figure 7 The user devices (e.g., user HMD 720) may differ in capabilities, and the host device may have the option to select from a pool of different user HMD devices based on needs.

[0082] The host device can establish a peer-to-peer connection (e.g., via radio frequency communication) with multiple user devices (e.g., user HMDs, such as AR / VR user HMDs) located within proximity of the host device (e.g., within communication range). The host device can maintain a connection with (e.g., maintain wireless communication with) the user HMD within a specified range (e.g., within twenty meters). In one or more examples, additional devices (e.g., similar to wireless router extenders) can be employed to operate as extenders to extend the range (e.g., communication range) of the host device. These additional devices (e.g., extenders) can increase the spatial range of the host device to accommodate XR rooms of any size.

[0083] The host device may assign different permissions to users (e.g., user HMDs) such as hosts, administrators, guests, viewers, etc. Users (e.g., user HMDs) assigned host permissions may have the ability to add and / or modify composite content within a scene (e.g., within an XR room).

[0084] A user device (e.g., a user HMD) may join or connect (e.g., via a wireless communication protocol) to a host device to operate as an administrator. When the user HMD operates as an administrator (e.g., in administrator mode), the user HMD may provide an image of a scene (e.g., a physical room) to the host device to create and / or update features (e.g., points of interest) and geometry of the scene. When operating as an administrator, the user device may add and / or modify synthetic content to the physical scene.

[0085] The host device may store (e.g., in a host database such as a feature database) the geometry of a 3D scene, which may include both physical content and synthetic content. When other user devices (e.g., HMDs of other users) join a 3D scene (e.g., an XR room), the synthetic content of the scene may be shared with the other user devices.

[0086] The host device may periodically receive camera feeds from other user devices to position the user relative to the 3D scene (e.g., XR room) in order to correct for any user positioning drift. In a more automated approach, the host device may also have a protocol for updating the 3D scene (e.g., XR room) using selected camera feeds from all user devices in one or more other user devices.

[0087] In one or more aspects, the user device may be an AR / VR HMD or AR glasses (or lenses) that may be connected to a single host device at a time. The user device (e.g., user HMD or user lenses) may estimate the eye positioning and head pose of its associated user at all times during operation. After receiving the synthetic content from the host device, the user device may render the synthetic content to the user via the user device's perspective view. The user device may need to know both natural and synthetic light structures in order to render realistic synthetic content for a physical scene (e.g., in an XR room).

[0088] As mentioned previously, currently, XR rooms typically require users of a specific app to join the XR room. XR rooms are typically hosted on the cloud (e.g., a cloud server), and users need to establish a connection to the cloud (e.g., a cloud server) via an app to experience the XR room. The disclosed systems and technologies employ a local host device (e.g., a wireless router) that can communicate with a user's device (e.g., the user's AR / VR HMD and / or the user's AR lenses) using a wireless protocol that allows for smooth transitions between the user and the XR room based on the proximity of the user's device to the host device.

[0089] Using a local host device instead of the cloud to host an XR room has many advantages. One advantage of using a local host device instead of the cloud to host an XR room is that using a local host device can provide better security than using the cloud. For example, a hacker could create multiple duplicate identifications (IDs) in an attempt to log into the cloud to access the XR room. A hacker generating many IDs could overload the system and crash the XR room. Systems and techniques overcome this problem in several ways by using a local host device. For example, systems and techniques can use multimodal data to robustly authenticate a user (and / or the user's device) to enable the user to join the XR room.

[0090] In one or more aspects of the systems and techniques, in order to be authenticated, a user (e.g., a user's device) needs to be physically within proximity of a host device (e.g., within twenty meters) and should be able to receive a broadcast message sent from the host device (e.g., within the communication range of the host device) to join the XR room. Since the communication between the host device and the user device (e.g., the user's HMD) is local, existing wireless protocols can be used to ensure data security.

[0091] In one or more examples, a user (e.g., a user's device) may be continuously authenticated by a host device using camera images from the user's device. The host device may match the camera images from the user's device with a 3D scene (e.g., an XR room) to authenticate the user (e.g., the user's device). In one or more examples of further authentication of the user (e.g., the user's device), the host device may direct the user to specifically change their viewing direction (e.g., move their head positioning or eye gaze) and use this new viewing direction to obtain additional camera images. The host device may then match these additional camera images with the 3D scene to further authenticate the user.

[0092] Another advantage of using a local host device instead of the cloud to host the XR room is that using a local host device allows for reduced bandwidth usage. Since the user device only needs to maintain a connection with the host device (e.g., not with the cloud), the communication data is not sent over the wired data fiber, which may require bandwidth usage. In the case of a scene update (e.g., for the XR room), the host device can broadcast the scene change to all user devices using only a single broadcast message. Broadcasting a single message requires less RF bandwidth than sending multiple messages. It is possible for the community to build a protocol for broadcasting a single message (e.g., a scene update for an XR room) to reduce bandwidth usage and ensure smooth communication.

[0093] Figure 5 and Figure 6 An example process 500, 600 for implementing a multi-user XR experience is shown. Specifically, Figure 5 is a flow chart illustrating an example of a process 500 for 3D reconstruction and generation of a feature database (e.g., a host database containing features). Figure 5 At block 510, a user device (e.g., a user AR / VR HMD such as Figure 7 The user HMD 7720, or the user AR lens) can be connected to or joined to a host device (e.g., a wireless router such as Figure 7 The user device may scan the entire 3D scene of its physical location (e.g., its physical environment) to obtain a camera image of the 3D scene. At block 530, the user device may share (e.g., send to the host device) the obtained camera image of the 3D scene and pose information of the user device (e.g., a 6DoF pose of the user device, such as a 6DoF head pose) with the host device.

[0094] At block 540, the host device may perform depth detection or generation (e.g., using a machine learning-based algorithm, a computer vision-based algorithm, or other depth generation algorithm) on the camera images of the 3D scene received from the user device to generate a depth map for each image. At block 540, the host device may then perform 3D scene reconstruction using the 6Dof depth map to reconstruct the 3D scene.

[0095] At block 560, the host device may perform feature detection and tracking (e.g., for multiple views) using the camera image of the 3D scene and the 6Dof from the user device to obtain 3D features and descriptors. At block 570, the host device may store (e.g., in at least one host database such as a feature database) the 3D features and descriptors across the 3D scene.

[0096] Figure 6is an example for using a head mounted device (HMD) (e.g., such as Figure 7 Flowchart of an example of a process 600 for authenticating and locating a user using a user HMD 720 of an XR room. At block 610, a user device (e.g., an HMD) located within proximity of an XR room (e.g., within twenty meters) may receive a request from a host device (e.g., Figure 7 The host device 730 of the user device receives the broadcast message (e.g., a wireless radio frequency message). At block 620, the user of the user device (and / or the user device itself) may optionally be authenticated using a password. For example, the user device may send a password associated with the user to the host device for the host device to authenticate the user. At block 630, the user device may share (e.g., send to the host device) a camera image of the 3D scene obtained by the user device (e.g., an HMD) with the host device.

[0097] At block 640, the host device may perform feature detection using a camera image of the 3D scene from the user device. At block 660, the host device may have previously stored (e.g., in at least one host database, such as a feature database) 3D features and descriptors across the 3D scene. At block 650, the host device may perform feature matching and bundle assignment using features detected from the camera image of the 3D scene from the user device and using 3D features and descriptors across the 3D scene from at least one host database. At block 670, the host device may determine whether the positioning of the user device (e.g., the user HMD) was successful. If the host device determines that the positioning of the user device was successful, the host device may use the 6Dof pose of the user relative to the scan. However, if the host device determines that the positioning of the user device was unsuccessful, at block 680, the host device may suggest to the user that the pose of the user device (e.g., the user HMD) be changed. Process 600 then proceeds back to block 630.

[0098] Figure 7 、 Figure 8 、 Figure 9 、 Figure 10 、 Figure 11 and Figure 12 Different examples and applications (e.g., different use cases) of systems and techniques for enabling multi-user XR experiences are shown. Specifically, Figure 7 is a diagram illustrating an example 700 where users located within proximity of an XR room 710 are able to join the XR room 710. Figure 7, an XR room 710 is shown as including a host device 730 (e.g., a wireless router), a host user device (e.g., a host user HMD 720), and a user device (e.g., a user HMD 740). A user operating as an administrator (who may be referred to as a host user) is shown wearing the host user device (e.g., the host user HMD 720). A user who wants to join the XR room 710 is shown wearing the user device (e.g., the user HMD 740).

[0099] In one or more examples, during operation of the protocol 712 for a host user to host an XR room, at block 722, a host device 730 (e.g., a wireless router) can be assigned to the XR room 710. At block 732, the host user device (e.g., the host user HMD 720) can join the host device 730 as an administrator (or host) and can provide the host device 730 with a camera scan of the entire scene obtained by the host user device (e.g., the host user HMD 720). At block 742, the host device 730 can process and / or process the camera scan offline (e.g., in a cloud such as a cloud server) to generate a 3D scene graph. At block 752, the host user device (e.g., the host user HMD 720) can insert synthetic content (e.g., the host user can add synthetic content 760, such as synthetic objects, such as synthetic lights 750 and / or synthetic lighting) into the 3D scene. At block 762 , the host device 730 may broadcast an XR room message (eg, which may include an invitation to join the XR room) to all users (eg, including the user HMD 740 ) within proximity to the host device 730 .

[0100] In one or more examples, during operation of the protocol 715 for a user to join an XR room, at block 725, a user (e.g., a user device, such as a user HMD 740) located within proximity 780 of the XR room 710 may receive a broadcast message sent from all host devices (e.g., host device 730) located within the proximity. At block 735, the user HMD 740 may be prompted 770 (e.g., via a broadcast message) to join the XR room 710, and the host device 730 may authenticate the user HMD. At block 745, the host device 730 may position the user HMD 740 relative to the 3D scene. At block 755, the host device 730 may share synthetic content (e.g., synthetic light 750) with the user HMD 740. At block 765, the user HMD 740 may render the synthetic content (e.g., synthetic light 750) based on the user's head pose and / or eye gaze.

[0101] Figure 8 is a diagram illustrating an example 800 of a business (e.g., a restaurant) having a corresponding XR room. Figure 8In Figure 1, three physical restaurants are shown. Figure 8 As shown, at block 860, in the future, each restaurant can create an XR room (e.g., XR Room 1 810a, XR Room 2 810b, XR Room 3 810c) to enhance the customer experience (e.g., by generating synthetic content for the customer, such as a synthetic interactive menu). At block 870, whenever a user is within proximity of an XR room, the user can be prompted with the option to join the XR room. At block 880, the user can seamlessly navigate through multiple XR rooms using a simple proximity-based algorithm.

[0102] Each restaurant in Figure 8 810a, XR Room 2 810b, XR Room 3 810c). Each XR Room is associated with a host device. Two users are also shown, each associated with a user device (e.g., user HMD 820a, 820b). User HMD 820a is within proximity (e.g., within twenty meters) of host device 830a and host device 830b. Therefore, XR Room 1 810a and XR Room 2 810b are available 840 for user HMD 820a to join. User HMD 820b is within proximity (e.g., within twenty meters) of host device 830c. Therefore, XR Room 3 810c is available 850 for user HMD 820b to join.

[0103] Figure 9 is a diagram illustrating an example 900 of users collaboratively hosting an XR room. Figure 9 , a shared living room 910 of a house is shown. Figure 9 As shown, at block 950, multiple users can collaboratively build an XR room in a shared physical space (e.g., such as a shared living room 910). At block 960, multiple users can join the XR room as hosts (e.g., operating as administrators or hosts) and add composite content to the 3D scene.

[0104] exist Figure 9 , a host device 930 and two users are shown as being located within a shared living room 910. One user may operate as a host user and be associated with a user device (e.g., host user HMD 920a). The other user may operate as a co-host user and be associated with a user device (e.g., co-host user HMD 920b). Both users within the shared living space 910 may operate as hosts (or administrators) 940 and, therefore, may simultaneously create and / or modify the composite content of the XR room of the shared living space 910.

[0105] Figure 10 is a diagram illustrating an example 1000 of a user interacting with composite content (e.g., composite content 1040a, 1040b, 1040c) within an XR room 1010. Figure 10 In FIG. 1 , an XR room 1010 of a business (e.g., a restaurant) is shown as including a host device 1030. Figure 10 As shown, at block 1060, an XR room (e.g., XR room 1010) may host interactive synthetic content (e.g., a synthetic menu) where users may interact with and / or modify the synthetic content. At block 1070, for example, a restaurant may have a visually interactive menu, and a waiting room (e.g., at a doctor's office) may have engaging synthetic content to entertain users while they wait for their appointment.

[0106] exist Figure 10 , users associated with user devices (e.g., wearing user AR lenses 1020) are shown in an XR room 1010 of a business (e.g., a restaurant). All users wearing user devices (e.g., user AR lenses, such as user AR lenses 1020) and joining the XR room 1010 can browse and interact with the restaurant's composite menu 1050. For example, the composite menu may include composite content 1040a, 1040b, 1040c, such as Figure 10 shown.

[0107] Figure 11 is a diagram illustrating an example 1100 of an XR classroom. Figure 11 , an XR classroom 1110 of a school is shown to include a host device 1130, a user (e.g., a teacher) associated with a user device (e.g., a user HMD 1120), and users (e.g., students) associated with user devices (e.g., user HMDs 1140a, 1140b, 1140c, 1140d, 1140e, 1140f). Figure 11 As shown, in the future, a classroom and / or presentation space may have an XR room (e.g., XR classroom 1110) at block 1080. At block 1190, a presenter (e.g., a teacher) may modify the composite content (e.g., composite content in the form of a heart 1150a and composite content in the form of a skeleton 1150b), while attendees (e.g., students) may experience the composite content.

[0108] exist Figure 11In the embodiment, a host user (e.g., a teacher) can add 1160 composite content (e.g., composite content 1170 including composite content 1150a and composite content 1150b) to the XR classroom 1110 by using their associated user HMD 1120. Students in the XR classroom 1110 can view and interact with the composite content 1170 (e.g., including composite content 1150a and composite content 1150b) by using their associated user HMDs 1140a, 1140b, 1140c, 1140d, 1140e, 1140f.

[0109] Figure 12 12 is a diagram illustrating an example 1200 of XR sub-rooms (eg, XR sub-room 1 1215a, XR sub-room 2 1215b, and XR sub-room 3 1215c) within an XR classroom 1210 (eg, a main AR room). Figure 11 In FIG, an XR classroom 1210 of a school is shown to include desks 1250a, 1250b, 1250c, a host device 1230, a user (e.g., a teacher) associated with a user device (e.g., a user HMD 1220), and users (e.g., students) associated with user devices (e.g., user HMDs 1240a, 1240b, 1240c). Each desk 1250a, 1250b, 1250c is associated with a corresponding XR sub-room.

[0110] like Figure 11 As shown, the possibility of having XR sub-rooms (e.g., XR sub-room 1 1215a, XR sub-room 2 1215b, and XR sub-room 3 1215c) within an XR room (e.g., XR classroom 1210) provides options for personalizing the XR room experience for users (e.g., personalizing the learning interaction with the composite content for students). Within the same main AR room (e.g., XR classroom 1210), different users can have different experiences based on the different XR sub-rooms to which they are assigned. Users can be assigned to different XR sub-rooms based on their attributes (e.g., their age, vision (color blindness), etc.). In one or more examples, the main AR room can be a classroom, a personal home, etc.

[0111] In another illustrative example of an application of the systems and techniques described herein, an interactive XR room is a grocery store, and users within proximity of the grocery store have XR devices that allow the user to join the XR room. After joining, if the user wants to know the physical location of a product, the user can utilize the XR device (e.g., an HMD) to interact with the XR room. When the user joins the XR room using their XR device, the XR room can render virtual directions to the product specific to the user only.

[0112] Figure 13is a flow chart illustrating an example of a process 1300 for implementing a multi-user XR experience. The process 1300 may be performed by a user device (or computing device apparatus) of a user or by a component or system (e.g., a chipset) of a computing device. The user device (or its component or system) may include or may be Figure 3 System 300, Figure 14 In some aspects, the user device is an XR device (e.g., AR / VR HMD, AR glasses, or AR lenses, etc.). The operations of process 1300 can be implemented as a processor (e.g., Figure 14 In addition, the communication may be performed, for example, via one or more antennas and / or one or more transceivers such as wireless transceivers (e.g., using Figure 14 The communication interface 1440 of the computing system 1400 of the embodiment of the present invention is used to enable the sending and receiving of signals by the first network entity in the process 1300.

[0113] At block 1310, a user device (or a component thereof) may receive a message from a host device including a prompt to join an XR room hosted by the host device. The user device may be within communication range of the host device. In some aspects, the host device is a wireless router (e.g., an 802.11x / WiFi router), and the user device is within communication range of the router. In some cases, the message is a broadcast message.

[0114] At block 1320, the user device (or a component thereof) may connect to the host device based on the message. In some aspects, the user device (or a component thereof) may send a password for authentication of the user to the host device. In such aspects, upon being authenticated, the user device may connect to the host device.

[0115] At block 1330, the user device (or a component thereof) may obtain an image of a three-dimensional (3D) scene of the physical environment. At block 1340, the user device (or a component thereof) may send the image to a host device.

[0116] At block 1350, the user device (or a component thereof) may receive synthetic content from the host device. A virtual representation of the user may be positioned relative to the XR room based on matching features in the image with features of the 3D scene of the physical environment. The synthetic content may include synthetic objects, one or more lighting effects, and / or other synthetic or virtual content.

[0117] At block 1360, the user device (or a component thereof) may render composite content of the XR room based on the pose of the user device. In some cases, the pose of the user corresponds to at least one of the pose of the user's head or the orientation of the user's eyes.

[0118] The user device (or computing device or apparatus) may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or a computing device of an autonomous vehicle, a robotic device, a television, and / or any other computing device with resource capabilities to perform the processes described herein, including process 1300 and / or other processes described herein. In some cases, the user device may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the user device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive data based on an Internet Protocol (IP) or other types of data.

[0119] The components of the user device may be implemented in circuitry. For example, the components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0120] Process 1300 is illustrated as a logical flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the process.

[0121] Additionally, process 1300 and / or other processes described herein may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes together on one or more processors, implemented in hardware, or implemented by a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that may be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0122] Figure 14 is a diagram illustrating an example of a system for implementing certain aspects of the present technology. Specifically, Figure 14 An example of a computing system 1400 is illustrated, which can be any computing device, for example, constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using a connection 1405. Connection 1405 can be a physical connection using a bus, or a direct connection to processor 1410, such as in a chipset architecture. Connection 1405 can also be a virtual connection, a networked connection, or a logical connection.

[0123] In some embodiments, computing system 1400 is a distributed system, in which the functionality described in this disclosure can be distributed within a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent a number of such components that each perform some or all of the functionality for which the component is described. In some embodiments, a component can be a physical device or a virtual device.

[0124] Example system 1400 includes at least one processing unit (CPU or processor) 1410 and connections 1405 that couple various system components including system memory 1415, such as read-only memory (ROM) 1420 and random access memory (RAM) 1425, to processor 1410. Computing system 1400 may include a cache 1412 of high-speed memory directly connected to, in close proximity to, or integrated as part of processor 1410.

[0125] Processor 1410 may include any general-purpose processor and hardware or software services, such as services 1432, 1434, and 1436 stored in storage device 1430, configured to control processor 1410 as well as a dedicated processor where software instructions are incorporated into the actual processor design. Processor 1410 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0126] To enable user interaction, computing system 1400 includes input device 1445, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice, etc. Computing system 1400 can also include output device 1435, which can be one or more of a plurality of output mechanisms. In some examples, a multimodal system can enable a user to provide multiple types of input / output to communicate with computing system 1400. Computing system 1400 can include communication interface 1340, which can generally govern and manage user input and system output.

[0127] The communication interface can perform or facilitate receiving and / or sending wired or wireless communications using wired and / or wireless transceivers, including utilizing audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, Wireless signal transmission, Low energy (BLE) wireless signal transmission, Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof.

[0128] The communication interface 1440 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of the computing system 1400 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There is no limitation to operating on any particular hardware arrangement, and thus the underlying features herein may be easily substituted for improved hardware or firmware arrangements as they are developed.

[0129] The storage device 1430 may be a non-volatile and / or non-transitory and / or computer-readable memory device and may be a hard disk or other type of computer-readable medium that can store computer-accessible data, such as a magnetic cassette, a flash memory card, a solid-state memory device, a digital versatile disk, a magnetic cassette, a floppy disk, a floppy disk, a hard disk, a magnetic tape, a magnetic stripe / tape, any other magnetic storage medium, flash memory, memory storage, any other solid-state memory, a compact disc read-only memory (CD-ROM) disc, a rewritable compact disc (CD) disc, a digital video disc (DVD) disc, a Blu-ray disc (BDD) disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a micro / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASH EPROM), a cache (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin-transfer torque RAM (STT-RAM), other memory chips or cartridges, and / or combinations thereof.

[0130] Storage device 1430 may include software services, servers, services, etc., which, when the code defining such software is executed by processor 1410, causes the system to perform functions. In some embodiments, hardware services that perform specific functions may include software components for performing functions stored in a computer-readable medium connected to necessary hardware components such as processor 1410, connection 1405, output device 1435, etc. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transient media in which data may be stored and does not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transient media may include, but are not limited to, disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer-readable media can store thereon code and / or machine-executable instructions that can represent a procedure, function, subroutine, program, routine, subroutine, module, software package, category, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, independent variables, parameters, or memory contents. Information, independent variables, parameters, data, etc. can be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network sending, etc.

[0131] In some embodiments, computer-readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media specifically excludes media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.

[0132] Specific details are provided in the description above to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that embodiments can be put into practice without these specific details. For clarity of explanation, in some instances, the present technology can be presented as including a separate functional block, which includes a device, device assembly, step or routine in a method embodied in a combination of software or hardware and software. Additional components other than those components shown in the accompanying drawings and / or described herein can be used. For example, circuits, systems, networks, processes and other components can be shown as components in block diagram form to avoid these embodiments becoming difficult to understand in unnecessary details. In other cases, known circuits, processes, algorithms, structures and techniques can be shown in order to avoid making each embodiment difficult to understand without necessary details.

[0133] Individual embodiments may be described above as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. Furthermore, the order of the operations may be rearranged. A process is terminated when its operations are completed, but a process may have additional steps not included in the accompanying drawings. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, termination of the process may correspond to the function returning to the calling function or main function.

[0134] The processes and methods according to the examples described above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, and the like.

[0135] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or add-in cards. By way of further example, such functionality may also be implemented on circuit boards among different chips or different processes executed on a single device.

[0136] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0137] In the foregoing description, various aspects of the present application have been described with reference to the specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although the exemplary embodiments of the present application have been described in detail herein, it is to be understood that the inventive concept can be embodied and adopted in various other ways, and the appended claims are intended to be interpreted as including such variations, unless limited by the prior art. The various features and aspects of the above-mentioned applications may be used individually or in combination. In addition, without departing from the broader essence and scope of this specification, the embodiments may be used in any number of environments and applications beyond the environments and applications described herein. Therefore, the description and the accompanying drawings should be considered as illustrative rather than restrictive. For illustrative purposes, each method is described in a specific order. It should be understood that in an alternative embodiment, each method may be performed in a different order than described.

[0138] One of ordinary skill in the art will appreciate that the less than ("<") and greater than (">") symbols or terms used herein may be replaced by less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of this specification.

[0139] Where a component is described as being “configured to” perform certain operations, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., a microprocessor or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0140] The phrase “coupled to” refers to any component being physically connected directly or indirectly to another component, and / or any component being in communication directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0141] Claim language or other language reciting "at least one of" a set and / or "one or more of" a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language reciting "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language reciting "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A, B, and C. The language "at least one of" a set and / or "one or more of" a set does not limit the set to the items listed in the set. For example, claim language reciting "at least one of A and B" or "at least one of A or B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0142] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been generally described in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the entire system. Technicians can implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be interpreted as departing from the scope of this application.

[0143] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, or integrated circuit devices with multiple uses, including applications in wireless communication devices and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium comprising program code, which includes instructions that, when executed, perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, or the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0144] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Thus, the term "processor," as used herein, may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.

[0145] Illustrative aspects of the present disclosure include:

[0146] Aspect 1. A method for providing an extended reality (XR) experience, the method comprising: receiving, by a user device associated with a user, a message from a host device including a prompt to join an XR room hosted by the host device; connecting to the host device based on the message; obtaining, by the user device, an image of a three-dimensional (3D) scene of a physical environment; sending, by the user device, the image to the host device; receiving, by the user device, synthetic content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on matching features in the image with features of the 3D scene of the physical environment; and rendering, by the user device, the synthetic content of the XR room based on a pose of the user device.

[0147] Aspect 2. The method according to aspect 1, wherein the user device is one of an augmented reality / virtual reality (AR / VR) head mounted device (HMD) or augmented reality (AR) lenses.

[0148] Aspect 3. The method according to any one of aspects 1 or 2, wherein the host device is a wireless router.

[0149] Aspect 4. The method according to any one of aspects 1 to 3, wherein the message is a broadcast message.

[0150] Aspect 5. The method according to any one of aspects 1 to 4, further comprising: sending, by the user device, a password for authentication of the user to the host device.

[0151] Aspect 6. The method according to any one of aspects 1 to 5, wherein the user device is located within the communication range of the host device.

[0152] Aspect 7. The method according to any one of aspects 1 to 6, wherein the posture of the user corresponds to at least one of the posture of the user's head or the direction of the user's eyes.

[0153] Aspect 8. The method according to any one of aspects 1 to 7, wherein the composite content comprises at least one of one or more composite objects or one or more lighting effects.

[0154] Aspect 9. An apparatus associated with a user for providing an extended reality (XR) experience, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: receive a message from a host device comprising a prompt to join an XR room hosted by the host device; connect to the host device based on the message; obtain an image of a three-dimensional (3D) scene of a physical environment; send the image to the host device; receive synthetic content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on matching features in the image with features of the 3D scene of the physical environment; and render the synthetic content of the XR room based on the pose of the apparatus.

[0155] Aspect 10. The apparatus of aspect 9, wherein the apparatus is one of an augmented reality / virtual reality (AR / VR) head mounted device (HMD) or an augmented reality (AR) lens.

[0156] Aspect 11. The apparatus according to any one of aspects 9 or 10, wherein the host device is a wireless router.

[0157] Aspect 12. The apparatus according to any one of aspects 9 to 11, wherein the message is a broadcast message.

[0158] Aspect 13. The apparatus of any one of aspects 9 to 12, wherein the at least one processor is configured to send a password for authentication of the user to the host device.

[0159] Aspect 14. The apparatus according to any one of aspects 9 to 13, wherein the apparatus is located within communication range of the host device.

[0160] Aspect 15. The apparatus according to any one of aspects 9 to 14, wherein the posture of the user corresponds to at least one of the posture of the user's head or the direction of the user's eyes.

[0161] Clause 16. The apparatus according to any one of clauses 9 to 15, wherein the composite content comprises at least one of one or more composite objects or one or more lighting effects.

[0162] Aspect 17. A non-transitory computer-readable medium of a user device associated with a user, the non-transitory computer-readable medium having instructions that, when executed by at least one processor, cause the at least one processor to: receive a message associated with the user from a host device including a prompt to join an XR room hosted by the host device; connect to the host device based on the message; obtain an image of a three-dimensional (3D) scene of a physical environment; send the image to the host device; receive synthetic content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on matching features in the image with features of the 3D scene of the physical environment; and render the synthetic content of the XR room based on the posture of the user device.

[0163] Aspect 18. The non-transitory computer-readable medium of aspect 17, wherein the device is one of an augmented reality / virtual reality (AR / VR) head mounted device (HMD) or augmented reality (AR) lenses.

[0164] Aspect 19. The non-transitory computer-readable medium of any one of aspects 17 or 18, wherein the host device is a wireless router.

[0165] Aspect 20. The non-transitory computer-readable medium of any one of aspects 17 to 19, wherein the message is a broadcast message.

[0166] Aspect 21. The non-transitory computer-readable medium of any one of aspects 17 to 20, wherein the instructions, when executed by the at least one processor, cause the at least one processor to send a password for authentication of the user to the host device.

[0167] Aspect 22. The non-transitory computer-readable medium of any one of claims 17 to 21, wherein the apparatus is located within communication range of the host device.

[0168] Aspect 23. The non-transitory computer-readable medium of any one of aspects 17 to 22, wherein the pose of the user corresponds to at least one of a pose of the user's head or an orientation of the user's eyes.

[0169] Aspect 24. The non-transitory computer-readable medium of any one of aspects 17 to 23, wherein the composite content comprises at least one of one or more composite objects or one or more lighting effects.

[0170] Aspect 25. An apparatus for providing an extended reality (XR) experience, the apparatus comprising one or more components for performing the operations according to any one of aspects 1 to 8.

[0171] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Accordingly, the claims are not intended to be limited to the aspects shown herein, but are to be accorded the full scope consistent with the language of the claims, wherein reference to an element in the singular is not intended to mean "one and only one" unless specifically stated otherwise, but rather "one or more."

Claims

1. A method for providing an extended reality (XR) experience, the method comprising: receiving, by a user device associated with a user, from a host device, a message including a prompt to join an XR room hosted by the host device; connecting to the host device based on the message; obtaining, by the user device, an image of a three-dimensional (3D) scene of a physical environment; The user device sends the image to the host device; receiving, by the user device, composite content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on matching features in the image with features of the 3D scene of the physical environment; and The composite content of the XR room is rendered by the user device based on the pose of the user device.

2. The method of claim 1, wherein the user device is one of an augmented reality / virtual reality (AR / VR) head mounted device (HMD) or augmented reality (AR) lenses. The method of claim 1 , wherein the host device is a wireless router. The method of claim 1 , wherein the message is a broadcast message.

5. The method according to claim 1, further comprising: The user device sends a password for authentication of the user to the host device. The method of claim 1 , wherein the user device is within a communication range of the host device. The method of claim 1 , wherein the user's posture corresponds to at least one of a posture of the user's head or a direction of the user's eyes.

8. The method of claim 1, wherein the composite content comprises at least one of one or more composite objects or one or more lighting effects.

9. An apparatus associated with a user for providing an extended reality (XR) experience, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory and configured to: receiving, from a host device, a message including a prompt to join an XR room hosted by the host device; connecting to the host device based on the message; Obtaining an image of a three-dimensional (3D) scene of a physical environment; sending the image to the host device; receiving composite content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on matching features in the image with features of the 3D scene of the physical environment; and The composite content of the XR room is rendered based on the pose of the device.

10. The apparatus of claim 9, wherein the apparatus is one of an augmented reality / virtual reality (AR / VR) head mounted device (HMD) or an augmented reality (AR) lens. The apparatus of claim 9 , wherein the host device is a wireless router. The apparatus of claim 9 , wherein the message is a broadcast message.

13. The apparatus of claim 9, wherein the at least one processor is configured to send a password for authentication of the user to the host device. The apparatus of claim 9 , wherein the apparatus is located within communication range of the host device. 15 . The device of claim 9 , wherein the user's posture corresponds to at least one of a posture of the user's head or a direction of the user's eyes.

16. The apparatus of claim 9, wherein the composite content comprises at least one of one or more composite objects or one or more lighting effects.

17. A non-transitory computer-readable medium of a user device associated with a user, the non-transitory computer-readable medium having instructions that, when executed by at least one processor, cause the at least one processor to: receiving, from a host device, a message including a prompt to join an XR room hosted by the host device; connecting to the host device based on the message; Obtaining an image of a three-dimensional (3D) scene of a physical environment; sending the image to the host device; receiving composite content from the host device, wherein a virtual representation of the user is positioned relative to the XR room based on matching features in the image with features of the 3D scene of the physical environment; and The synthesized content of the XR room is rendered based on the pose of the user device.

18. The non-transitory computer-readable medium of claim 17, wherein the device is one of an augmented reality / virtual reality (AR / VR) head mounted device (HMD) or augmented reality (AR) lenses.

19. The non-transitory computer-readable medium of claim 17, wherein the host device is a wireless router.

20. The non-transitory computer-readable medium of claim 17, wherein the message is a broadcast message.

21. The non-transitory computer-readable medium of claim 17, wherein the instructions, when executed by the at least one processor, cause the at least one processor to send a password for authentication of the user to the host device.

22. The non-transitory computer-readable medium of claim 17, wherein the apparatus is within communication range of the host device.

23. The non-transitory computer-readable medium of claim 17, wherein the pose of the user corresponds to at least one of a pose of the user's head or an orientation of the user's eyes.

24. The non-transitory computer-readable medium of claim 17, wherein the composite content comprises at least one of one or more composite objects or one or more lighting effects.