Shared 3d view viewing of assets in augmented reality systems

By implementing shared perspective viewing and avatar heap mode in XR environment, the problem of maintaining shared perspective and spatial relationships in multi-user collaboration is solved, and collaboration efficiency and user experience are improved.

CN119987536APending Publication Date: 2025-05-13AUTODESK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411485445.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-10
Filing Date
2024-10-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In a multi-user collaboration environment, it is difficult for the prior art to effectively maintain all users' shared perspective on content while maintaining spatial relationships between users, resulting in a decline in collaboration efficiency and user experience.

Method used

By implementing shared perspective viewing in an extended reality (XR) environment, users can view assets from the same perspective and maintain spatial relationships between users through avatar heap mode. This technology ensures that all users have the same view of the perspective in the XR environment through perspective transformation and avatar repositioning, while the user's avatar looks distant in space.

Benefits of technology

It realizes efficient collaboration between multiple users, improves the user experience, enables all users to view content from the same perspective, while maintaining spatial relationships between users, and enhancing the immersive experience of collaboration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987536A_ABST
    Figure CN119987536A_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including medium-encoded computer program products, for computer-aided design of physical structures, the methods, systems, and apparatus comprising: rendering a view of a first user of an augmented reality environment to a display device of the first user, the first user view is generated using a first gesture tracked for the first user; identifying that a first avatar associated with the first user has entered a shared view mode of a second avatar associated with a second user of the XR environment, wherein the shared view mode has a single frame of reference shared by the first user and the second user in the XR environment; and performing the rendering to the display device of the first user when in the shared viewing angle mode.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present specification relates to shared perspective viewing of content between multiple users.In addition, the present specification relates to physical model data used in computer graphics applications (such as computer-generated animation of physical structures and / or computer-aided design and / or other visualization systems and techniques).

[0002] Computer systems and applications may provide a virtual reality (VR) environment that enables users to perform individual or collaborative tasks. The VR environment may be supported by different types of devices that may be used to provide interactive user input (e.g., goggles, joysticks, sensor-based interaction support tools, pens, touchpads, gloves, etc.). By utilizing interactive techniques and tools supported by the VR environment (e.g., a VR collaboration platform), users may be able to work independently or collaboratively with other users in the same and / or different three-dimensional spaces on one or more projects.

[0003] The display device of a user collaborating with other users in a VR environment may render assets in respective views and avatars of respective other users participating in the VR environment. In some cases, the avatars may move their positions and perform collaborative actions with respect to the assets. The avatar of one user may be rendered in the view of another user based on positional data indicating an identification of a position in the VR environment (e.g., based on a tracked position associated with a VR controller of the other user). Summary of the invention

[0004] The present specification relates to shared perspective viewing of one or more assets between users in an extended reality (XR) environment. Shared perspective viewing may be associated with viewing of an asset, which may be a set of content including 2D and / or 3D content. In some instances, the set of content may be, for example, a model object, which may be a 3D object. In some cases, an asset that may be viewed in an XR environment may have a 3D orientation in the XR environment. For example, 2D content (such as a presentation of information on a virtual screen in an XR environment) may be rendered in a user's view of the XR environment, where the 2D content may have a 3D orientation and may be repositioned to change its orientation. When users in the XR environment are provided with a shared perspective view of an asset, they may also be provided with a view of an avatar of other users to maintain a spatial relationship between viewers sharing the perspective. An avatar may be rendered as a 2D or 3D object in the XR environment and may have a position and orientation for viewing space in the XR environment. Shared perspective views can be performed in the context of collaborative work on projects (such as architectural design projects, including construction projects, product manufacturing projects, design projects (e.g., design of machines, buildings, construction sites, factories, civil engineering projects, etc.)) as well as in virtual reality (VR) games and other immersive experience interactions between users viewing the same content.

[0005] In some implementations, the XR collaboration platform may provide tools and techniques for generating, maintaining, and operating an interconnected environment across multiple devices, where assets may be presented from a shared perspective while maintaining spatial relationships between avatars. In this way, the view of an asset may be maintained as the same when presented to multiple users, while at the same time, the user may view other users as spatially distant avatars to maintain an understanding of multi-user viewing. In some implementations, a shared perspective view may be provided by rotating the asset based on a shared view to match the perspective of each user. In some implementations, a shared perspective view may be provided by changing the position of the user's avatar to match the shared perspective so that the user can perceive the model from the same perspective. In those implementations, users may be gathered in an avatar stack, which may be visualized as an avatar tray included in a shared perspective mode in front of a view generated for the respective user. The tray may include an avatar corresponding to the user in the shared perspective mode. The avatars in the station may be rendered as 3D avatars (one or more of them and combined with one or more of them as 2D avatars) with 3D positions and 3D orientations, where the avatars in the pile will be rendered with a certain offset based on the 3D position and orientation of the pile (and optionally the number of avatars to be included in the station) so that the avatars (in the station) can be included in a shared view mode, where the avatars have the same viewpoint. In the avatar pile mode, since all avatars have the same viewpoint, for each of the avatars in the pile, the other avatars are rendered in their view with a certain offset from their position so that the other avatars can be seen. The other avatars are rendered by being repositioned according to a predefined offset from the shared position, where their gaze and pointing towards content in the XR environment are reoriented as adjustments are needed in view of the offset from the avatar's viewpoint.

[0006] In some implementations, when a first user has a shared perspective view of a model object in a collaboration session, an avatar of a second user identified as an active speaker in the collaboration session and having a shared perspective toward the model object may be rendered as a 3D avatar and may have a 3D orientation and position adjusted to be included in the first user's view.

[0007] Multiple interconnected environments may be provided in an XR environment of an XR system, which may include a VR environment, an augmented reality (AR) environment, a desktop environment, a web preview environment, a mobile native application environment, and other example environments that may be supported by different types of devices and corresponding interactive user input devices (e.g., goggles, joysticks, sensor-based interactive support tools, pens, touchpads, gloves, etc.) used by users entering the XR environment. In such an interconnected environment, a group of users may be gathered in a shared perspective mode. For example, when a group of users indicates that they want to collaborate on an asset (such as an XR object), and the object is rendered from the same perspective to each of the users with avatars having different positions in the XR environment, grouping the users in such a shared perspective mode may be triggered. Even if the avatars are substantially far away from each other in the XR environment, the users of those avatars may have a shared perspective of a set of content in the XR environment and be provided with views of the avatars of other users sharing the same perspective view. In another example, grouping of users may be triggered based on identifying that two or more user avatars are in close proximity in the XR environment (below a certain distance threshold, e.g., within half of the minimum single-dimensional size of the avatars), and then grouping them in a shared mode where the avatars are associated with the same location for viewing. When a user wants their avatar to join a group of other avatars, the avatar is repositioned to the location of the group (pile). When multiple user avatars are piled up, they are positioned at the same location in the XR environment to provide the same perspective in the XR environment; note that each user may still be allowed to control their own view and orientation, such as through movement of the head, through gestures, or through instructions received at a controller or sensor associated with the respective user. When a view of the XR environment is rendered for one of the avatars, the view includes presentation of other avatars (from the pile) having the same viewing location in the XR environment. The other avatars are presented in the view as they are repositioned and reoriented to be rendered in the user's view of the XR environment. By entering the pile, the avatar can see other avatars that share the location in the XR environment and can see where the other avatars are looking and / or pointing.

[0008] In some implementations, grouping may be performed based on an active indication from a user to join a sharing mode with other users (e.g., a group that can be viewed as a user station from a third-person perspective). In some cases, moving an avatar away from the group may trigger a modification to the grouping to exclude avatars that have been relocated further away from the group. In some cases, users may navigate the position of their avatars so that, through interaction with a user input device, they switch positions, join a group, or move to another group. By utilizing interactive techniques and tools supported by the XR collaboration platform, users are able to have a shared perspective view of assets even if their avatars are spatially far away from each other. The interactive techniques and tools of the present invention support an immersive experience in which each group of users shares their perspective of viewing a portion of the content in the XR environment (e.g., content representing an XR object) and is still part of a multi-user viewing experience in which the avatars of other users can be seen. Each user has a different perspective view of the other avatars. In some instances, users may have a shared perspective view (a shared perspective view of a set of content (such as 3D content)) in an XR environment, where transformations may be applied to: (a) the set of content (e.g., a 3D model) to reorient and reposition the set of content for presentation to each avatar in shared perspective mode, or (b) other avatars of other users in shared perspective mode to reorient and reposition those avatars when they are rendered in each user's view when the avatars are stacked and have a shared position so that the avatars are visible to the users.

[0009] In some implementations, the content may be rotated for each user so that each user who is part of a shared perspective view may have the same 3D perspective for viewing assets in the XR environment and different and specific views of other user avatars in the XR environment. With this immersive experience, users may have an improved view of shared content regardless of the user's exact location or the number of users participating in the experience. Although the 3D perspective for viewing the content may be shared, since the user may be viewing the content from the same direction as another user in the XR environment, the users still maintain an experience of having spatial distance relative to each other because the user's view may include the rendered avatars of the other users offset based on the shared reference frame used to view the content. When a user interacts with the content (e.g., points to a portion of the shared content), the pointing may be handled so that other users can see the pointing to the exact place in the previewed model as the place / position that the first user was pointing to when looking at his view on his device. A user viewing shared content may see a disposition from another avatar (or the avatar's hand) pointing toward that content, where the disposition may be rendered based on applying a directional offset to the avatar's hand (when the avatar is a 3D avatar) to mirror the content visible in the real world to match the direction of the point on the model toward which the user is aiming.

[0010] In some implementations, to share a perspective view of content in an XR environment, an avatar stack may be created so that users in close proximity to each other may be grouped in the stack and included in a shared mode, where the position of the avatar for viewing may be changed instead of rotating the content to be viewed, so that multiple avatars are placed in a shared perspective mode associated with a specific position in the XR environment. The avatar stack is formed and has a position and orientation in the 3D space of the XR environment, and a visual table may be rendered in the user's view to include the avatars of other users in the shared perspective view in the XR environment. The table with the avatars of the users in the XR environment may be shown in front of the shared perspective of the stack of each user. The avatars as shown in the table may be presented in the table differently in the view of each respective user because they will be offset based on the 3D position of the stack and the 3D orientation of the stack (a movement orientation that defines the movement direction of the avatar if the user's input for advancing is provided). Thus, in those cases, the user's adjusted perspective allows a group of users to have a shared 3D perspective of content in the XR environment while being provided with visualization of the other users sharing the view in the rendered view. The positioning of the avatars in the view allows the gestures and pointing of the users to be seen, as well as the active speaker or presenter of the group to be identified. Such groups of stacked avatars may be viewed from a third person perspective as a group comprising a defined set of users represented by their avatars, names, or a combination thereof, wherein the third person perspective may allow users out of range to determine who is speaking at any moment, as well as see how avatars from the pile are pointing to portions of the shared content, wherein a pointer presented by one avatar from the pile may be accurately presented to the view of a user outside the pile who is seeing the content from a different perspective.

[0011] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the invention will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 An example of a system that can be used to support collaboration between users sharing views of an asset in a shared view mode in an XR environment is shown.

[0013] Figure 2A An example desktop review mode provided for display to multiple users accessing an XR environment through different devices viewing an asset in a shared view mode is shown in accordance with implementations of the present disclosure.

[0014] Figure 2BAn example reorientation of a 3D object in a desktop review mode provided for display to four users accessing an XR environment through different devices viewing the asset in a shared view mode is shown in accordance with implementations of the present disclosure.

[0015] Figure 2C Example viewing modes provided for display to multiple users accessing an XR environment through different devices viewing an asset in a shared view mode in accordance with implementations of the present disclosure are described.

[0016] Figure 2D An example viewing mode of a user interacting with an XR environment providing an option to join a shared view mode of an avatar pile is shown in accordance with implementations of the present disclosure.

[0017] Figure 2E An example interaction of an avatar with a pointer to an asset rendered in a shared view mode is shown in accordance with implementations of the present disclosure.

[0018] Figure 2F An example user interface of a collaboration application implementing a shared perspective mode according to an implementation of the present disclosure is shown.

[0019] Figure 2G An example user interface including a first-person view from a collaborative application implementing a shared perspective mode is shown in accordance with implementations of the present disclosure.

[0020] Figure 3 An example of a process for providing a shared perspective view of assets in an XR environment in accordance with implementations of the present disclosure is shown.

[0021] Figure 4A An example of a process for providing a shared view mode for shared viewing of XR objects in an XR environment in accordance with implementations of the present disclosure is shown.

[0022] Figure 4B An example scheme for performing perspective transformation when rendering an object in a shared perspective mode according to an implementation of the present disclosure is shown.

[0023] Figure 5A An example of a one-to-one scale collaborative review mode of users provided with an avatar stack perspective is shown in accordance with implementations of the present disclosure.

[0024] Figure 5B An example of a process for providing an avatar stack mode for shared perspective viewing of objects in an XR environment in accordance with implementations of the present disclosure is shown, wherein user avatar tables are rendered.

[0025] Fig. 6AAn example user view including an avatar stack mode with an active speaker is shown in accordance with implementations of the present disclosure.

[0026] Figure 6B An example of an avatar stack mode providing representations of gestures and gazes of an active speaking avatar is shown in accordance with implementations of the present disclosure.

[0027] Figure 7 An example avatar stack mode from a third user perspective is shown in accordance with implementations of the present disclosure.

[0028] Figure 8 is a schematic diagram of a data processing system including a data processing device that can be programmed as a client or a server.

[0029] Like reference numbers and designations throughout the various drawings indicate like elements. DETAILED DESCRIPTION

[0030] This disclosure describes various tools and techniques for providing users with shared perspective views of 3D content in an XR environment.

[0031] When users are engaged in a meeting or other collective activity involving viewing content, such as assets (e.g., documents, computer models (2D or 3D models), objects, motion pictures, structures, textures, materials, groups of user avatars in 3D space, etc.) in an XR (VR, AR, or mixed reality) environment, the users may be provided with the same perspective view of at least some of the content simultaneously, even if the users' positions and perspectives, as identified by the devices they use to access the environment, are associated with different viewpoints.

[0032] Currently, when users interact in a collaborative space, each has their own viewpoint that is discrete from all other users, which matches the viewing experience in the real world. However, in cases where sharing the same view is relevant to interactive activities, or where presenting the same view is relevant to the user's experience (e.g., participating in a 3D movie), this viewing from discrete viewpoints may impose limitations on collaborative work on a project. Having discrete viewpoints may provide some users with a worse viewing position than other users, or may limit users from viewing parts of a shared asset that are relevant to an interaction by another user (e.g., pointing to a section of a model that is not visible from another user's viewpoint). For example, a movie theater has multiple seats, of which typically only one or two seats have the "best view." In addition, if user avatars are positioned in the same position to be provided with exactly the same view of the content, the users will not see the other avatars in the XR environment, and therefore they will not be provided with an interactive experience with other users (e.g., they will not see how they point to the content from the shared view). Thus, having all users in a group share the same perspective view while still maintaining the experience of activities in a group (e.g., in a movie theater or a workshop or at a table) with others who may be considered spatially distant may improve the user experience. According to implementations of the present disclosure, perspective transformations may be applied to shared content or to avatars presented in a user's view to improve the user's experience. In some instances, tools and techniques may be provided for presenting content (e.g., 2D or 3D content) in a shared mode, where the content is presented to each user from the same shared perspective. With those tools and techniques, a view of an asset may have a shared perspective for all users, while the user may still view the avatars of other users in the 3D space as spatially distant (and seen from different angles in each of the user's views) to support an immersive experience where a portion of a 3D environment is seen by all users from a shared perspective, while the views of the user's avatars are offset based on each user's respective reference frame to integrate into each view.

[0033] Synchronous remote collaboration typically requires multiple tools and techniques to support interactions between users during a creative process where there is no fixed order of operations defined for each user. Therefore, users expect collaborative applications to provide flexible and seamless tools to integrate into the interactive process through common shared documents. Mixed focus collaboration is key to work processes such as brainstorming, design review, complex content creation, decision making, script making, etc. Mixed focus collaboration may involve simultaneous editing of a 3D model of an object by multiple users at multiple locations and / or at the same location in an interactive manner. Therefore, in collaborative mode, the model is viewed by users from multiple different remote instances, where one or more users can edit the model. According to an implementation of the present disclosure, the 3D model can be presented to the user on the user's device from the same shared perspective with a corresponding rendered view of each of the users.

[0034] In some instances, users can interact with shared assets (e.g., data, objects, avatars of other users) using different interfaces viewed from different devices or platforms (including desktop devices, tablets, VR and / or AR devices, and other examples). In some implementations, VR and AR devices may be referred to as mixed reality (MR) or XR devices. Computer graphics applications include different software products and / or services that support the generation of representations of three-dimensional (3D) objects that can be used for visualization of object models, animation, and video rendering. Computer graphics applications also include computer animation programs and video production applications that generate 3D representations of objects and views during collaborative user reviews. 3D computer animations can be created in a variety of scenarios and in the context of different technologies. For example, 3D models of objects (such as manufacturing plants, buildings, physical structures) can be animated for display on user interfaces of native device applications, web applications, VR applications, AR applications, and other examples. Prototype models of objects can be executed in different environments, including VR environments and based on VR technology, AR environments, display user interfaces, remote video streaming, and the like.

[0035] In some implementations, users interact with shared content (such as a 3D model of an object) in an XR environment, where the interaction can be supported by various environments that users access through different devices. XR collaboration can include video teleconferencing, audio teleconferencing, or video and audio teleconferencing, virtual and / or augmented reality collaborative interaction environments, or combinations thereof. In some instances, one user can participate in a VR interactive session, and another user can access a teleconference connection to the VR interactive session via a video and audio connection (e.g., relying on video and audio tools and settings for interaction), where the user can receive a real-time streaming of the VR interactive session, and the other user can participate only through a video connection (e.g., relying only on video tools and settings (e.g., cameras) to stream the video feed of the participant during the interaction with other users). Users can join different sessions based on devices that support access to different sessions. In some instances, according to implementations of the present disclosure, a VR interactive session may include multiple users, and the multiple users can view a displayed 3D model of an object from a shared perspective.

[0036] People can use different user interface applications to access the XR collaboration platform for real-time, high-bandwidth communication and use real-time groupware to synchronously work on assets (such as 3D models of objects, interface designs, construction projects, machine designs, combinations of 3D and 2D data, and other example assets). According to implementations of the present disclosure, the XR collaboration platform can be implemented and configured to enable remote collaboration that can produce a smooth and flexible workflow supported by efficient tools that reduce the number of user interactions with collaborative applications, because the shared content will be presented from a shared perspective, while at the same time each collaborative user can see other avatars and the locations where those avatars are looking or pointing at the shared content, thereby improving the timeliness and efficiency of the process flow.

[0037] In some implementations, the XR environment may be accessed through a VR device (such as a VR headset, a head-mounted display, a smartphone display, or other device), and the VR display may present a collaborative XR environment that may be rendered to multiple users wearing VR headsets to provide them with an immersive experience in a 3D space, where assets are shown in a shared perspective mode. For example, users may collaborate on a 3D model of an object, which is rendered to multiple users in a review mode within the collaborative VR environment by rotating the 3D model of the object according to the reference frames of the respective users.

[0038] In some implementations, an avatar's frame of reference may be established when the avatar enters a shared view mode with one or more other avatars. Establishment of the frame of reference may be performed using positional information from the tracked posture of the first user upon entry. In some cases, the user of the avatar in the shared view mode may continue to turn their head, thereby changing the tracked posture, thereby slightly changing the rendering perspective. However, specific threshold ranges for such movement and posture changes may be established to allow the user to move while still continuing to be part of the shared view mode. By moving the head slightly, the avatar's view may change slightly without the avatar exiting the shared view mode.

[0039] In some implementations, a user may be represented in an XR environment by displaying avatars that may be rendered as different types of avatars (e.g., 2D or 3D avatars with different shapes, colors, annotations, etc.). For example, the type of avatar that may be used to represent a user may be based on the type of device the user uses to connect to the VR environment, based on settings defined for the user (e.g., gender, age, selected appearance, etc.), or based on current interactions or instructions performed by the user in the XR environment (e.g., a determination that the user is the active speaker, a user selection of a button for modifying the user's avatar, etc.). The XR environment may be configured to receive input and / or user interaction from a VR controller of one or more computer devices connected to the user via a wired or wireless connection. Users interacting with the XR environment may be evaluated to determine whether they are authorized to access and view shared assets in the XR environment.

[0040] Figure 1 An example of a system 100 that can be used to support collaboration between users who share views of an asset in a shared view mode configured in an XR environment is shown. A computer 110 includes a processor 112 and a memory 114, and the computer 110 can be connected to a network 140, which can be a private network, a public network, a virtual private network, etc. The processor 112 can be one or more hardware processors, and the one or more hardware processors can each include multiple processor cores. The memory 114 can include both volatile memory and non-volatile memory, such as random access memory (RAM) and flash RAM. The computer 110 may include various types of computer storage media and devices, which may include the memory 114 to store instructions for programs running on the processor 112, the programs including VR programs, AR programs, teleconferencing programs (e.g., video teleconferencing programs, audio teleconferencing programs, or combinations thereof) that support view sharing and editing functions.

[0041] Computer 110 includes application 116, which includes implemented logic that allows users to interact with shared content in VR space. For example, application 116 can help users who are engaged in a common task (e.g., working on a common shared 3D model in parallel and / or synchronously) achieve their goals by providing them with a shared perspective view of the shared content. Application 116 can be groupware designed to support group intention processes and can provide software functionality to facilitate these processes.

[0042] The application 116 may be executed locally on the computer 110, remotely on a computer of one or more remote computer systems 150 (e.g., one or more server systems of one or more third-party providers accessible by the computer 110 via the network 140), or both locally and remotely. In some implementations, the application 116 may be an access point for accessing services running on an XR collaboration platform 145 that supports viewing of shared data (e.g., object model data stored at a database 160 on the XR collaboration platform 145) and / or other assets in an XR environment (e.g., structures, textures, user avatars, other objects, or portions thereof).

[0043] In some implementations, application 116 may obtain 3D model 161 (e.g., architectural design documents, construction designs, floor plans, etc.) from database 160. 3D model data such as stored at database 160 may be accessed as a single source of truth when providing representations of the model to users in the XR environment in various review modes. In some implementations, database 160 may store multi-dimensional document data, including charts, spreadsheets, diagrams, 3D models, image data, floor plans. In some instances, displaying data (including the content of the XR environment) on display device 120 to generate a view for user 190 may be based on obtaining 3D and 2D data from database 160.

[0044] A user 190 may interact with the application 116 to initiate an XR environment (e.g., a computer-generated representation of the physical world including physical and pseudo-structures (such as an architect's office) that represents a collaborative space or representation of the interior space of a physical structure (such as a building) or join an already initiated XR environment to view shared assets 130 (such as a 3D model 161) and share the same perspective view with other users associated with the XR collaboration platform 145 and the XR environment. In some implementations, the application 116 may provide a user interface 132 (e.g., including a portion of the shared assets presented in review mode) on a connected display device 120 to display the shared assets to a group of users in a shared perspective mode. The remaining users may receive a shared view of the assets at a different instance of the collaboration application 116. A shared view may be provided by rotating a shared asset when generating a view for each of the users in a shared view mode or by moving the users' viewpoints (e.g., when the spatial distance between the users is below a certain threshold) to match a single reference frame, while providing a multi-user experience to each of the users by rendering other avatars at a certain spatial distance from each other and with respective orientations corresponding to the particular user's view. The presentation of the other users may be performed by rendering their avatars at a certain offset based on a shared reference frame or shared viewpoint of the stack of users viewing the XR environment.

[0045] In some implementations, the application 116 may be operated using one or more input devices 118 (e.g., a keyboard and a mouse) of the computer 110. Figure 1110, but the display device 120 and / or the input device 118 may also be integrated with each other and / or with the computer 110, such as in a tablet computer (for example, a touch screen may be an input device 118 / output device 120). In addition, the computer 110 may include a VR or AR system or may be part of a VR or AR system. For example, the input device 118 / output device 120 may include a VR / AR input controller, a glove or other manual manipulation tool 118a, and / or a VR / AR head-mounted device 120a. In some instances, the input / output device may include a sensor-based hand tracking device that tracks movement and recreates interactions as if performed using a physical input device. In some implementations, the VR and / or AR device may be a stand-alone device that may not need to be connected to the computer 110. The VR and / or AR device may be a stand-alone device with processing capabilities and / or an integrated computer (such as the computer 110), for example, with input / output hardware components (such as controllers, sensors, detectors, etc.). VR and / or AR devices connected to the computer 110 or as standalone devices integrated with a computer (having a processor and memory) may communicate with the XR collaboration platform 145 and immerse users connected through these devices into a virtual world, where 3D models of objects may be presented in a simulated real physical environment (or a substantially similar environment) from a shared perspective, and users are represented with corresponding avatars to replicate real interactions.

[0046] In some implementations, the system 100 may be used to display data from 3D and 2D documents / models that may be used to generate an XR environment (which may be a VR environment for one or more first users, an AR environment for one or more second users, and a non-VR or non-AR environment for one or more third users) that is presented in a corresponding interface on the display device 120, thereby allowing the user to use the data to navigate, modify, adjust, search, and perform other interactive operations that may be performed on the data presented in the 3D space of the virtual world.

[0047] In some implementations, users, including user 190, may connect, access, and interact in the same virtual world as may be displayed by system 100, relying on different rendering technologies and different user interactive input / output environments. Thus, XR collaboration platform 145 may support interconnected user interfaces that may render virtual worlds in which multiple users interact, either at the same geographic location or at remote locations.

[0048] The XR collaboration platform 145 may be a cloud platform that may be accessed to provide cloud services (such as visualization services 147 related to data visualization in different review modes) and may support tools and techniques for interacting with content in a shared perspective mode in a cross-device setting to navigate and modify content when accessing data from the database 160 from multiple points associated with different devices. A user may be provided with a view in an XR environment, where assets in the XR environment may be presented to groups of users simultaneously from a shared perspective. In some cases, a user interacting with a virtual world may have the user's 190 view rendered by different types of devices supporting different graphics rendering capabilities. For example, a display of a rendering of a virtual world may be provided to some users via a VR device, a display of a rendering of a virtual world may be provided to other users via an AR device, and a display of a rendering of a virtual world may be provided to yet other users via a desktop monitor, a laptop, or various types of devices. It should be understood that a user's view of a shared asset on the user interface 132 may change while the user is viewing the asset, such as based on other users manipulating, interacting, modifying, or otherwise changing the view of the asset (e.g., movement of another user's avatar in shared mode viewing).

[0049] The systems and techniques described herein are applicable to any suitable application environment that can graphically render any portion of a virtual world (including objects therein). Thus, in some implementations, the XR collaboration platform 145 supports view sharing in the same viewing mode while maintaining the presentation of avatars of other users who remain spatially distant in the shared view, e.g., as described with respect to Figure 2A , Figure 2B , Figure 2C , Figure 2D , Figure 3 Figure 4 Figure 5A and Fig. 6A described.

[0050] Figure 2A An example desktop review mode is shown provided for display to multiple users accessing the XR environment 200 through different devices viewing the asset 210 in a shared view mode in accordance with implementations of the present disclosure.

[0051] In some implementations, a user may access the system (such as Figure 1 The XR environment is accessed by a system 100, and the asset 210 viewed by the user can be viewed from an XR platform such as Figure 1 The database of the XR collaboration platform 145) is obtained.

[0052] In some implementations, a user may generate, create, or upload a model of an object at an XR application, which may be accessed through various access points (such as different user interfaces accessible through a desktop device, a mobile device, a VR device, an AR device, etc.). The model of an object may include the 3D geometry of the object, metadata, and a 2D floor plan, such as a construction site, a building, and an architectural design configuration, as well as other examples of technical configuration and / or design data for a construction or manufacturing project. There may be multiple hierarchies of model definitions for an object based on the spatial distribution of the model (e.g., the bounding volume of each part of the model, starting from the exterior), and based on the metadata (e.g., building->floor->room, or column->beam->slab, etc.).

[0053] In some instances, using sensors on a user's head mounted device interacting with the XR environment, the XR collaboration application may map the user's real world table to their virtual table within the XR experience when displaying one or more 3D models. This will act as (passive) haptic feedback to facilitate hand-based interaction with their personal virtual table. The user may authenticate to a cloud system storing their architectural model or at an XR application when the model is accessible (e.g., as stored at the XR application's storage device or at an external storage device communicatively coupled to the XR application). The virtual table allows the user to browse the cloud storage device or another storage device accessible from the VR interface (e.g., a hub, project, folder, etc.) and select a model (uploaded by the user or another user but still accessible) for review in the XR space.

[0054] In some implementations, when Figure 2A As shown, when a 3D model is presented in a desktop view, some or all of the 3D data associated with the model may be downloaded and used for rendering. For example, different levels of detail associated with rendering a 3D model may affect the amount of data downloaded for data visualization. Users interacting with the virtual world of an XR environment may be rendered to represent the physical world in a desktop review mode and / or other modes in which a 3D model of an object is displayed and may be viewed as an avatar within a corresponding representation of the virtual world provided to each user based on an interface input and output device for participating in the virtual world (e.g., a VR device may be used to participate in a VR experience, while an AR device such as an IPAD may be used to render the virtual world on a screen of the IPAD as a display stream to connect the VR experience of a user on the VR device with a user on the AR device). Connection of users through different user interfaces is supported on a device configured to access an XR collaboration platform (e.g., Figure 1The XR collaboration platform 145 of the present invention is used to provide an XR collaboration experience on different types of devices that participate in virtual collaboration, and the virtual collaboration is rendered with rendering tools and techniques corresponding to the devices used by the users and the viewing characteristics of the review mode on those devices. Viewing within the collaboration session between users can be performed based on rendering the content of the XR environment from a shared perspective view while maintaining the spatial distance between the users, and the user experience of multi-user activities can be maintained.

[0055] In some implementations, the desktop review mode allows a user to immerse in a virtual world that can simulate a meeting style set up around a dollhouse scale model on a central pedestal. The user can import any 3D model from the XR collaboration platform (e.g., a cloud storage device, a database, or a hard drive, among other example storage devices) to review here in the review mode. Additionally, the user interface of the VR environment can include a whiteboard visualized on a virtual wall that can present 2D data related to a model of an object presented in the desktop review mode in the middle of the virtual table (on the central pedestal).

[0056] In some implementations, users may wish to collaborate on an asset that may represent content, such as a 3D model of an object that may be shared with a group of users in the XR environment 200. If users are provided with a view of the 3D model of the object when they enter the XR environment, each user will have a different view of the 3D model that may match their posture (e.g., as recognized by their respective user devices (e.g., goggles, VR headsets, controllers, sensors, etc.) used to access the XR environment). If the views are provided based on the position of the users' avatars in the XR environment, the users will have different views of the 3D model when their respective avatars are at different positions relative to the 3D model in the XR environment. Thus, even if users are provided with options for viewing content in the XR environment 200, and they may see the same object, they will have different perspectives on these views, and therefore the content they see will be different. For example, such differences in viewing perspectives may not be desirable when a group of users are working in a 3D space and want to view and manipulate objects in a collaborative work task.

[0057] To address these shortcomings of sharing asset views without providing a shared perspective, the XR environment 200 may support a shared perspective view mode for a group of user avatars that wish (or are defined to) share a view of an asset in the XR environment 200. In such an implementation, each of the users will be presented with the same perspective view of the asset 210, which may be the view of one of the users in the group, or may be a view that may be preconfigured as a view for sharing a common perspective with the users. For example, some assets may be assigned a "best" perspective view that may be used as a default perspective for sharing with other users in the XR environment 200. When sharing a perspective, each user's view is generated using a posture of the user tracked by the user's device (e.g., the 3D position and orientation of a head mounted device used by the user to enter the XR environment 200) and a perspective transform applied to the shared asset 210 based on a shared reference frame for the asset 210 (i.e., a shared specific perspective view) and a reference frame established for the corresponding user's avatar. Each user's view may be generated and rendered to the user's display device, such as Figure 2B shown.

[0058] Figure 2B An example perspective transformation of a 3D model of an object in a desktop review mode provided for display to four users accessing an XR environment through different devices viewing the asset in a shared perspective mode in accordance with implementations of the present disclosure is shown. The perspective transformation may be applied to the 3D model so that each respective user has a shared perspective mode in which the 3D model is reoriented (e.g., a rotation about a single axis is applied in a simple desktop view mode, or a rotation about more than one axis is applied) and / or content is repositioned (e.g., as described below) Figure 2C ). When a user enters a Figure 2A When the avatar is in a collaborative space of an XR environment 200, each user has their own viewpoint that is discrete and matches their view to the XR environment (their viewpoint and view orientation). The viewpoint is the position of the avatar in the XR environment, and the position can change based on user input, for example, based on the user moving in the physical world, based on user joystick input, based on the user pointing at something in the XR environment and clicking a button to reposition the avatar within the XR environment, etc. The avatar's view orientation is the 3D direction of the view of the XR environment. In some instances, this view orientation can be changed based on user actions, such as when the user turns their head while wearing a VR headset or based on other user instructions.

[0059] In order to provide an experience in this XR environment (e.g., VR or AR environment) where the view of the rendered asset is maintained from the same perspective for all users (e.g., like a video conference call where the display is shared and maintained the same for all participants at the same time), a shared perspective may be defined with respect to the asset being viewed while still having different spatial relationships between the avatars of the users experiencing the shared viewing. The shared perspective may be managed so that each avatar is provided with a combination of a fixed perspective of the content (asset) they are viewing and a variable perspective of the avatar with which each avatar (person) is viewing the asset.

[0060] 2, four users 221, 222, 223, and 224 are provided with a desktop review mode of an asset 230 that is a model of an object. When viewing asset 230, the users enter a shared perspective mode, with each of the four users viewing the asset from the same perspective.

[0061] The shared perspective mode can provide a shared viewpoint of an asset (e.g., a 3D model of an object) to multiple users, in this case four users in a virtual environment. Thus, at 220, which is a view of user 223, user 223 is viewing asset 230 from a given perspective, which can be defined as a perspective to be shared with three other users 221, 222, and 224. Thus, such perspective views of asset 230 can be provided at views 225, 226, and 227 of the other three users. Figure 2B In the example of , a round table and the 3D model on it are surrounded by user avatars. In order to provide a shared perspective mode, the model is rotated around the axis of the table so that all users can view the model from the same angle. In order to provide the same perspective view as shown at each of 220, 225, 226 and 227 to the users, the asset 230 is rendered at each of the views by applying a perspective transformation based on a shared reference system (i.e., the reference system of user 223). Each of the views 220, 225, 226 and 227 is associated with a different user from the four users in the desktop review mode. It can be understood that when the four users are together in the XR environment and they have a shared perspective mode, each of them will have a view as shown at 220, 225, 227, 226 about the asset 230. In such a scenario, the transformation of the model to be rotated toward each respective user can be performed based on a transformation operation that facilitates the rotation of the model. For example, the center of the table can be used as the reference center (the (0,0,0) position of the XR environment). The rotation of the model can be considered a 2D problem because the upward direction remains constant during the rotation. Figure 4A and Figure 4B Further details for performing the transformation are provided.

[0062] With the shared view mode, all users can look at the shared assets in the XR environment from the same direction from their subjective perspective (to provide the same perspective view), while providing them with a spatial relationship relative to each other in the group that matches their real-world physical location. When users interact with shared content, they can point to shared assets. When a user points to an asset in shared view mode, other users can see the first user's pointing in a way that is similar to their pointing but offset in direction from their hand or gesture in the physical world to align with their relative position in the view viewed by the first user. Users in the XR environment are rendered based on technical details gathered for the users, including their gestures and interactions with content in the virtual world. These technical details are used to create a custom rendering of a representation of the XR world for each individual, which provides a shared view of the assets in the world while also allowing users to be viewed in the XR world, including presenting the user's avatar and showing their gestures.

[0063] Figure 2C Example viewing modes provided for display to multiple users accessing an XR environment through different devices viewing an asset 240 in a shared view mode in accordance with implementations of the present disclosure are shown.

[0064] In an auditorium, if multiple users are positioned next to each other, even if they are close, the best viewpoint is provided only to one person whose viewing angle matches the dead center on the 2D screen presenting the asset 240. In order to provide each user with the experience of viewing from the dead center of the screen while also maintaining the perception of viewing the screen from an auditorium with other users, the best viewpoint may be provided as described with respect to Figure 2A , Figure 2B The asset 240 may be rotated as discussed in connection with FIG. 4 , or the asset 240 may remain fixed for presentation while the viewpoint of the avatar of each of the users may be moved and aligned with a particular shared perspective. Figure 2C240. In the example of , the shared view may be the view that the user 245 in the center is experiencing, and each of the other users 246 and 247 next to the center user may have a reoriented viewpoint toward the asset 240. In this particular example, the user 244 may be a presenter, in which case the auditorium is used as a workshop where the user 244 is presenting material to a group of users 245, 246, and 247. When the three avatars 245, 246, and 247 in the XR environment as shown at scene 241 enter the shared view mode, the viewpoints of the avatars will be changed, and for example, the avatar of user 246 will have a changed viewpoint to match the viewpoint of avatar 245, as shown at scene 248. In scene 241, the viewpoint of avatar 245 is shared with avatars 246 and 247. For example, the viewpoint for sharing may be determined based on identifying a viewpoint that corresponds to a "best" viewpoint for asset 240. When the avatar is about to enter the shared mode, each of avatars 245, 246, and 247 is repositioned to the shared viewpoint (in this example, this is the middle seat that matches the position of avatar 245) to generate display output for their respective users. In this way, each user will see asset 240 from the same viewpoint without changing the position of asset 240. Additionally, at scene 248, avatar 244 is presented with a perspective transformed in the view of avatar 246, such as the user corresponding to avatar 246 and avatar 248 corresponding to the same user but with a changed position when entering the shared viewpoint mode of asset 240. Avatar 244, when displayed in the view of avatar 246 (in shared mode 248), will be rendered in a transformed manner as viewed by avatar 246 from the shared viewpoint. Since avatar 244 was oriented to look at avatar 246 at scene 241 (given the relative positions of avatar 246 and avatar 244 at scene 241 prior to entering shared perspective mode), avatar 244 is reoriented at scene 248 in shared mode to look at avatar 246. In this manner, each of avatars 246, 245, and 247 in its shared perspective view after entering shared mode is provided with a display of avatar 244 having a different view orientation relative to asset 240.

[0065] In this example, the viewpoint of the avatar of user 246 will be changed, as shown at scene 248, and this will be associated with a modification of the representation of the avatar of user 244 (the presenter) such that avatar 244 as an asset in the XR environment will be offset to align with the change in viewpoint of user 246 in their new viewpoint in scene 248.

[0066] It should be considered that moving the viewpoint of an avatar in a VR / AR experience may be associated with the movement of a user associated with the avatar or the creation of VR sickness. Therefore, movement around the position of the avatar may be performed within a certain threshold range of movement and reorientation, which will reduce the chance of creating such VR sickness. However, the movement of the user's camera in the XR environment is not entirely within the control of the system because the user will also move, however, changes in the user's position that may be made to provide a shared perspective view of the asset may be controlled within certain boundaries to balance between aligning the user's view and avoiding VR / motion sickness. In some implementations, the position of the avatar may be moved by performing a teleportation of the avatar from one location to another to improve the user's experience and reduce motion sickness that may occur if the avatar must be repositioned within a distance above the threshold range.

[0067] In some implementations, while the user is viewing in the XR environment and is part of a shared view mode, the user may turn their head (e.g., in 3D space) and / or move their position slightly. Such movement of the user may be detected as part of tracking the user's posture or based on receiving information about the user's tracked posture while remaining in the shared view mode. During the shared view viewing experience in the XR environment, the user may continue to change their posture (position and orientation) and the view rendering of the XR environment may be updated in real time or near real time, thereby avoiding VR sickness. In some cases, the position movement of the avatar within the XR environment may be large enough, for example, to be repositioned outside a threshold range around a position associated with the shared view mode. In those cases, the avatar may be removed from the shared view mode. For example, if the user's avatar wants to exit the shared view mode, the user may provide instructions for repositioning the avatar to a position outside a defined threshold range or perimeter to "jump out" of the shared mode. In some cases, if a user makes a change to their position and / or orientation while in shared view mode, and thereafter does not move for a period of time and still has their avatar in shared view mode along with other users' avatars, then based on detecting that the user has not moved, the user's rendered view may be readjusted, and the system may slowly pull the user's rendered view back to the shared view that was changed due to the change in position / orientation.

[0068] Figure 2D An example viewing mode of a user interacting with an XR environment 250 providing an option to join a shared view mode of an avatar pile is shown in accordance with implementations of the present disclosure.

[0069] The XR environment 250 may be accessible to multiple users to join. Users may join the environment through VR and / or AR devices and may be provided for viewing in different review modes (such as from Figure 2AThe XR environment 250 may be configured to support a shared view mode, which may be initiated when a user is viewing XR content in a particular review mode. For example, if an XR object is rendered in the XR environment, two users may be grouped to have a shared view of the XR object, where they can look at the object in desktop review mode.

[0070] In some implementations, groups of user avatars may have a shared view mode, where avatars may be grouped based on identifying that the distance between the locations of the avatars in the VR space is within a certain threshold. When the avatars are identified as being close to each other (e.g., within a threshold distance), the avatars may be prompted to opt into the shared view mode. In some implementations, the shared view mode may be generated to render a view for the user that presents the XR content from the same viewpoint. In some implementations, the XR content is rendered from the same viewpoint because the viewpoint of each avatar from the group is the same, and the user's rendered view is generated by applying a viewpoint transform to other avatars of other users, so that the other avatars (with the same viewpoint) are still visible within the views of the other users. Thus, the spatial relationship between avatars is maintained by providing an avatar stack that allows each user to experience a multi-user activity in which other avatars participate in the viewing.

[0071] exist Figure 2D , a user with a 3D avatar 260 named Shyam is navigating within an XR environment 250 and may be presented with views of other user avatars that are part of the XR environment 250. During interaction and / or navigation in the XR environment 250, groups of avatars may be formed, such as a group 255 that includes three users and is presented as an avatar stack of three avatars (users) collaborating in the group. Such a group 255 may be a group that shares the same perspective view when viewing an asset (such as an XR object) by repositioning the avatars' viewpoints so that the avatars enter the stack, while the users in the group 255 may see each other's avatars in a shifted manner (i.e., offset to align with the change in the viewpoint of the respective user due to the repositioning) in their respective views. The avatar 260 is provided as a view of the group 255 of the avatar stack object that includes multiple user avatars, so that the avatars 260 may view the grouped avatars without sharing their viewpoints. Thus, avatar 260 is provided with an external (or third party) view of the avatar stack, and avatar 260 can join the group by repositioning closer to avatar stack group 255 or by having the user of avatar 260 provide instructions for jumping to the location of group 255 and being provided with a shared view of the other avatars in group 255.

[0072] In some implementations, avatar 260 may be provided with a view of the position of the avatar stack, where the avatars may be shown as 3D avatars, and their positions and orientations may be aligned and offset to align with their perceived orientation toward shared content in the shared view group. In some implementations, avatar 260 may be provided with a display of avatar 256 of another user in the XR environment in his view. Avatar 256 is represented as a 3D avatar associated with the position and orientation in the XR environment (as a hemisphere including an icon and labeled with a name). In some implementations, avatar 260 and avatar 256 may enter a shared view mode (e.g., an avatar stack mode with respect to group 255), which may be associated with a shared reference frame in the XR environment, where both avatars will be repositioned to a shared viewpoint (e.g., by teleporting or otherwise repositioning) to enter the shared view mode. In the event that avatars 260 and 256 enter a shared perspective mode and the shared perspective mode is associated with a shared viewpoint, the view of the XR environment displayed for avatar 260 may be generated by applying a perspective transform based on a shared reference frame to avatar 256. In this manner, even though both avatars are associated with the same location in the XR environment, avatar 256 is displayed in the view of avatar 260. Avatar 256 rendered at the view of avatar 260 has an offset position and is reoriented to compensate for the adjustment caused by the offset.

[0073] In some implementations, in response to recognizing that avatar 260 moves or navigates (e.g., teleports, jumps, or otherwise) to a location in the XR environment that is within a threshold distance of an area occupied by the avatar in the shared view mode in the XR environment, it may be recognized that avatar 260 has entered (or provided an indication of a request to enter) a shared view mode with group 255. When avatar 260 joins the shared view mode, rendering of a new fourth user view may be performed on a display device of a user of avatar 260 by repositioning the viewpoint of avatar 260 to match a single reference frame associated with the shared view mode of group 255 and by applying a viewpoint transform to the avatar in the shared view mode based on the single reference frame, a reference frame established for avatar 260 (e.g., a reference frame associated with the avatar's viewpoint if the avatar has not been adjusted and moved to the shared viewpoint) and a corresponding reference frame established for the avatar in the shared view mode.

[0074] In some implementations, if an avatar from the group 255 initiates movement or navigation to another location in the XR environment that is outside of a threshold region defined for the shared view mode of the avatar stack 255, the avatar may be identified as leaving the shared view, and thus a new user view of the user associated with the avatar may be generated and rendered. The new user view may be generated by repositioning the viewpoint of the removed avatar to match the reference frame established for the avatar based on the user's device and the new location.

[0075] In some implementations, an avatar navigating within an XR environment and interacting with content may be transferred to an already defined avatar pile, and even if the avatar moves, the avatar does not have to leave the avatar pile and the shared perspective view. Movement may be allowed within a threshold distance or predefined area.

[0076] In some implementations, when avatars are identified as occupying the same space, they may be automatically stacked into an avatar stack, and then other avatars may join the group, or one or more of the avatars may leave the stack. In some implementations, stacking and unstacking may be performed based on the navigation of the avatars within the space and the identification of their positions and relative distances.

[0077] Figure 2E An example interaction of an avatar with a pointer to an asset rendered in a shared view mode according to an implementation of the present disclosure is shown. Figure 2E The example interactions presented may be performed in the context of rendering assets in shared view mode, where Figure 2B The method for applying a perspective transform to an asset.

[0078] For example, when one avatar of user 270 points with his hand to a location on a model of an object, another user in the group, user 275, may not understand (or correctly see, as it may be in a hidden location) what the location on the object is unless the second user has a perspective that includes the location (e.g., if the first avatar is pointing to the back of a building and the second user is looking at the front of the building, the second user will not see the location).

[0079] In some implementations, such perspective differences may not be desirable when the "best" viewpoint for presenting a model of an object is a particular viewpoint, and it may be impractical for all users to be physically located in the same place. For example, there may be a "best" seat in a movie theater, and only one user may have the "best" viewpoint when viewing content in an XR environment. To address such shortcomings and / or limitations when sharing content in an XR environment, a user may generate a video clip such as a video clip of a scene or scene. Figure 2E , where a group of users have a perspective view created by rotating the content (objects on a table) to match the user's perspective in 3D space. In some cases, the user's perspective can be moved relative to the content rendered in the XR environment, e.g. Figure 2C shown.

[0080] In some implementations, to provide a shared perspective view to users associated with viewing a given asset (e.g., XR content, such as an object, structure, or group thereof, or another user's avatar) when the user enters a shared perspective mode, a respective user view may be generated and rendered at the user's respective device. The respective user view may be generated based on a gesture tracked for the user, a reference frame of the asset to be shared, and a perspective transform applied to the asset.

[0081] In some implementations, content that can be shared in an XR environment includes assets that can be XR objects or user avatars. XR objects can include "computer models" (e.g., 3D models of objects) that can be rendered for viewing from different perspectives inside and / or outside the asset.

[0082] In some implementations, when user avatars are in a shared perspective view, the first user's view may include other user avatars that may be rendered in an offset position in the view, and their orientations may be adjusted to compensate for the offset in the position. In some implementations, during a collaborative activity between users with avatars sharing a shared perspective mode, the user may provide instructions for the avatar to interact with the rendered 3D content, such as instructions for pointing to a portion of the content. For example, when avatar 270 points to point 271 of an object with his hand (e.g., based on user-provided instructions for initiating pointing), avatar 275 may be provided with the same perspective view of the object, in which he may see avatar 270 pointing but with a representation 280 of a hand that is pointing to the location. In avatar 275's view, the hand of avatar 280 is offset to align with the viewpoint of avatar 275.

[0083] In some instances, an avatar of the user may be rendered based on the user's pose and perspective established for the XR environment. In some implementations, the shared content may be more than one asset and may include a combination of a 3D model of an object and / or an avatar.

[0084] Figure 2F An example user interface 290 of a collaboration application implementing a shared perspective mode according to implementations of the present disclosure is shown. The user interface 290 shows a third-person view of an avatar stack having five avatars, wherein each of the avatars in the stack may have a shared perspective view of an XR environment (shared perspective view not shown). In some implementations, a collaboration application providing a hybrid XR collaboration experience as discussed in the present disclosure may be implemented to provide a shared perspective view of multiple user avatars of an XR environment.

[0085] In some implementations, a shared view mode is implemented to provide a shared view of an XR environment by multiple avatars when the multiple avatars are placed in a pile. In some implementations, a first avatar of a first user is placed in an avatar pile in a shared view mode with a second avatar of a second user in response to the first avatar moving within a distance threshold of the second avatar of the second user in the XR environment. In some implementations, the movement of one avatar closer to or away from another avatar can be based on a request received from one or more users associated with the XR environment. In some cases, a first avatar in the XR environment can be smoothly moved to another second avatar (e.g., based on an instruction of a user associated with the first avatar). In some cases, the first avatar can be transferred to another avatar, where the transfer can be performed based on a selection of a location associated with the other avatar or based on a selection of the name of the other avatar (e.g., from a list of names of user avatars associated with the XR environment). For example, the transfer can be performed based on a direct instruction to the name of the avatar, which may not require a selection from a list, but can be a voice instruction or other gesture input. In some cases, the movement of the first avatar to another avatar may be based on a request to gather two or more avatars at a certain location in the XR environment. Whether the location of the avatars to be gathered may be defined as a location in the XR environment associated with a particular avatar, or may be defined as an input location defined by one of the avatars to be piled or by a master avatar defined for the XR environment. In some cases, piling the avatars into a location may be based on an instruction received from a user of the first avatar to be repositioned to another avatar (e.g., a request to perform an operation "Go to" the user, which may be provided as a user interactive element at the user interface 290). In some implementations, substantially similar movements based on different requests and / or instructions may be performed such that the user may move away from the avatar pile and thereby exit the shared view mode.

[0086] In the example user interface 290, avatar 295 (named "Shyam Prathish") may be an avatar of a first user, which may be moved to a position in a stack and clustered with four other avatars, as shown. In some implementations, in a shared view mode, each user view generated for each user associated with an avatar in the stack includes a graphical representation of each other avatar in the stack at the position in each respective user view. For example, in the view of avatar 295, graphical representations of the other avatars may be included, such as may be positioned at predefined positions in the user view. In some cases, the graphical representation may include an icon for the avatar, which may be an image or other graphic indicating the user. About Figure 2G More details about the first person view of the XR environment by the first avatar in the stack are presented. In some implementations, Figure 2F The user view of the avatar 295 may be as follows Figure 2G shown.

[0087] In some implementations, the first avatar of the first user may be moved from a current position to a position of the avatar pile at a rate slow enough to prevent motion sickness in the XR environment. In some implementations, the first avatar of the first user may be teleported from the current position of the first avatar to the position of the second avatar to join the pile. The teleportation may be in response to a command received from the first user. In some implementations, the avatar may be teleported to the position of the pile based on receiving a selection of the user's avatar from a list of avatars at a user interface component rendered at a user view of the second user by the second user. In this manner, the second user may select to teleport another avatar to add to the already stacked avatar. The second avatar may determine the avatar to be teleported from a group of avatars associated with the XR environment (e.g., avatars within the same experience, the same room, or other criteria), which may be identified by name, for example, presented at the user interface component and selected by the second avatar.

[0088] In some implementations, two or more avatars may be gathered into the avatar's location in response to a second user's user input identifying the two or more avatars to be gathered. For example, an avatar, such as avatar 295, may send instructions to gather four other avatars at his location. The other avatars may be selected by avatar 295, for example, based on name or based on presenting a list of icons associated with available avatars to join the shared view mode. In some implementations, one or more avatars may be dragged away from a current location associated with an avatar stack to change the shared view of the XR environment by the dragged avatars of the stack. In some instances, dragging the avatar may be based on the movement of the avatar of the stack designated as the owner.

[0089] In some implementations, in the shared view mode, the avatar of the active speaker in the pile is rendered with a three-dimensional model representation of the avatar of the active speaker at a predefined position in each respective user's view. For example, avatar 295 is rendered with a 3D model of his avatar to indicate that at this point in time, avatar 295 is the active speaker in a group of avatars in the pile. Avatar 295 is rendered with a 3D avatar representation to be seen from a third-person view of the pile. The 3D model of the avatar of the active speaker is rendered at the position of the pile in the user's view of other users associated with the avatars in the XR environment, the other users not in the shared view mode and the pile within their field of view. Within the view of the user of the avatar from within the pile, the avatar of the user who is the active speaker may also be shown with the 3D model to indicate to the avatars in the pile who is the active speaker.

[0090] Figure 2GAn example user interface 297 including a first-person view from a collaborative application implementing a shared perspective mode is shown in accordance with an implementation of the present disclosure. For example, the user interface 297 is from Figure 2F The user interface 297 includes a first-person view of a first user associated with a first avatar, the first-person view being a shared view within the XR environment by each of the five users associated with the five avatars in the XR environment within the stack. The user interface 297 includes an avatar for each of the avatars portion of the stack, wherein each avatar is presented with a distinguishable graphical representation (in this non-limiting example, this is a photograph). The user interface 297 may render the graphical representations of the avatars in the stack at a particular location in the view. For example, they may be arranged in a row in sequence or included in a graphical representation of a table. In some implementations, when a user speaks, the user's corresponding avatar may become visually modified, for example, enlarged, highlighted, dimmed, decorated, or otherwise recognizable as a graphical representation of a different style than other avatars. In some implementations, and as shown at the user interface 297, a user interface text component may be rendered in close proximity to the graphical representation of the avatar, wherein a message is provided to identify the name of the user who is speaking (in this example, this is Lucas Alves). In some implementations, and not shown in user interface 297, a 3D model of the active speaker's avatar may be rendered at user interface 297 to indicate who is the current active speaker, in addition to or in lieu of any other indication of other avatars shown with a graphical representation (e.g., an icon with a photo). In some implementations, within the user view of the active speaker, the avatars of the other users in the pile may be shown, and the user may be identified as the active speaker, for example, by including a visual text component of a message, or by including the name of the user who is the active speaker, or by adding a highlighted indication of the avatar representation of the user's avatar within the user view.

[0091] Figure 3 An example of a process 300 for providing a shared perspective view of assets in an XR environment is shown in accordance with implementations of the present disclosure.

[0092] At 310, a first user's view is rendered to a display device of the first user. The first user's view is of the XR environment, such as Figure 1 , Figure 2A , Figure 2B , Figure 2C , Figure 2DAs discussed. The XR environment may include a VR environment, or an AR environment, or a combination thereof. The view of the first user is generated using a first gesture tracked for the first user. For example, the first gesture includes a 3D position established for a user device or controller (e.g., a head mounted device and / or a VR controller) and movements (e.g., up and down, left and right, and rotation) that can be identified within the XR environment by tracking the movements of the user device or controller.

[0093] At 320, it is identified that a first avatar associated with a first user has entered a shared view mode with a second avatar associated with a second user of the XR environment. When entering the shared view mode, the avatar may enter a Figure 2B The shared desktop review mode shown may either enter an avatar stack where user avatars have their viewpoints reoriented to align with a single reference frame for viewing assets in the XR environment, or may enter a 1:1 scale collaborative review mode where user avatars are immersed in the VR experience with one avatar's viewpoint including a shared perspective of the asset and / or provided with a view of the avatar table reoriented for the respective avatar.

[0094] When the avatars enter shared view mode, they share a view of the asset (e.g., an XR object such as a 3D model or an individual's avatar in shared mode) that is reoriented for each of the users in the group. In the case where the avatars not only share a single reference frame for viewing objects in the XR environment but are also provided with reoriented views of the avatars of other users in the group, the positions of the other avatars are reoriented to adjust the presentation of the other avatars in view of the adjusted viewpoint of the respective avatar in the respective shared view. In this case, the offsets that need to be performed to generate the view of each individual avatar will be different because the users have different initial viewpoints that move to the common viewpoint (e.g., with respect to the avatars of the group). Figure 2C and Figure 5A The user avatar can be Figure 5A and Figure 5B The avatar gaze and gestures will mimic the actions performed by the user with the user's device (e.g., controller, or sensed by user sensors on a head mounted device or other device) as shown in the presentation, and for each user, the avatar gaze and gestures will mimic the actions performed by the user with the user's device (e.g., controller, or sensed by user sensors on a head mounted device or other device). The shared view mode has a single reference frame shared by the first user and the second user in the XR environment.

[0095] At 330, rendering to the display of the first user is performed while the users are in the shared view mode. By using the first gesture tracked for the first user and based on the single reference frame and for the first avatar (e.g., as shown in FIG. 1 ) when entering the shared view mode. Figure 2B ) or a second incarnation (e.g., Figure 5A and Figure 5BThe additional reference system established by (as shown) is applied to the perspective transformation of the asset to generate a first user view and perform rendering.

[0096] In some implementations, rendering to the first user's display can be performed in the context of presenting the asset as reoriented content (e.g., as Figure 2B In those cases, the additional reference frame used for the perspective transform on the asset corresponds to the position of the first avatar when the shared perspective mode is entered. The perspective transform is applied to the asset so that the asset can be viewed by the first user from a determined shared perspective associated with the single reference frame.

[0097] In some implementations, the first avatar and the second avatar may be part of an avatar stack in which the avatars have a shared viewpoint. For a rendered view of the first user, the avatar of the second user may be rendered as an asset in the view, where the second avatar may be rendered as a 3D avatar. When rendering the view for the first user, a perspective transform is applied to the 3D avatar so that the first user may see the second avatar of the second user, where both avatars are associated with the same viewpoint to the XR environment (stacked at the same point in the XR environment). In those instances, the additional reference frame corresponds to a predefined offset from the shared position of the avatar stack when the shared view mode is entered.

[0098] In some implementations, the first avatar and the second avatar may be provided with a shared view mode in which their views are repositioned, e.g. Figure 2C In those cases, the additional reference frame corresponds to at least the position of the first avatar when the first avatar entered the shared view mode, and the perspective transformation can be applied to reposition the viewpoint of the first avatar and can be applied to another avatar visible in the view of the first avatar (e.g., having a position visible from the shared viewpoint in the XR environment (e.g., from Figure 2C )) so that the orientation of the first avatar toward the other avatar can be relatively maintained to match the orientation of the first avatar toward the other avatar before entering the shared view mode.

[0099] According to implementations of the present disclosure, a shared view mode providing user avatars in an XR environment may be performed by reorienting content rendered on each user's display in the shared mode. In some instances, the reorientation may be performed on each user's identified content (such as an XR object shared with a group of avatars) in the shared mode. In some instances, the reorientation may be performed individually for each user and for the avatars of other users in the shared view mode, for example, as described with respect to Figure 5A , Figure 5B , Fig. 6A and Figure 6B described.

[0100] In some instances, an avatar stack may be presented in the user's view as a stack of 2D and 3D avatars to identify that the viewing user is performing the viewing in a group setting. The avatar stack may be considered an asset that may be rendered for each individual user's view, where reorientation must be applied differently for each avatar to offset the viewpoint of the avatar viewing the XR environment (e.g., in a 1:1 scale review mode or in a desktop review mode for 3D models, the view may include XR objects). Each user's avatar (except for the avatar of the viewing avatar) is shown in a stage generated for each individual and may be rendered as a visual representation showing movement and gestures (such as pointing with hands at other shared content in the XR environment). In this way, the avatars are presented as dynamic assets that may be updated based on dynamic changes in the user's interaction with the user interface and the user's device to change the interaction of their avatar in the XR environment.

[0101] In some instances, because presentation of a shared perspective mode can be accomplished using different methods, options for presentation can be provided as options for selection. For example, a user can select and initiate a shared mode where a group of users want to maintain a shared perspective mode while they work on a model and review the model from a desktop (e.g. Figure 2A and Figure 2B As shown, the user can view the model in a shared perspective mode, and when the user wants to navigate within the volume of the object, they can switch to a shared perspective mode that immerses their avatar in a mode where their viewpoint is changed and they can see other avatars in the avatar stack. There may be other options for switching or selecting between these two options, where the option may be dynamically selected based on the settings and location of the recognized user (e.g., when the user is in close proximity to meet distance criteria, the user may be immersed in avatar stack mode). By utilizing the avatar stack mode to provide a shared perspective mode, users can more easily adapt to different scenarios and angles for viewing XR objects, for example, moving from a desktop review mode to an immersive 1:1 scale review mode. When you want to view and work on a shared asset in desktop review mode, you may select a shared perspective mode in which the shared asset is reoriented (e.g., with respect to Figure 2A , Figure 2B and Figure 4A It should be understood that these two options can co-exist and be configured to be used based on user selection, dynamic determination, or other predetermined or default configurations for different users and use case scenarios.

[0102] Figure 4A An example of a process 400 for providing a shared view mode for shared viewing of XR objects in an XR environment is shown in accordance with implementations of the present disclosure.

[0103] At 410, a first user view of the XR environment is rendered to a first display device of a first user. For example, the first user may be provided with a desktop view of a model of an object, such as a Figure 2A The view of the first user is generated using a first gesture tracked for the first user. In some implementations, the first gesture as tracked includes a position and orientation of the first user relative to the XR environment.

[0104] At 420, a second avatar associated with a second user of the XR environment is identified as entering a shared perspective mode with a first avatar associated with the first user. A shared perspective may be defined for at least some content displayed in an XR environment (e.g., including a VR environment and / or an AR environment) while having a personalized perspective of at least one or more other avatars or other XR objects in the XR environment.

[0105] In some implementations, entering the shared view mode may be performed based on an active request from one or both of the first user and the second user, or may be initiated by a third party (e.g., another user in the XR environment, an administrator, or through a collaboration application of the users (such as Figure 1 The collaboration application 116 (an external service that configures and / or schedules collaboration sessions between users) is initiated. When it is recognized that the second avatar enters a shared view mode with the first avatar, the first avatar establishes a first reference frame in the XR environment.

[0106] At 430, a second user view of the assets of the XR environment is generated by applying a perspective transform to the assets of the XR environment based on a shared reference system of the shared perspective mode, a second reference system established for the second user in the XR environment, and a second posture tracked for the second user. In some cases, if the first user has a view to be shared, the shared reference system can be the first reference system. In some other cases, the shared reference system is different from the first reference system and can be a reference system of another user's perspective to be shared, including the reference system of the second user. In some instances, the shared reference system can be a reference system that none of the avatars in the shared perspective mode directly have. Therefore, when multiple avatars enter the shared perspective mode, a rendered view of the assets of the XR environment can be provided to all of them, and the rendered view was not a view of the assets by any of the avatars before entering the shared perspective mode.

[0107] At 440 , a second user view of the XR environment is rendered to a second display device of the second user and while in a shared view mode.

[0108] Figure 4B An example scheme 450 for performing perspective transformation when rendering objects in shared perspective mode is shown according to implementations of the present disclosure.

[0109] In some implementations, a shared perspective mode can be provided to two users to view a model of an object from a shared perspective, such as, for example, Figure 1 , Figure 2A , Figure 2B and Figure 4A The model of the object may be rendered in the XR environment in a desktop review mode, with two avatars, Avatar A and Avatar B, positioned around a table, as shown. The object may be rendered in the XR environment as object 455, and may be associated with a perspective view as a fixed scene, which may be associated with a fixed reference frame. For example, this perspective view as a fixed scene may be considered a "best" viewing perspective for object 455.

[0110] When avatars A and B enter the XR environment and opt into a shared perspective mode with respect to object 455, a view of the two users of avatar A and avatar B may be generated in which object 455 may be repositioned so that each of the users sees it from the same perspective. In some implementations, the view of the model of the object to the two avatars in shared perspective mode is generated based on applying a perspective transform to a fixed scene so that the model of the object is rotated and positioned for viewing by each of the two avatars from the same perspective as the perspective associated with the fixed reference frame.

[0111] In some implementations, a different reference frame is established for each of the users and used to determine how to rotate the model of the object, as long as the center of the table is the reference center (0,0,0) location in 3D space. Rotation of the model of the object to be rendered in both users' views presents a 2D problem because the rotation occurs on the x- and y-axes (the rotation occurs on the surface of the table in a 2D coordinate system), and the upward direction z remains constant.

[0112] In some implementations, a different reference frame can be established for each of the users in the XR environment by using position information from tracked postures of each user when entering shared view mode. In some cases, the reference frame can be referred to as a position and orientation considered in the XR environment, which can be defined relative to a coordinate system or transformation (in 3D graphics) defined for the XR environment. Therefore, the reference frame established for a user is the user's local coordinate system. Transformations applied based on such a coordinate system take into account differences in the position, rotation, and scale of the presentation of objects to different users. When rendering assets (such as XR objects) in the user's view, a perspective transform can be applied to provide a way to change between positions and orientations in the XR environment for different reference frames. Fixed scenes (such as Figure 4B The transformation applied (as shown) can be represented mathematically and relies on the principle of rotation to determine the corresponding position, rotation and scale of the object in the corresponding reference frame for a given user (for avatar A or avatar B).

[0113] In some implementations, each user may have their own table-centric (e.g., Figure 4B 455 ) and facing the user (e.g., in the X direction). The object may then be "attributed" or "attached" to the reference frame of the respective user, wherein when the model of the object is rendered in the user's view, the model of the object is rendered based on adjusting the orientation of the object to provide the same perspective viewpoint. The two users enter a shared perspective mode for the object in their own perspective reference frames. In some implementations, in response to one of the avatars manipulating the position of the model (e.g., changing the position, such as rotating a house to view its north wall), new relative coordinates of the model relative to the user's reference frame may be obtained and communicated to the other user (and / or multiple users, if the users have entered a shared perspective mode). Then, when a new view is rendered for another second user (after the position of the object is modified by the first user), the rendering is performed by generating a view in which the transformation is applied to the object based on relative coordinates with respect to the reference frame as determined for the first user and the shared reference frame of the fixed scene 455. The transformation is applied by transferring the modifications from the reference frame of the first user to the shared reference frame of the shared fixed scene and then reapplying them to the model of the object as viewed in the second reference frame for the second user. In this way, the second user is provided with a view of the model of the object that is rendered relative to the second user from the same angle (having the same viewing angle) as the first user in their view.

[0114] In some implementations, the view with the perspective transformation applied may be provided to the remaining users in a shared perspective mode. Rendering may be performed based on sharing of information about modifications to the model of the object by the respective users in their respective reference frames. Communicating new relative coordinates of the repositioned model of the object associated with the first reference frame associated with the first user by the first user to other users in a shared perspective view may flexibly support unification of views between users. Communication of relative coordinates may be provided to all users positioned around a table, where it may be appreciated that the table may take different shapes, including circular, rectangular, square, etc. In some implementations, sharing of relative coordinates of a model of an object as viewed by one user to other users in a shared perspective mode may be performed when the user is immersed in an XR environment and looking at the model of the object while the object is not positioned on a table or table but rather manipulated with the user's hands in virtual reality.

[0115] In some implementations, when two users enter a shared view mode, after rendering their views by applying a transformation from the fixed scene 455 to their reference frames, two avatars A and B may be presented with views of the model of the object with the same view, namely 490 and 495, respectively. In some implementations, the transformation may be performed when one or more of the two users performs an interaction other than modifying the model of the object, such as pointing to a portion of the model. For example, avatar A may point to a location on the model 490 of the object with a cursor 480. Such pointing may be transferred to the view of avatar B, where after applying a transformation to move from avatar A's reference frame to the reference frame of the fixed scene 455 and then to avatar B's reference frame, avatar B may be presented with a view of a pointer 485 that matches the cursor position and orientation of cursor 480.

[0116] In some cases, constantly aligning the view reference frame with the position of the users' eyes in the shared view when rendering views for the first user and the second user can limit the users' ability to change views and look around the object model to better understand its 3D structure. In some implementations, a fixed point in real space close to the user's physical position can be defined and used to align and rotate the model of the object toward that point. This determination of a point that is not exactly aligned with the eyes but still close to the user's position can support a better user experience while viewing the 3D model of an object in an XR environment in a shared view mode with other users.

[0117] In some implementations, when the model 490 of the object is rendered for avatar B, the rendering may need to account for the avatar's gaze and, optionally, the position at which the avatar is pointing. In some cases, there may be an "ideal" or "optimal" position defined for the model (typically, such a position is used to display the model in a workshop / presentation mode, such as with respect to Figure 2C Once such a position is determined, it can be used to define for the fixed scene 455, and then when rendering for different users, the model of the object can be rotated towards the other user based on the position of the other user (as determined based on the user device user entering the XR environment).

[0118] In some implementations, for each user entering the shared view mode, a VR origin is determined, which is a posture tracked by the user's device for the user, such as the location of a head mounted device when the user launches a cross-reality collaboration application to enter an XR environment. Figure 4BAs shown, the two users have VR origins 470 and 460, and those gestures are used to perform a 3D transform for the respective users' positions and orientations to determine how to represent a pointing performed by one avatar's hand (e.g., pointer 480 by avatar A's hand) to a point on a model of an object for avatar B (as shown, pointer 485 is a representation of pointer 480 rendered with a 3D transform to account for the difference in the VR origins of the two avatars). In some examples, a transform (e.g., a matrix multiplication) may be applied to the pointer to obtain the position and orientation of the pointer for avatar A's user and to obtain the perspective of the pointer in avatar B's user's view.

[0119] In some implementations, when assets are rendered in a user's view in shared perspective mode, those assets may be models or avatars of the object, and when they are presented from a shared perspective, modifications to the position and orientation of the assets are applied based on the position and orientation of another avatar being viewed. In some implementations, the XR object may be rotated to provide a shared perspective view, where avatars of other users using shared perspective mode may also be rendered, and those avatars may be rotated to support the user's understanding of the presence of other users in the XR environment. These other users may be provided as 3D avatars whose positions and orientations are offset to align with their presentation in another user's view. In this way, the assets have different positions and orientations for each user so as to provide the user with a common perspective on the model of the object. When the view includes both the model of the object and the avatar, the model may be maintained with a shared perspective view, while the avatars may be rendered from different perspectives of each of the other users based on their relative positions in the real world.

[0120] Figure 5A An example of a one-to-one scale collaborative review mode 500 of a user provided with an avatar stack perspective according to an implementation of the present disclosure is shown. In some implementations, the one-to-one scale collaborative review mode 500 may be substantially similar to Figure 2D , where in review mode 500, the view is the interior of a rendered XR model from an object. In this example, the one-to-one scale collaborative review mode 500 is the interior of a building (e.g., such as Figure 2A In some implementations, a group of users can interact with an object (such as a building, such as a Figure 2A ) and is provided as Figure 2A , Figure 2B , Figure 3 and Figure 4AA shared perspective view of the building in question. In some implementations, if a group of users begins interacting with a building and chooses to be provided a one-to-one scale collaborative review model at a particular location of the building (e.g., at a given floor or room of the building), the group of users can be immersed in a one-to-one scale collaborative review mode 500, which can still maintain a shared perspective mode for all users and also provide an avatar stack (stage 510 and list 550) to support the experience of shared collaborative tasks with other avatars interacting with shared content. In the case of switching from a shared perspective mode to a one-to-one scale collaborative review including avatars, the stack can be triggered when the users in the group are in close proximity to each other (e.g., the physical distance between each other does not exceed a threshold or occupies an area within a predefined perimeter, among other examples).

[0121] Figure 5A 5 shows how a user may be immersed in an XR experience from inside a model of an XR object (e.g., a building) in a one-to-one scale review mode, with an avatar stack presented as user interface elements in the form of a table 510 and a list 550. Table 510 is a user interface object that includes a 3D avatar positioned inside the table, and list 550 includes the names of other users (inside a group and not presenting 3D avatars) who are presented with 2D avatars (in some cases only name identifications (not shown)).

[0122] In some implementations, stage 510 may be configured to present a predefined number of avatars as 3D avatars before continuing to present avatars of other users (e.g., users who join the shared view mode at a later point in time, or there may be an order of avatars, such as a name order) in list 550. In the one-to-one scale review mode, the avatars' view is from inside the model of the XR object, where they can interact with the model of the object and have the same perspective view of the interior of the XR model of the object, while being provided with an avatar stack 510 that includes all users participating in the shared view.

[0123] Figure 5AIt is shown how a user under an avatar stack in a one-to-one scale collaborative review mode can communicate with other avatars of the group from the perspective of an XR model of a shared object (e.g., a building) via two-way screen interaction supported by providing interactive 2D data 530. Interactive 2D data 530 may be rendered within view 500 and may be rendered to each of the avatars within view 500. It should be understood that view 500 includes a view of the XR model of the object that appears the same to each of the avatars, while table 510 may appear slightly different to each avatar. The avatars in table 510 may be rendered with some offset because the user's view 500 allows each user to understand that other users are spatially far away from them even though they actually occupy the same position in the XR environment. In this way, in view 500, the orientation and position of the user's avatar from the physical world may be reflected in the way the avatar is presented in view 500. For example, an avatar of another user in view 500 may be rendered by applying a perspective transform (applied when generating view 500) that includes moving and reorienting the avatar based on a shared reference frame of the XR model of the object and a reference frame established for the other user.

[0124] The shared view mode of the avatar stack 500 may support interaction between users whose avatars are positioned near each other in the XR environment and share views of the model (e.g., such as Figure 2C The same viewpoint of the content inside the model is the same viewpoint of the content inside the model, where the view of the model is from the inside. However, it should be understood that the Figure 2C The same perspective view technology can be applied in Figure 5A Shown in the context of a view of the model from within.

[0125] As users immerse themselves in an interactive experience in avatar stack model 500, participating user avatars may be rendered as 3D and 2D avatars in stage 510. For example, avatar 520 of user Melissa C. may be marked (e.g., with a visual indication, with a color, or with a change in the type or style of rendering the user's name, among other examples) to indicate that this is the avatar in the avatar stack model 500. Figure 5A500. In some instances, station 510 may include avatars of users whose views are presented on the avatar stack 500. Additionally, station 510 may include some avatars that are 3D avatars, presented with bodies that represent the postures, gazes, and gestures of their respective users, as shown for Maddy Y, Melissa C, and Gerard S. Additionally, station 510 may include other avatars presented in list 550, the list including 2D avatars annotated with their respective user names. The 2D avatars may be images, logos, diagrams, or other 2D graphical representations, where the 2D avatars may be selected by each user through user interaction with user settings, or may be automatically assigned to each user upon entering a shared review mode of avatar stack 500. In some instances, station 510 may be defined for a set number of avatars to be presented as 3D avatars (e.g., three in this example), while other avatars added to the view may be presented as 2D avatars in the list, as shown at list 550.

[0126] In some implementations, the station 510 may be configured to render only one avatar as a 3D avatar that will correspond to the active speaker in the interactive sharing session. For example, when Melissa C. is detected as the active speaker, the station 510 may only include Melissa C.'s avatar 520 in the view of the user (one or more of the users in the interactive avatar stack mode 500). In some implementations, if another user (such as Gerard S.) begins speaking, the change in active speaker may be detected, and the station 510 may be rendered to include Gerard S.'s avatar as a single avatar presented as a 3D avatar.

[0127] In some implementations, when one or more avatars are presented as 3D avatars in table 510 in the user's avatar stack view 500, the 3D avatar representations are rendered in the user's view 500 by applying a perspective transform based on a single reference frame defined for a shared perspective mode of view 500, a reference frame established for the avatars of the user of view 500, and a reference frame established for the avatars rendered in the table of avatar stack view 500.

[0128] In some implementations, in the case of the user's view 500, Melissa C.'s 3D avatar 520 may have hands connected to the body (as shown), and a ray may be emitted from the user Melissa C's controller, where the ray may fall at a point in the view, for example, at the location 560 of the 2D interactive data 530. The orientation of Melissa C.'s 3D avatar may be transformed based on the user Melissa C.'s pose, Melissa C.'s reference frame as established by her device and / or controller, and the shared viewpoint of the avatar stack view 500. Melissa C.'s avatar 520 may be highlighted to indicate that Melissa C. is the active speaker at this time, where if another user becomes the active speaker, the highlighted avatar will be changed. In some instances, if another user associated with the 2D avatar is identified as the active speaker, the presentation of the 3D avatar in the station may be updated to include the user's avatar and render it as a 3D avatar. In some cases, other visual indications may be used instead of highlighting to indicate an avatar corresponding to a user identified as an active speaker. In some implementations, in order for a user to be identified as an active speaker, the user must be detected as an active speaker within a threshold period of time (e.g., 10 seconds), so that detection of such a speaker can be used as a trigger to provide an indication of the active speaker for a corresponding avatar, or to change an existing indication of an already identified active speaker.

[0129] Figure 5B An example of a process 570 for providing an avatar stack mode for shared perspective viewing of a model of an object in an XR environment is shown, wherein a table of user avatars is rendered, in accordance with implementations of the present disclosure.

[0130] Process 570 can be provided in the context of creating a group collaborative interaction, where content in the XR environment is seen by each user from a common viewpoint. Avatars of the user portions of the group can be associated with a common location in the XR environment to have a shared perspective view.

[0131] In some implementations, process 570 may be performed at an XR collaboration application, which may support a shared perspective view mode in which the XR content is moved or the viewpoints of the avatars are aligned to provide the same perspective view for each user in the group. In some implementations, the process 570 may be performed in the context of a desktop review mode of the XR content (e.g., such as Figure 2C as shown) or as Figure 5A Process 570 is performed in the context of an immersive one-to-one scale review mode as shown.

[0132] At 575, a first user view is rendered to a display device of the first user. The view is of the XR environment, for example, Figure 2C and Figure 5A The view of the first user is generated by using a first gesture tracked for the first user.

[0133] At 580, a first avatar of a first user is identified to have entered a shared view mode with a second avatar associated with a second user of the XR environment. The shared view mode has a single reference frame to be shared by the first user and the second user in the XR environment. In this way, the same view of a shared model of an object, such as, for example, about Figure 2C and Figure 5A As discussed above, the views are provided by changing the viewpoints of the two users to have a shared view of the model that matches a single reference frame as defined for the shared view mode. In some instances, the shared view can be the view that the first user had before entering the shared view mode, can be the view that the second user had before entering the shared view mode, or can be another view that can be dynamically determined, provided as input from a user or other application, or configured for the rendered model (e.g., an optimal view of the model can be set by default).

[0134] At 585, when in the shared view mode, the first user is provided with a view of the XR environment that is the same as the second user's view of the XR environment by generating a first user view and rendering avatars of both users (e.g., in a stage presented in front of the first user view, such as Figure 5A ), performing rendering to a display device of the first user. A first user view is generated by applying a perspective transformation to a second avatar of the second user based on a single reference system, a second reference system established for the second avatar, and a first reference system established for the first avatar when entering the shared view mode. In some implementations, based on the applied perspective transformation, the second avatar is rendered within the first user view at an orientation and a position that is offset from the orientation and position tracked for the second user when the second user is identified as entering the shared view mode to adjust the position of the second avatar. This offset is provided to maintain a visual experience of the spatial relationship between the two users just as in the physical world. Since their viewpoints are changed, but their physical positions are not changed, the way the user is oriented, where he / she is looking, or how he / she gestures (as tracked by a user device (such as a head-mounted device, a controller, a sensor (e.g., for movement, tracking eyes, etc.)) must be aligned with the change in perspective so that when the avatar is presented, it has the same relative orientation and approximate rotation (roll, pitch, and yaw) toward the content.

[0135] Fig. 6A An example user view 600 including an avatar stack mode with an active speaker is shown in accordance with implementations of the present disclosure.

[0136] In some implementations and as with respect to Figure 2C , Figure 5A and Figure 5B As discussed, a user may be immersed in a shared view mode provided with a stack of avatars, wherein the viewpoints of the avatars are moved so that they are aligned with a single reference frame, rather than manipulating content to render different views of the user.

[0137] In the first user view 600, XR content is rendered to the first user and an avatar of the first user is presented in front of the user view 600, i.e., avatar Ajay M.610. Avatar Ajay M.610 may be presented as a 3D avatar because Ajay is the active speaker in an interactive collaboration session between multiple users (nine (9) users, as indicated by the number of avatars and avatars 610 in list 620). In some instances, avatar Ajay M.610 may be the avatar of the first user whose view 600 is presented. In some other instances, avatar Ajay M.610 is the avatar of a second user who is identified as the active speaker when in a group collaboration activity in the XR environment. In some instances, tracking of the active speaker may be performed and the active speaker may be rendered with a 3D avatar in the station, as described above with respect to Figure 5A The platform is also presented in a substantially similar Figure 5A 560 is associated with other avatars in list 620. List 620 includes avatars of all other users in a shared view mode with the avatar of Ajay M 610, wherein the avatars in list 620 are presented as 2D avatars.

[0138] In some implementations, a shared perspective view of the XR content (i.e., inside the XR content of an XR building object) is provided to the first user by providing a table with avatars and a list of avatars (or other combinations thereof that allow the user to distinguish who is the active speaker and / or provide an offset orientation of the avatars to align with the first user's reference frame), while also including a visualization of the avatars in front of the first user's view 600 to allow understanding that other users are participating in the shared perspective viewing, even though all avatars have the same viewpoint (and can be viewed as looking from the same point, as if they were positioned one above the other).

[0139] In some implementations, the first user view 600 includes an avatar 610 identified as an active speaker, e.g., the first user is the active speaker, wherein the 3D avatar is positioned and oriented by applying an offset to the avatar to be included in the first user view 600 to align with a reference frame corresponding to the first user, with a reference frame shared for a shared view mode (matching a shared viewpoint), and with a reference frame of the avatar of the user identified as the active speaker (e.g., the same reference frame as the reference frame corresponding to the first user if this is the first user's view and the first user is the active speaker).

[0140] In some implementations, rendering of the shared view mode at the display device of another user from the user portion of the shared view mode with Ajay M. 610 may be performed by repositioning the viewpoint of the second avatar of the second user to match a single reference frame (defined for the shared view mode to look at the model from the inside) and by applying a viewpoint transform to the first user's avatar Ajay M. 610 in the shared view mode based on the single reference frame, the first reference frame established for the avatar 610, and the second reference frame established for the second avatar. If the first user is identified as speaking while in the shared view mode with the second user, the avatar of the first user may be rendered as a 3D avatar in the view of the other avatars. In some implementations, the 3D avatar is rendered in an orientation that may be different for each user in the group, wherein the 3D avatar has an orientation that is offset from an orientation and position tracked for the first user having a first pose associated with the first reference frame established for the avatar 610 to adjust the orientation and position of the avatar 610 to align with the single reference frame and also take into account the reference frame rotation relative to the particular user whose view the avatar is to be rendered.

[0141] Figure 6B An example of an avatar stack mode 650 providing representations of gestures and gazes of an active speaking avatar is shown in accordance with implementations of the present disclosure.

[0142] As about Fig. 6A As described above, users can be provided with Figure 6B The avatar stack mode 650 corresponds to an avatar stack mode, such as the avatar stack mode 600, and provides a shared perspective of the interior of a 3D model of an object to a group of users, while presenting them with a table and a list to provide a view of the avatars of the users participating in the shared perspective mode. As previously described, the avatar stack mode 650 can be rendered to the display device of the first user, including the rendered table 655, i.e., where the avatar 660 of Ajay M corresponding to the active speaker is presented. It should be understood that the mode 650 may be presented to the first user who is the active speaker and corresponds to the avatar Ajay M.660.

[0143] In some implementations, a first user viewing the avatar pile mode 650 on their display device may also be provided with a list 665 of other avatars in the shared view mode, such as a list of avatars in the shared view mode. Fig. 6A In some instances, 2D interactive data may be rendered in a shared perspective view, such as Figure 5A and Fig. 6A The first user may use his controller at 670 to point to the rendered 2D interactive data 667 in the shared perspective view, as previously described. The user of avatar Ajay M 660 interacts with the 2D interactive data 667 and may point to locations within the 2D interactive data 667. For example, each or some of the users in the avatar pile may be provided with the ability to update, modify, adjust, enhance, or delete data from the 2D interactive data 667.

[0144] In some implementations, the user of avatar Ajay 660 may point to a location where pointing may be performed using a controller device associated with the user's device or other sensors that can track the user's gestures. The pointing performed may be associated with a specific location at the 2D interactive data as shown in the user's view, where the location may be tracked as performed by the user of Ajay 660 and then transferred to the first user view including the avatar pile mode 650. The pointing performed by Ajay 660 may be transferred to a point 675 at the first user's 2D interactive data 667. The point 675 may be labeled with the name of the user performing the pointing, and the 3D avatar of Ajay 660 may be presented in the first user view at an offset position and orientation and pointing to the location 675 as shown in the view.

[0145] In some implementations, each user, such as a first user or a second user (Ajay), may have at least a first posture and a second posture for user tracking. The first posture may include a first position and a first orientation of a device corresponding to the user's head. The second posture may include a second position and a second orientation of another device or interactive tool corresponding to the user's pointer. When rendering the first user's first user view together with the content pointed to by the active speaker as the second user (Ajay), the rendering of pointing is performed by reorienting the graphical representation of the pointer of the second user in the XR environment based on the second user's second posture, the second reference system established for the second avatar, and the single reference system associated with the shared perspective mode (and corresponding to the shared viewpoint). In some implementations, the 3D representation of the avatar associated with the user performing the pointing is reoriented to include a reorientation of the presentation of the second avatar's hand and gaze based on the offset of the second avatar to be rendered in the stage 655 in front of the first user view. As shown, avatar 660 points with his hand to location 675, wherein the presented pose of avatar 660 is an adjusted pose of a pose established for the second user according to a single reference frame shared in the shared perspective view and a reference frame (or viewpoint) corresponding to the second user.

[0146] Figure 7 An example third person perspective 700 of an avatar stack is shown in accordance with implementations of the present disclosure.

[0147] In some implementations, the third person perspective 700 is a view mode provided to a third user to include an avatar stack defined in the XR environment that can be seen as a group of avatars at a certain location in the XR world without seeing the exact shared perspective view seen by each of the avatars in the stack. In some implementations, the avatar stack can be, for example, Figure 2C , Figure 2D , Figure 5A , Figure 5B , Fig. 6A and Figure 6B Defined as described.

[0148] For example, avatar piles 710 and avatar piles 720 may be defined in an XR environment, and when previewed in a user view of an external user, those avatar piles may be viewed as groups of users in a third-person perspective 700. The example avatar pile mode 700 is viewed by a third-person user who is not participating in either of the avatar piles 710 and 720. For example, the view may be substantially similar to Figure 2D 260. In the example avatar stack model 700, two groups are defined as a shared viewing mode of XR content in an XR environment. The two stacks can be rendered as a bunch of avatars one on top of the other (or other visual representations that allow avatars to be shown as part of a group, such as Figure 2D ). However, it will be appreciated that other visualizations of the avatar stack from a third person perspective may be provided.

[0149] In some implementations, a third person perspective of an avatar stack, such as the third person perspective of avatar stack 710, may include representations of the avatars as 2D avatars and / or may include representations of one or more of the avatars as 3D avatars that may be shown gesturing and pointing, e.g., by rendering the movement of the hands of the avatars in the stack as presented to an external user. However, a user viewing third person perspective 700 may not be provided with the same perspective view shared by users in avatar stack 710 or avatar stack 720. In some implementations, avatar stack 710 includes 3D avatars that indicate active speakers in a collaborative session of users associated with the avatars in the stack.

[0150] In some implementations, two avatar piles may be merged into one, e.g., by transferring an avatar from avatar pile 710 to avatar pile 720 or otherwise, based on a user instruction from a third person viewing perspective 700. In some implementations, a user associated with perspective 700 may configure groups of avatars and may define a shared view to be associated with each avatar pile. In some implementations, the view of a given avatar pile may be controlled by one or more of the users associated with the pile.

[0151] Figure 8 8 is a schematic diagram of a data processing system including a data processing device 800, which can be programmed as a client or a server. The data processing device 800 is connected to one or more computers 890 via a network 880. Figure 8Only one computer is shown as a data processing device 800, but multiple computers may be used. The data processing device 800 includes various software modules that may be distributed between the application layer and the operating system. These may include executable and / or interpretable software programs or libraries, including tools and services of an XR collaboration application 804 with a shared perspective mode, which includes a user interface that allows the display of XR content to users of the collaboration application 804. The view of the content may be rendered as a shared perspective view or an avatar stack, both of which allow users in a shared interactive session to be presented with the same perspective representation of the XR content in the XR environment. The collaboration application 804 may provide an audit mode, such as a desktop audit mode or a one-to-one collaborative audit mode, as described throughout this disclosure. In addition, the collaboration application 804 may implement teleconferencing functions, content sharing and editing, determining the location pointed to by other users within the shared content, and visualizing user avatars and identifying active speakers, as well as other functions as described throughout this disclosure. The number of software modules used may vary depending on the implementation. In addition, the software modules may be distributed on one or more data processing devices connected by one or more computer networks or other suitable communication networks.

[0152] The data processing device 800 also includes a hardware or firmware device, including one or more processors 812, one or more additional devices 814, a computer-readable medium 816, a communication interface 818, and one or more user interface devices 820. Each processor 812 is capable of processing instructions for execution within the data processing device 800. In some implementations, the processor 812 is a single-threaded or multi-threaded processor. Each processor 812 is capable of processing instructions stored on a computer-readable medium 816 or a storage device (such as one of the additional devices 814). The data processing device 800 uses a communication interface 818 to communicate with one or more computers 890, for example, via a network 880. Examples of user interface devices 820 include displays, cameras, speakers, microphones, tactile feedback devices, keyboards, mice, and VR and / or AR devices. The data processing device 800 may store instructions for implementing operations associated with the above-mentioned programs on, for example, a computer-readable medium 816 or one or more additional devices 814, such as one or more of a hard disk device, an optical disk device, a tape device, and a solid-state memory device.

[0153] The embodiments of the subject matter and functional operations described in this specification may be implemented in digital electronic circuits, or in computer software, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented using one or more computer program instruction modules encoded on a non-transitory computer-readable medium for execution by a data processing device or for controlling the operation of a data processing device. The computer-readable medium may be a manufactured product, such as a hard drive in a computer system or an optical disk sold through a retail channel, or an embedded system. The computer-readable medium may be obtained separately and later encoded with one or more computer program instruction modules, for example, after transmitting one or more computer program instruction modules through a wired or wireless network. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, or a combination of one or more of them.

[0154] The term "data processing apparatus" encompasses all devices, means and machines for processing data, including, for example, a programmable processor, a computer or a plurality of processors or computers. In addition to hardware, the apparatus may also include code that generates an execution environment for the computer program in question, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, a runtime environment, or a combination of one or more of them. In addition, the apparatus may employ a variety of different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.

[0155] Computer programs (also referred to as programs, software, software applications, scripts, or codes) may be written in any suitable form of programming language, including compiled or interpreted languages, declarative or procedural languages, and may be deployed in any suitable form, including as independent programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple collaborative files (e.g., files that store portions of one or more modules, subroutines, or codes). A computer program may be deployed to execute on one computer, or on multiple computers located at one location or distributed across multiple locations and interconnected by a communication network.

[0156] The processes and logic flows described in this specification may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and the device may also be implemented as, a special purpose logic circuit, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC).

[0157] Processors suitable for executing computer programs include, for example, both general-purpose and special-purpose microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. In general, a computer will also include one or more mass storage devices for storing data, such as a magnetic disk, a magneto-optical disk, or an optical disk, or operatively coupled to receive data from or transfer data to the one or more mass storage devices, or both. However, a computer does not need to have such a device. In addition, a computer may be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), to name a few. Devices suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and storage devices, including, for example, semiconductor memory devices, such as EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.

[0158] To provide interaction with a user, embodiments of the subject matter described in this specification may be implemented on a computer having: a display device, such as a liquid crystal display (LCD) device, an organic light emitting diode (OLED) display device, or another monitor for displaying information to a user; and a keyboard and pointing device, such as a mouse or trackball, by which a user can provide input to the computer. Other kinds of devices may also be used to provide interaction with a user; for example, the feedback provided to the user may be any suitable form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user in any suitable form may be received, including acoustic, voice, or tactile input.

[0159] A computing system may include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by a computer program that runs on a corresponding computer and has a client-server relationship with each other. The embodiments of the subject matter described in this specification may be implemented in a computing system, which includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a browser user interface that a user can use to interact with the implementation of the subject matter described in this specification), or any combination of such back-end, middleware or front-end components. The components of the system may be interconnected by digital data communication (e.g., a communication network) of any suitable form or medium. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., self-organizing peer-to-peer networks).

[0160] Although this specification contains many implementation details, these should not be interpreted as limitations on the scope of what is claimed or may be claimed, but rather as descriptions of features that are peculiar to a particular embodiment of the disclosed subject matter. Certain features described in this specification in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually in multiple embodiments or in any suitable sub-combination. In addition, although features may be described above as acting in certain combinations and even initially claimed for this purpose, one or more features from a claimed combination may be separated from the combination in some cases, and a claimed combination may involve a sub-combination or a variation of a sub-combination.

[0161] Similarly, although operations are depicted in a particular order in the accompanying drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in a continuous order, or that all of the operations shown be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can be generally integrated together in a single software product or packaged into multiple software products.

[0162] Thus, certain embodiments of the invention have been described. Other embodiments are within the scope of the following claims. Additionally, the actions recited in the claims can be performed in a different order and still achieve desirable results.

[0163] Example

[0164] Although the present application is defined in the appended claims, it will be appreciated that the invention may also (additionally or alternatively) be defined according to the following examples:

[0165] Shared 3D viewing of assets in an extended reality system

[0166] Example 1. A computer-implemented method comprising:

[0167] rendering a first user view of an extended reality (XR) environment to a display device of a first user, the first user view generated using a first gesture tracked for the first user;

[0168] recognizing that a first avatar associated with the first user has entered a shared view mode with a second avatar associated with a second user of the XR environment, wherein the shared view mode has a single frame of reference shared by the first user and the second user in the XR environment; and

[0169] When in the shared view mode, the rendering to the display device of the first user is performed by generating the first user view using the first posture tracked for the first user and a perspective transformation applied to an asset based on the single reference frame and an additional reference frame established for the first avatar or the second avatar when entering the shared view mode.

[0170] Example 2. The method of Example 1, wherein the asset is a set of content in the XR environment, the additional reference frame is established for the first avatar based on at least a position of the first avatar indicated by the first gesture when entering the shared view mode, and applying the view transform during generation of the first user view comprises:

[0171] The set of content in the XR environment is rotated to align the single reference frame with the additional reference frame established for the first avatar.

[0172] Example 3. The method of Example 2, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to a head of the user, the second pose comprising a second position and a second orientation of another physical object corresponding to a pointer of the user, and the method comprising when in the shared view mode:

[0173] The graphical representation of the pointer of the first user in the XR environment is re-orientated based on the second posture of the first user, the additional reference frame established for the first avatar, and the single reference frame to render the re-oriented graphical representation of the pointer of the first user at a display device of the second user providing a second-user view of the XR environment.

[0174] Example 4. A method as described in any of the preceding examples, wherein the asset is the second avatar, the additional reference frame is established for the second avatar using the shared perspective mode based at least on a predefined offset between the first avatar and the second avatar, and applying the perspective transform during the generation of the first user view comprises:

[0175] The second avatar is moved and re-orientated from the single reference frame to the additional reference frame.

[0176] Example 5. The method of Example 4, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to a head of the user, the second pose comprising a second position and a second orientation of another device corresponding to a pointer of the user, and the method comprising when in the shared view mode:

[0177] The graphical representation of the pointer of the second user in the XR environment is re-orientated based on the second posture of the second user, the additional reference frame established for the second avatar, and the single reference frame to render the re-oriented graphical representation of the pointer of the second user at the display device of the first user.

[0178] Example 6. The method of Example 4, comprising:

[0179] Rendering avatar stations, each avatar corresponding to a user among the users in the shared view mode; and

[0180] The first avatar in the station is rendered, wherein the first avatar is rendered as a 3D avatar in the XR environment based on identifying that the first user is an active speaker.

[0181] Example 7. The method of Example 4, comprising:

[0182] identifying that a third avatar of a third user is within a predefined distance threshold of the first avatar and the second avatar in the XR environment; and

[0183] The third avatar associated with the third user is placed in the shared view mode with the first avatar and the second avatar in response to the identifying.

[0184] Example 8. The method of Example 7, comprising:

[0185] identifying a fourth avatar of a fourth user entering the XR environment; and

[0186] A fourth user view is generated by using a fourth posture tracked for the fourth user and a perspective transformation applied to the first and second avatars in the shared perspective mode based on the single reference system and the reference system established for the fourth avatar when the fourth avatar enters the XR environment, and rendering is performed to a display device of the fourth user, wherein the fourth user view includes a representation of the avatar stand of the user in the shared perspective mode.

[0187] Similar operations and processes as described in Examples 1 to 8 may be performed in a system including at least one process and a memory, the memory being communicatively coupled to at least one processor, wherein the memory stores instructions that, when executed, cause at least one processor to perform operations. In addition, a non-transitory computer-readable medium may also be implemented that stores instructions that, when executed, cause at least one processor to perform operations as described in any one of Examples 1 to 8.

[0188] Shared perspective

[0189] Example 1. A computer-implemented method comprising:

[0190] rendering a first user view of an extended reality (XR) environment to a first display device of a first user, the first user view generated using a first gesture tracked for the first user;

[0191] recognizing that a second avatar associated with a second user of the XR environment enters a shared view mode with a first avatar associated with the first user;

[0192] generating a second user view of the asset of the XR environment by applying a perspective transform to the asset of the XR environment based on a shared reference frame for the shared perspective mode, a second reference frame established for the second user in the XR environment, and a second pose tracked for the second user; and

[0193] The second user's view of the XR environment is rendered to a second display of the second user while in the shared view mode.

[0194] Example 2. A method as described in Example 1, wherein the asset is a three-dimensional (3D) model of an object rendered in the shared perspective mode of the first user and the second user when viewing the XR environment, wherein the first user view and the second user view render the asset from the same angle toward the first posture of the first user and the second posture of the second user, respectively.

[0195] Example 3. A method as described in any of the preceding examples, wherein generating the second user view comprises:

[0196] The shared reference frame in the XR environment of the shared view mode to be shared with the first user and a third user viewing the XR environment is determined, the shared reference frame corresponding to the third user.

[0197] Example 4. A method as described in any of the preceding examples, wherein generating the second user view comprises:

[0198] The perspective transform is applied to the asset of the XR environment based on a first reference frame established as the shared reference frame of the shared perspective mode for the first avatar and a second pose tracked for the second user.

[0199] Example 5. The method of any of the preceding examples, wherein generating the second user view of the asset of the XR environment comprises:

[0200] The perspective transform is applied based on the second reference frame and the second pose tracked for the second user to rotate the asset of the XR environment to align with the shared reference frame.

[0201] Example 6. The method of Example 5, comprising:

[0202] While in the shared view mode, the rendering to the first user is performed by generating a new first user view using the first gesture tracked for the first user and a view transformation applied to the asset based on the second reference frame and a shared reference frame established for the first avatar of the first user when entering the shared view mode.

[0203] Example 7. The method of any of the preceding examples, wherein generating the second user view of the asset of the XR environment comprises:

[0204] A view of the first avatar of the first user in the XR environment is generated based on offsetting the 3D orientation of the first avatar according to the single reference frame for sharing the view of the asset of the XR environment.

[0205] Example 8. The method of any of the preceding examples further includes:

[0206] receiving an instruction from the first user to change or move the asset as rendered in the first user view;

[0207] In response to receiving the instruction, rendering a second user view of the second user by applying the perspective transformation to the asset as moved based on the shared reference system of the shared perspective mode and a second reference system established for the second user based on tracking the second posture of the second user when entering the shared mode.

[0208] Example 9. A method as in any of the preceding examples, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to the user's head, the second pose comprising a second position and a second orientation of another physical object corresponding to the user's pointer, and the method comprising when in the shared view mode:

[0209] The graphical representation of the pointer of the first user in the XR environment is re-orientated based on the second pose of the first user, a first reference frame established for the first avatar, and the shared reference frame.

[0210] Similar operations and processes as described in Examples 1 to 9 may be performed in a system including at least one process and a memory, the memory being communicatively coupled to at least one processor, wherein the memory stores instructions that, when executed, cause at least one processor to perform operations. In addition, a non-transitory computer-readable medium may also be implemented that stores instructions that, when executed, cause at least one processor to perform operations as described in any one of Examples 1 to 9.

[0211] Avatar stacks for 3D viewing of assets in a shared extended reality system

[0212] Example 1. A computer-implemented method comprising:

[0213] rendering a first user view of an extended reality (XR) environment to a display device of a first user, the first user view generated using a first gesture tracked for the first user;

[0214] recognizing that a first avatar associated with the first user has entered a shared view mode with a second avatar associated with a second user of the XR environment, wherein the shared view mode has a single reference frame to be shared by the first user and the second user in the XR environment and corresponds to a shared viewpoint of the first avatar and the second avatar of the second user; and

[0215] When in the shared perspective mode, the rendering to the display device of the first user is performed by generating the first user view by applying a perspective transformation to the second avatar of the second user based on the single reference system, the second reference system established for the second avatar and the first reference system established for the first avatar when entering the shared perspective mode.

[0216] Example 2. The method of example 1, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to a head of the user, the second pose comprising a second position and a second orientation of another device or interactive tool corresponding to a pointer of the user, and the method comprising when in the shared view mode:

[0217] The graphical representation of the pointer of the second user in the XR environment is re-orientated based on the second gesture of the second user, the second reference frame established for the second avatar, and the single reference frame.

[0218] Example 3. The method of Example 2, wherein the method comprises:

[0219] In response to recognizing that the first user is speaking while in the shared view mode, rendering the first avatar of the first user in the first user view as a three-dimensional (3D) avatar having an orientation and position based on the shared viewpoint offset, the orientation and the position defining a 3D position and 3D orientation in the XR environment.

[0220] Example 4. A method as described in any of the preceding examples, wherein the method comprises:

[0221] Rendering avatar stations to the display device of the first user, each avatar corresponding to one of the users in the shared view mode; and

[0222] Rendering the second avatar in the station as a 3D avatar with a 3D orientation in the XR environment to the display device of the first user based on identifying that the second user is the active speaker, and wherein the 3D position of the 3D avatar has an offset and a 3D orientation to be adjusted to the shared viewpoint.

[0223] Example 5. A method as described in any of the preceding examples, wherein the avatar station is rendered to include a 3D avatar and / or a 2D avatar, wherein when another avatar of another user is identified as an active speaker, the avatar of the other user can be updated to be rendered as a 2D avatar, and a 3D avatar can be rendered for the other user as the active speaker.

[0224] Example 6. A method as described in any of the preceding examples, wherein the method comprises:

[0225] while in the shared view mode, performing rendering to a display device of a second user the second user's view by generating a second user view of the XR environment based on the single reference system and the first reference system established for the first avatar and the second reference system established for the shared view mode for the second avatar by applying a view transform to the first avatar of the first user to match the single reference system;

[0226] In response to identifying that the first user is speaking while in the shared view mode, rendering the first avatar of the first user in the second user view as a 3D avatar having an orientation that is offset from the orientation and position tracked for the first user having the first posture to adjust the orientation and position to be aligned with the single reference frame.

[0227] Example 7. A method as in any of the preceding examples, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to the user's head, the second pose comprising a second position and a second orientation of another physical object corresponding to the user's pointer, and the method comprising while in the shared view mode:

[0228] The graphical representation of the pointer of the second user in the XR environment is re-orientated based on the second gesture of the second user and the 3D representation of the hand and gaze of the second avatar of the second user based on the offset of the second avatar rendered in a stage in front of the first user's view.

[0229] Example 8. A method as described in any of the preceding examples, wherein the method comprises:

[0230] identifying that a third avatar of a third user is within a predefined distance threshold of the first avatar and the second avatar in the XR environment; and

[0231] The third avatar associated with the third user is placed in the shared view mode with the first avatar and the second avatar in response to the identifying.

[0232] Example 9. The method of Example 8, wherein the method comprises:

[0233] identifying a fourth avatar of a fourth user entering the XR environment; and

[0234] A fourth user view is generated by using a fourth posture tracked for the fourth user and a perspective transformation applied to the first avatar and the second avatar as stacked in the shared perspective mode, and rendering is performed to a display device of the fourth user, wherein the first avatar and the second avatar are rendered in the fourth user view based on a reference system established for the fourth avatar when the fourth avatar enters the XR environment.

[0235] Example 10. The method of Example 9, wherein the method comprises:

[0236] In response to identifying that the fourth avatar moves to a position in the XR environment that is within a threshold distance from an area in the XR environment occupied by the first avatar and the second avatar in the shared view mode, identifying that the fourth avatar enters the shared view mode with the first avatar and the second avatar; and

[0237] When in the shared perspective mode, a new fourth user view is generated by repositioning the viewpoint of the fourth avatar of the fourth user to match the single reference system and by applying a perspective transformation to the avatar in the shared perspective mode based on the single reference system, a fourth reference system established for the fourth avatar and a corresponding reference system established for the avatar in the shared perspective mode, and rendering the new fourth user view to the display device of the fourth user is performed.

[0238] Example 11. A method as described in any of the preceding examples, wherein the method comprises:

[0239] In response to identifying that the first avatar moves to a position in the XR environment that is outside a threshold distance from an area in the XR environment occupied by the first avatar and the second avatar when the shared view mode was entered, identifying that the first avatar leaves the shared view mode with the second avatar; and

[0240] Generating a new first-user view by repositioning a viewpoint of the first avatar of the first user to match a reference frame established for the first avatar based on a first user device accessing the XR environment, and rendering the new first-user view to the display device of the first user is performed.

[0241] Similar operations and processes as described in Examples 1 to 11 may be performed in a system including at least one process and a memory, the memory being communicatively coupled to at least one processor, wherein the memory stores instructions that, when executed, cause at least one processor to perform operations. In addition, a non-transitory computer-readable medium may also be implemented that stores instructions that, when executed, cause at least one processor to perform operations as described in any one of Examples 1 to 11.

[0242] In and Out of the Heap

[0243] Example 1. A system comprising:

[0244] a non-transitory storage medium having stored thereon instructions for a collaborative application for an extended reality (XR) environment; and

[0245] One or more data processing devices, wherein the one or more data processing devices are configured to execute the instructions of the collaborative application so that the one or more data processing devices implement a shared perspective mode, in which a first avatar of a first user is placed in an avatar stack in the shared perspective mode with a second avatar of a second user in the XR environment in response to the first avatar moving within a distance threshold of the second avatar of the second user.

[0246] Example 2. A system as described in Example 1, wherein the instructions of the collaborative application implement the shared view mode by causing the one or more data processing devices to perform the following operations: moving the current position of the first avatar to the position of the avatar stack at a rate slow enough to prevent motion sickness in the XR environment.

[0247] Example 3. A system as described in any of the preceding examples, wherein the instructions of the collaborative application implement the shared perspective mode, in which each user view generated for each user associated with an avatar among the avatars in the stack includes a graphical representation of the position of each other avatar among the avatars in the stack in each corresponding user view.

[0248] Example 4. A system as described in Example 1, wherein the instructions of the collaborative application implement the shared perspective mode by causing the one or more data processing devices to perform the following operations: teleporting the first avatar from the current position of the first avatar to the position of the second avatar to join the pile, wherein the teleporting is in response to a command received from the first user.

[0249] Example 5. A system as described in any of the preceding examples, wherein the instructions of the collaborative application implement the shared view mode by causing the one or more data processing devices to perform the following operations: based on the second user's received selection of an avatar from a list of avatars rendered at a user interface component at the second user's user view, teleporting the avatar to the position of the second avatar to join the pile, a group of avatars associated with the XR environment and identified by name at the user interface component.

[0250] Example 6. A system as described in any of the preceding examples, wherein the instructions of the collaborative application implement the shared perspective mode by causing the one or more data processing devices to perform the following operations: in response to user input of the second user identifying two or more avatars to be gathered, clustering the two or more avatars into the position of the second avatar.

[0251] Example 7. The system of Example 6, wherein the user input requests to move the avatar to the location of the second avatar based on a username of an avatar.

[0252] Example 8. A system as described in any of the preceding examples, wherein the instructions of the collaboration application implement the shared view mode by causing the one or more data processing devices to perform the following operations: dragging the avatar in the stack away from a current position associated with the avatar stack to change the shared view of the XR environment of the avatar of the stack, wherein the dragging of the avatar is based on the movement of the avatar of the stack that is designated as the host

[0253] Example 9. A system as described in any of the preceding examples, wherein the instructions of the collaborative application implement the shared perspective mode, in which the avatar of the active speaker in the stack is rendered at a predefined position in each corresponding user view using a three-dimensional model representation of the avatar of the active speaker.

[0254] Example 10. A system as described in any of the preceding examples, wherein the instructions of the collaborative application implement the shared view mode, in which the avatar of the active speaker is rendered with a three-dimensional model of the avatar of the active speaker in the stack at the location of the stack in the user view of other users associated with the avatars in the XR environment, wherein the stack is within the field of view of each of the avatars of the other users.

[0255] Example 11. A system as described in any of the preceding examples, wherein the instructions of the collaborative application implement the first avatar's exit from the shared perspective mode by causing the one or more data processing devices to perform the following operations: allowing the first avatar to leave the avatar stack in response to the first user moving the first avatar a threshold distance away from the center position of the avatar stack in the XR environment.

Claims

1. A computer-implemented method comprising: rendering a first user view of an extended reality (XR) environment to a display device of a first user, the first user view generated using a first gesture tracked for the first user; recognizing that a first avatar associated with the first user has entered a shared view mode with a second avatar associated with a second user of the XR environment, wherein the shared view mode has a single frame of reference shared by the first user and the second user in the XR environment; as well as When in the shared view mode, the rendering to the display device of the first user is performed by generating the first user view using the first posture tracked for the first user and a perspective transformation applied to an asset based on the single reference frame and an additional reference frame established for the first avatar or the second avatar when entering the shared view mode.

2. The method of claim 1 , wherein the asset is a set of content in the XR environment, the additional reference frame is established for the first avatar based on at least a position of the first avatar indicated by the first pose when entering the shared view mode, and applying the view transform during generation of the first user view comprises: The set of content in the XR environment is rotated to align the single reference frame with the additional reference frame established for the first avatar.

3. The method of claim 2, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to the user's head, the second pose comprising a second position and a second orientation of another physical object corresponding to the user's pointer, and the method comprising when in the shared view mode: The graphical representation of the pointer of the first user in the XR environment is re-orientated based on the second posture of the first user, the additional reference frame established for the first avatar, and the single reference frame to render the re-oriented graphical representation of the pointer of the first user at a display device of the second user providing a second-user view of the XR environment.

4. The method of any one of the preceding claims, wherein the asset is the second avatar, the additional reference frame is established for the second avatar using the shared perspective mode based at least on a predefined offset between the first avatar and the second avatar, and applying the perspective transform during generation of the first user view comprises: The second avatar is moved and re-orientated from the single reference frame to the additional reference frame.

5. The method of claim 4, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to the user's head, the second pose comprising a second position and a second orientation of another device corresponding to the user's pointer, and the method comprising when in the shared view mode: The graphical representation of the pointer of the second user in the XR environment is re-orientated based on the second posture of the second user, the additional reference frame established for the second avatar, and the single reference frame to render the re-oriented graphical representation of the pointer of the second user at the display device of the first user.

6. The method of claim 4, comprising: Rendering avatar stations, each avatar corresponding to a user among the users in the shared view mode; as well as The first avatar in the station is rendered, wherein the first avatar is rendered as a 3D avatar in the XR environment based on identifying that the first user is an active speaker.

7. The method of claim 4, comprising: identifying that a third avatar of a third user is within a predefined distance threshold of the first avatar and the second avatar in the XR environment; as well as The third avatar associated with the third user is placed in the shared view mode with the first avatar and the second avatar in response to the identifying.

8. The method of claim 7, comprising: identifying a fourth avatar of a fourth user entering the XR environment; as well as A fourth user view is generated by using a fourth posture tracked for the fourth user and a perspective transformation applied to the first and second avatars in the shared perspective mode based on the single reference system and the reference system established for the fourth avatar when the fourth avatar enters the XR environment, and rendering is performed to a display device of the fourth user, wherein the fourth user view includes a representation of the user's avatar stand in the shared perspective mode.

9. A computer-implemented method comprising: rendering a first user view of an extended reality (XR) environment to a first display device of a first user, the first user view generated using a first gesture tracked for the first user; recognizing that a second avatar associated with a second user of the XR environment enters a shared view mode with a first avatar associated with the first user; generating a second user view of the asset of the XR environment by applying a view transform to the asset of the XR environment based on a shared reference frame for the shared view mode, a second reference frame established for the second user in the XR environment, and a second pose tracked for the second user; as well as The second user's view of the XR environment is rendered to a second display device of the second user while in the shared view mode.

10. A method as claimed in claim 9, wherein the asset is a three-dimensional (3D) model of an object rendered in the shared perspective mode of the first user and the second user when viewing the XR environment, wherein the first user view and the second user view render the asset from the same angle toward the first posture of the first user and the second posture of the second user, respectively.

11. The method of claim 9 or claim 10, wherein generating the second user view comprises: The shared reference frame in the XR environment of the shared view mode to be shared with the first user and a third user viewing the XR environment is determined, the shared reference frame corresponding to the third user.

12. The method of any one of claims 9 to 11, wherein generating the second user view comprises: The perspective transform is applied to the asset of the XR environment based on a first reference frame established as the shared reference frame of the shared perspective mode for the first avatar and a second pose tracked for the second user.

13. The method of any of claims 9-12, wherein generating the second user view of the asset of the XR environment comprises: The perspective transform is applied based on the second reference frame and the second pose tracked for the second user to rotate the asset of the XR environment to align with the shared reference frame.

14. The method of claim 13, comprising: While in the shared view mode, the rendering to the first user is performed by generating a new first user view using the first gesture tracked for the first user and a view transformation applied to the asset based on the second reference frame and a shared reference frame established for the first avatar of the first user when entering the shared view mode.

15. The method of any of claims 9-14, wherein generating the second user view of the asset of the XR environment comprises: A view of the first avatar of the first user in the XR environment is generated based on offsetting a 3D orientation of the first avatar according to the shared reference frame for sharing the view of the asset of the XR environment.

16. The method of any one of claims 9 to 15, further comprising: receiving an instruction from the first user to change or move the asset as rendered in the first user view; as well as In response to receiving the instruction, rendering a second user view of the second user by applying the perspective transformation to the asset as moved based on the shared reference system of the shared perspective mode and a second reference system established for the second user based on tracking the second posture of the second user when entering the shared mode.

17. The method of any of claims 9-16, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to the user's head, the second pose comprising a second position and a second orientation of another physical object corresponding to the user's pointer, and the method comprising when in the shared view mode: The graphical representation of the pointer of the first user in the XR environment is re-orientated based on the second pose of the first user, a first reference frame established for the first avatar, and the shared reference frame.

18. A computer-implemented method comprising: rendering a first user view of an extended reality (XR) environment to a display device of a first user, the first user view generated using a first gesture tracked for the first user; recognizing that a first avatar associated with the first user has entered a shared view mode with a second avatar associated with a second user of the XR environment, wherein the shared view mode has a single reference frame to be shared by the first user and the second user in the XR environment and corresponds to a shared viewpoint of the first avatar and the second avatar of the second user; as well as When in the shared perspective mode, the rendering to the display device of the first user is performed by generating the first user view by applying a perspective transformation to the second avatar of the second user based on the single reference system, the second reference system established for the second avatar and the first reference system established for the first avatar when entering the shared perspective mode.

19. The method of claim 18, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to the user's head, the second pose comprising a second position and a second orientation of another device or interactive tool corresponding to the user's pointer, and the method comprises when in the shared view mode: The graphical representation of the pointer of the second user in the XR environment is re-orientated based on the second gesture of the second user, the second reference frame established for the second avatar, and the single reference frame.

20. A method as claimed in claim 18 or claim 19, wherein the method comprises: In response to recognizing that the first user is speaking while in the shared view mode, rendering the first avatar of the first user in the first user view as a three-dimensional (3D) avatar having an orientation and position based on the shared viewpoint offset, the orientation and the position defining a 3D position and 3D orientation in the XR environment.

21. The method of any one of claims 18 to 20, wherein the method comprises: Rendering an avatar station to the display device of the first user, each avatar corresponding to a user among the users in the shared view mode; as well as Rendering the second avatar in the station as a 3D avatar with a 3D orientation in the XR environment to the display device of the first user based on identifying that the second user is the active speaker, and wherein the 3D position of the 3D avatar has an offset and a 3D orientation to be adjusted to the shared viewpoint.

22. The method of any one of claims 18-21, wherein the avatar station is rendered to include a 3D avatar and / or a 2D avatar, wherein when another avatar of another user is identified as an active speaker, the avatar of the other user can be updated to be rendered as a 2D avatar, and a 3D avatar can be rendered for the other user as the active speaker.

23. The method of any one of claims 18 to 22, wherein the method comprises: while in the shared view mode, performing rendering to a display device of a second user the second user view by generating a second user view of the XR environment based on the single reference system and the first reference system established for the first avatar and the second reference system established for the shared view mode for the second avatar by applying a view transform to the first avatar of the first user to match the single reference system; as well as In response to identifying that the first user is speaking while in the shared view mode, rendering the first avatar of the first user in the second user view as a 3D avatar having an orientation that is offset from the orientation and position tracked for the first user having the first posture to adjust the orientation and position to be aligned with the single reference frame.

24. The method of any of claims 18-23, wherein each user has at least a first pose and a second pose tracked for the user, the first pose comprising a first position and a first orientation of a device corresponding to the user's head, the second pose comprising a second position and a second orientation of another physical object corresponding to the user's pointer, and the method comprising when in the shared view mode: The graphical representation of the pointer of the second user in the XR environment is re-orientated based on the second gesture of the second user and the 3D representation of the hand and gaze of the second avatar of the second user based on the offset of the second avatar rendered in a stage in front of the first user's view.

25. The method of any one of claims 18 to 24, wherein the method comprises: identifying that a third avatar of a third user is within a predefined distance threshold of the first avatar and the second avatar in the XR environment; as well as The third avatar associated with the third user is placed in the shared view mode with the first avatar and the second avatar in response to the identifying.

26. The method of claim 25, wherein the method comprises: identifying a fourth avatar of a fourth user entering the XR environment; as well as A fourth user view is generated by using a fourth posture tracked for the fourth user and a perspective transformation applied to the first avatar and the second avatar as stacked in the shared perspective mode, and rendering is performed to a display device of the fourth user, wherein the first avatar and the second avatar are rendered in the fourth user view based on a reference system established for the fourth avatar when the fourth avatar enters the XR environment.

27. The method of claim 26, wherein the method comprises: in response to identifying that the fourth avatar moves to a position in the XR environment that is within a threshold distance from an area in the XR environment occupied by the first avatar and the second avatar in the shared view mode, identifying that the fourth avatar enters the shared view mode with the first avatar and the second avatar; as well as When in the shared perspective mode, a new fourth user view is generated by repositioning the viewpoint of the fourth avatar of the fourth user to match the single reference system and by applying a perspective transformation to the avatar in the shared perspective mode based on the single reference system, a fourth reference system established for the fourth avatar and a corresponding reference system established for the avatar in the shared perspective mode, and rendering the new fourth user view to the display device of the fourth user is performed.

28. The method of any one of claims 18 to 27, comprising: in response to identifying that the first avatar moves to a position in the XR environment that is outside a threshold distance from an area in the XR environment occupied by the first avatar and the second avatar when the shared view mode was entered, identifying that the first avatar leaves the shared view mode with the second avatar; as well as Generating a new first-user view by repositioning a viewpoint of the first avatar of the first user to match a reference frame established for the first avatar based on a first user device accessing the XR environment, and rendering the new first-user view to the display device of the first user is performed.

29. A system comprising: a non-transitory storage medium having stored thereon instructions for a collaborative application for an extended reality (XR) environment; as well as One or more data processing devices, wherein the one or more data processing devices are configured to execute the instructions of the collaborative application so that the one or more data processing devices implement a shared perspective mode, in which a first avatar of a first user is placed in an avatar stack in the shared perspective mode with a second avatar of a second user in the XR environment in response to the first avatar moving within a distance threshold of the second avatar of the second user.

30. The system of claim 29, wherein the instructions of the collaboration application implement the shared view mode by causing the one or more data processing devices to move the current position of the first avatar to the position of the avatar stack at a rate slow enough to prevent motion sickness in the XR environment.

31. A system as described in claim 29 or claim 30, wherein the instructions of the collaborative application implement the shared perspective mode, in which each user view generated for each user associated with an avatar among the avatars in the stack includes a graphical representation of the position of each other avatar among the avatars in the stack in each corresponding user view.

32. A system as described in any of claims 29-31, wherein the instructions of the collaborative application implement the shared perspective mode by causing the one or more data processing devices to perform the following operations: teleporting the first avatar from the current position of the first avatar to the position of the second avatar to join the pile, wherein the teleporting is in response to a command received from the first user.

33. A system as described in any of claims 29-32, wherein the instructions of the collaborative application implement the shared view mode by causing the one or more data processing devices to perform the following operations: based on the second user's received selection of an avatar from a list of avatars rendered at a user interface component at the second user's user view, the avatar is transferred to the position of the second avatar to join the pile, a group of avatars are associated with the XR environment and identified by name at the user interface component.

34. A system as described in any of claims 29-33, wherein the instructions of the collaborative application implement the shared perspective mode by causing the one or more data processing devices to perform the following operations: in response to user input of the second user identifying two or more avatars to be gathered, the two or more avatars are gathered into the position of the second avatar.

35. The system of claim 34, wherein the user input requests to move the avatar to the location of the second avatar based on a username of an avatar.

36. A system as described in any of claims 29-35, wherein the instructions of the collaborative application implement the shared perspective mode by causing the one or more data processing devices to perform the following operations: dragging the avatar in the stack away from the current position associated with the avatar stack to change the shared view of the XR environment of the avatar of the stack, wherein dragging the avatar is based on the movement of the avatar designated as the host of the stack.

37. A system as described in any of claims 29-36, wherein the instructions of the collaborative application implement the shared perspective mode, in which the avatar of the active speaker in the stack is rendered with a three-dimensional model representation of the avatar of the active speaker in the stack at a predefined position in each corresponding user view.

38. A system as described in any of claims 29-37, wherein the instructions of the collaborative application implement the shared view mode, in which the avatar of the active speaker is rendered with a three-dimensional model of the avatar of the active speaker in the stack at the location of the stack in the user view of other users associated with the avatars in the XR environment, wherein the stack is within the field of view of each of the avatars of the other users.

39. A system as described in any one of claims 29-38, wherein the instructions of the collaborative application implement the first avatar's exit from the shared perspective mode by causing the one or more data processing devices to perform the following operations: allowing the first avatar to leave the avatar stack in response to the first user moving the first avatar away from the center position of the avatar stack in the XR environment by a threshold distance.