Gaze-based coexistence system

By determining a target area based on user gaze and managing data streams through a central server, the computational costs and resource demands of transmitting avatar data in extended reality environments are reduced, optimizing bandwidth and power usage.

JP2026510760APending Publication Date: 2026-04-10APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
APPLE INC
Filing Date
2024-03-11
Publication Date
2026-04-10

Smart Images

  • Figure 2026510760000001_ABST
    Figure 2026510760000001_ABST
Patent Text Reader

Abstract

Techniques for transmitting data in a coexistence environment include initiating a virtual communication session between a local device and a remote device in a shared coexistence environment, with each of the multiple transmitting devices transmitting a transmit-quality data stream within the virtual communication session. The target area of ​​the local device, which includes a portion of the coexistence environment, is determined. The local device subscribes to a first-quality data stream for remote devices represented in the target area and a second-quality data stream for remote devices not represented in the target area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to image processing. More specifically, but not limited thereto, the present disclosure relates to techniques and systems for improving power and data usage when transmitting avatar data.

Background Art

[0002] Some devices can generate and present an extended reality (XR) environment. The XR environment can include a fully or partially simulated environment that people perceive and / or interact with via an electronic system. In XR, a subset of a person's body movements or their representations are tracked, and in response, one or more characteristics of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. Some XR environments enable multiple users to interact with each other within the XR environment. However, transmitting such avatar data can be computationally costly.

Brief Description of the Drawings

[0003] [Figure 1] A flowchart of a technique for transmitting avatar data to a receiving device according to one or more embodiments is shown.

[0004] [Figure 2] A flowchart of a technique for selectively transmitting avatar data of different qualities according to one or more embodiments is shown.

[0005] [Figure 3] A diagram of a technique for selectively requesting avatar data of different qualities according to one or more embodiments is shown.

[0006] [Figure 4] A flowchart of a technique for determining a region of interest according to one or more embodiments is shown.

[0007] [Figure 5] An exemplary network diagram, representing one or more embodiments, is shown in block diagram format.

[0008] [Figure 6] A mobile device according to one or more embodiments is shown in block diagram form. [Modes for carrying out the invention]

[0009] The embodiments described herein relate to techniques for generating and transmitting avatar data. Specifically, the embodiments described herein describe techniques for manipulating avatar data on a sender device for efficiency, and the avatar is provided to a server at different quality levels depending on the gaze of the user on the receiver device.

[0010] The techniques described herein relate to methods, systems, and computer-readable media for efficiently representing avatar data on a target area in a receiving device. Specifically, the techniques described herein include determining a target area in a receiving device based on the viewing direction of the user of the receiving device in a virtual communication session. One or more remote users represented within the target area are then identified. A first-quality data stream is requested for these remote devices represented within the target area of ​​the receiving device. For example, a data stream containing avatar data for users corresponding to the sending device may be requested at a high quality level. One or more remote devices that are active in the virtual communication session but not represented within the target area can be identified, and a lower-quality data stream can be requested for those sending devices. Thus, the receiving device can receive a full-quality data stream for avatar data of avatars viewed by the user of the receiving device, while avatar data for avatars outside the target area is generated using a lower-quality data stream.

[0011] The techniques described herein relate to various technical improvements. For example, the techniques described herein reduce the total data transmitted to a receiving device that transmits only the older data stream of avatar data represented in the target area. Furthermore, the power requirements on the receiving device are reduced. For example, by reducing the frame rate of avatar data received from some of the remote sender devices, the receiving device requires fewer resources to process the received image data. Since virtual communication sessions can contain a large number of users, the technical improvements are amplified. For example, by degrading the quality of data transmission for some of the device's hair, the overall bandwidth requirements are reduced.

[0012] The following description provides numerous specific details for illustrative purposes to enhance understanding of the disclosed concepts. As part of this description, some of the drawings in this disclosure represent structures and devices in block diagram form to avoid obscuring novel embodiments of the disclosed concepts. Also, for clarity, not all features of actual implementations are described herein. Furthermore, as part of this description, some of the drawings in this disclosure may be provided in flowchart form. Boxes in any particular flowchart may be presented in a specific order. However, it should be understood that any particular sequence in any flowchart is used only to illustrate one embodiment. In other embodiments, any of the various components shown in the flowchart may be omitted, or the illustrated sequence of operations may be performed in a different order or simultaneously. In addition, other embodiments may include additional steps not shown as part of the flowchart. Furthermore, the language used in this disclosure may be chosen primarily for readability and explanatory purposes, and not to limit or restrict the subject matter of the invention, and it is necessary to rely on the claims to determine such subject matter of the invention. In this disclosure, any reference to “one embodiment” means that a particular feature, structure, or characteristic described in relation to the embodiment is included in at least one embodiment of the disclosed subject matter, and any multiple references to “one embodiment” should not necessarily be understood as all referring to the same embodiment.

[0013] In actual implementation development (such as software and / or hardware development projects), developers must make numerous decisions to achieve their specific goals (e.g., compliance with system and business-related constraints), and these goals may vary depending on the implementation. Furthermore, while such development efforts can be complex and time-consuming, they are still routine work for those skilled in the art who are engaged in the design and implementation of multimodal processing systems that are of interest to this disclosure.

[0014] This document describes various embodiments of electronic systems and techniques for using such systems in relation to various technologies.

[0015] As used herein, the physical environment refers to the physical world that people can perceive and / or interact with without the aid of electronic devices. The physical environment may include physical features such as physical surfaces or physical objects. For example, the physical environment corresponds to a physical park, including physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses such as sight, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through electronic devices. For example, an XR environment may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, etc. In an XR system, a subset of a person's bodily movements or their representation is tracked, and in response, one or more properties of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. As an example, an XR system can detect a person's head movements and, in response, adjust the graphic content and sound field presented to that person in the same way that such views and sounds would change in the physical environment. As another example, an XR system can detect the movements of an electronic device presenting an XR environment (e.g., a mobile phone, tablet, laptop) and, in response, adjust the graphic content and sound field presented to that person in the same way that such views and sounds would change in the physical environment. In some situations (e.g., for accessibility reasons), an XR system can adjust the characteristics of the graphic content within the XR environment in response to a representation of bodily movement (e.g., a voice command).

[0016] The existence of a wide variety of electronic systems enables people to perceive and / or interact with various XR environments. Examples include head-mountable systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mountable system may have one or more speakers and an integrated opaque display. Alternatively, a head-mountable system may be configured to accept an external opaque display (e.g., a smartphone). A head-mountable system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mountable system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eye. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In some implementations, the transparent or translucent display may be configured to be selectively opaque. A projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface.

[0017] Figure 1 illustrates an exemplary technique for selectively transmitting data of different qualities within a virtual communication session, according to one or more embodiments. Specifically, the diagram in Figure 1 shows an exemplary virtual communication session 100 in which a receiver 105 interacts with three senders 110A, 110B, and 110C. Specifically, the different senders 110A, 110B, and 110C may be associated with devices that transmit avatar data representing the corresponding users. The avatar data may include, for example, data that allows the receiver device to reconstruct a visual representation of the user of the sender device. As an example, the avatar data may include data for generating face texture, alpha component, depth component, and / or pose data. In some embodiments, the avatar data is transmitted in the form of video image data, such as a series of image frames. It should be understood that each of these users may be associated with devices that transmit and receive data. For example, in a coexistence application with multiple devices, each device may encode an avatar representation and transmit it to each of the other devices (and thus also receive and decode avatar representations from each of the other devices). In some embodiments, the sender device 110 and the receiver device 105 may interact in an extended reality (XR) environment, such as a coexistence session or a communication session. For example, receiver 105 may be associated with a device that receives avatar data from devices corresponding to senders 110A, 110B, and 110C, and also transmits its own avatar data to each of the devices corresponding to senders 110A, 110B, and 110C. Similarly, devices corresponding to senders 110A, 110B, and 110C may receive avatar data from the device corresponding to receiver 105. However, for clarity, this description is given in relation to a specific receiver 105 in relation to a sender in a virtual communication session.

[0018] According to one or more embodiments, each sender is associated with a sender device. For example, sender 110A is associated with sender device 120A, sender 110B is associated with sender device 120B, and sender 110C is associated with sender device 120C. In some embodiments, each of sender devices 120A, 120B, and 120C can transmit sender data, such as avatar data, to receiver device 140. In some embodiments, receiver device 140 is a device associated with receiver 105.

[0019] In some embodiments, a central server 130 can be used to manage transmissions between a sender device 120 and a receiver device 140. For example, as presented, each sender device 120 can send sender data at a first quality level to the server 130. From there, the server 130 can determine the quality level at which to send the sender data to the receiver device 140. In this way, the sender device 120 does not need to decide at what quality level to send its own data to a particular receiver device. Rather, the server 130 determines and manages the various transmissions during a virtual communication session. Thus, the sender device 120 can send a single sender data stream at a single quality level.

[0020] In the current exemplary virtual communication session 100, receiver 105 is seeing sender 110B. Therefore, receiver device 140 (associated with receiver 105) can request sender B data 125 from sender 120B (associated with sender 120B) at a first quality level. For example, receiver device can request sender 110B's avatar data at a higher quality level for avatar data when the avatar is not in the target area, such as sender 110A and sender 110C. In some embodiments, the first quality level may be associated with the original quality level at which sender data 105 is generated by sender device 120. In contrast, since senders 110A and 110C are outside the target area in the virtual communication session 100, receiver device 140 can receive sender A data 125A from sender device 120A (associated with sender 110A) and sender C data 125C generated by sender device 120C (associated with sender 110C) at a reduced quality level. Thus, receiver device 140 is shown to have received a reduced frame rate for sender A data 135A and a reduced frame rate for sender C data 135C. However, receiver device 140 has received the full frame rate for sender B data 135B.

[0021] In some embodiments, the receiver device 140 can receive different data transmissions from the sender device based on a requested quality level or sender data. In some embodiments, the receiver device 140 can request a quality level for a given data transmission based on whether the user represented by the device is represented within a target area. The target area may be determined based on the gaze direction of a receiver, such as receiver 105. In some embodiments, the receiver device 140 can be enabled using eye-tracking technology, which includes one or more sensors capable of determining gaze vectors or other gaze information or eye-tracking data. Gaze information may be used to determine the target area. As an example, the target area may include a portion of the graphic representation of a virtual communication session in which the receiver's attention is focused. This can be determined based on gaze information or other contextual information (such as active content) in the virtual communication session. For example, the field of view or a portion of the field of view of receiver 105 can be determined based on gaze information and used as the target area. A determination can be made as to whether a particular sender is within the target area. For these senders within the target area, such as sender 110B, a request can be made to transmit avatar data or other data generated by the associated sender device at a first quality level. In contrast, if a particular sender is not within the target area, a request can be made for the avatar data or other data generated by the associated sender device to be transmitted at a second quality level, which may be lower than the first quality level.

[0022] In some embodiments, the requests may be sent directly to the sender device or to a central server 130 that can manage data transmissions. In some embodiments, the server 130 can receive sender data, such as sender A data 125A, sender B data 125B, and sender C data 125C, at a first quality level and be configured to generate a data stream for the recipient device 140 at the requested quality level. For example, in some embodiments, the frame rate of the sender data can be reduced by dropping frames from the sender data prior to transmission to the recipient device. In some embodiments, the technique for reducing the frame rate of the sender data can be enhanced by the sender device that pre-marks the frames of the sender data in a manner such that the server 130 can identify which frames should be dropped to reach a particular predetermined frame rate. For example, a virtual communication session may be associated with a set of default frame rates for quality levels, and the sender device 120 may pre-indicate to the server 130 which frames should be dropped and / or how to determine which frames should be dropped to reach each of the target quality levels for the virtual communication session.

[0023] FIG. 2 shows a flowchart of a technique for selectively transmitting avatar data of different qualities according to one or more embodiments. For purposes of explanation, the following steps are described in the context of FIG. 1. However, it should be understood that various actions may be taken by alternative components. Further, the various actions may be performed in different orders. Further, in accordance with various embodiments, some actions may be performed simultaneously, some may not be required, or other actions may be added.

[0024] Flowchart 200 begins at block 205, where a virtual communication session is initiated in a coexistence environment. As described above, a virtual communication session can be presented on a local device within an extended reality environment where multiple users communicate with a common virtual object from different physical devices. Thus, users can be located in the same location, remote locations, or any combination thereof. A virtual communication session may include content sharing between at least some users and other users. For example, users may share avatar data representing users on devices so that their virtual representations are presented in a common configuration across devices within the virtual communication session. In addition, a virtual communication session may include common virtual objects and applications that users can interact with or view in the virtual environment. In some embodiments, a virtual communication session may be presented in the form of an augmented reality environment, a virtual reality environment, or other extended reality environment.

[0025] Flowchart 200 proceeds to block 210 where a data stream is received from each device at a first quality level. According to some embodiments, a central server can receive data streams from each device at a first quality level. These data streams can include, for example, avatar data representing the user from whom the data stream is received. That is, the data stream can be used to convey to the receiving device how to render a graphical representation of the user associated with the device transmitting the data such that at least some characteristics of the user's appearance or reactions are conveyed by the avatar data. The data stream can additionally or alternatively include other data such as user interface content, media items, etc. In some embodiments, the data streams can be received from all devices at the same quality level or at different quality levels. For example, a user communicating in a virtual communication session from a device having substantial computing resources may be able to generate and stream data at a higher quality level than a device having limited resources.

[0026] As described above, in some embodiments, the data stream can be received in the form of a video data stream including a series of frames. In this embodiment, the initial quality level can be associated with a particular frame rate. This initial frame rate can be a global initial frame rate at which each sender is expected to transmit sender data. Alternatively, the initial frame rate can be device-specific, for example, when the original quality levels are different between devices.

[0027] Flowchart 200 proceeds to block 215, where a stream quality request for data streams from the remaining devices is received by one device. As described above, in embodiments of a coexistence environment, each device may be sending and receiving data. However, for clarity, this description describes the technique in relation to a specific device that receives data generated by other devices in a virtual communication session. As described below, a stream quality request may be based on whether the data of a particular device is represented within the area of ​​interest with respect to the requesting device. A stream quality request may be based on other parameters, such as whether the receiving device can handle a high-quality stream based on resource availability. A stream quality request can indicate the quality level at which the receiving device wants to receive the sender data stream. Therefore, a quality request may indicate the quality levels of multiple sender devices in a communication session, or multiple quality requests may be received by the server.

[0028] Flowchart 200 proceeds to block 220, where a data stream for the requesting device is generated at a certain quality level based on the area of ​​interest. As described above, the data stream can be generated for each device in the virtual communication session for which the data stream is requested. Therefore, the stream may include, for example, avatar data or other media data to be presented in the virtual communication session. In some embodiments, the request may be received for devices represented in the area of ​​interest with respect to the receiving device. Therefore, the request may be associated with a high-quality data stream request. According to some embodiments, the data stream may indicate how to generate the data stream at a reduced quality level for a particular data stream. As an example, the sender may send instructions for frames to be dropped in order to reach a specific quality level of one or more. This may be done, for example, by marking individual frames in the header file of the data stream. In some embodiments, the sender device may send data with instructions for one or more target quality levels. The target quality level may be a global target quality level, such as a default frame / second. Alternatively, the target quality level may be defined with respect to the original quality level of the transmission. For example, a reduced quality level could involve dropping every other frame from the original quality level, regardless of the original transmission's frame rate.

[0029] In some embodiments, the server can determine the remaining active devices in a virtual communication session where no request has been received. That is, if the receiving device does not specify a particular sending device identified in the area of ​​interest, the server can identify the remaining devices and decide to send a lower quality data stream from those devices. Thus, in block 225, the server determines whether each of the remaining devices is in the area of ​​interest of the requesting device. Then, in block 230, the server reduces the frame rate for transmissions from each of the remaining devices outside the area of ​​interest. This may be based on decisions made by the server, requests from the receiving device, etc.

[0030] Flowchart 200 ends at block 235, when the central server sends the stream to the requesting device. As described above, the stream may be specific to a particular sending device, and the central server may send the stream to the requesting device for each of the other devices (or one or more additional devices) active in the virtual communication session.

[0031] Figure 3 illustrates a technique for selectively requesting avatar data of different qualities according to one or more embodiments. According to one or more embodiments, the technique described in Figure 3 is performed by a receiving device. However, it should be understood that various functions may be performed by additional and / or alternative devices. Furthermore, various actions may be performed in different orders. Moreover, according to various embodiments, some actions may be performed simultaneously, some may not be necessary, or other actions may be added.

[0032] Flowchart 300 begins at block 305, where the location information of a local device (i.e., a receiver device) in the coexistence environment is determined. According to some embodiments, the local device may be associated with a specific set of locations or coordinates that the representation of the user of the local device is visible to other users in the shared coexistence environment. In some embodiments, the location information may be associated with coordinates in the virtual coexistence environment, or with other location information that can determine the relative position and / or orientation of the device among shared virtual items in the coexistence environment.

[0033] Flowchart 300 proceeds to block 310, where the area of ​​interest is determined for the local device within the coexisting environment. As described above, in some embodiments, the area of ​​interest may be determined based on the portion of the coexisting environment that the user's attention is focused on. The area of ​​interest may be determined based on, for example, the local user's gaze information, active content within the coexisting environment, etc. In some embodiments, the area of ​​interest may be associated with the field of view or a portion of the field of view of the coexisting environment from the viewpoint of the local device. The determination of the area of ​​interest will be described in more detail below with reference to Figure 4.

[0034] In block 315, flowchart 300 includes identifying the location of one or more remote devices in a coexisting environment with respect to the area of ​​interest. For the purposes of flowchart 300, the locations of remote devices are determined one at a time. However, it should be understood that the locations of remote devices can be identified simultaneously. According to one or more embodiments, the location of a remote device in a coexisting environment may be identified in order to determine whether a user representation of the remote device is present in the area of ​​interest. As another example, the location of a remote device in a coexisting environment may be identified based on the presentation of content provided by the remote device in the coexisting environment.

[0035] The flowchart proceeds to block 320, where a determination is made as to whether the current device is within the area of ​​interest. More specifically, a determination is made as to whether the content provided by the remote device is represented within the area of ​​interest from the perspective of the receiving device (i.e., the local device). According to some embodiments, determining whether the current device is within the area of ​​interest may include determining the relative prominence of the remote devices. For example, instead of defining the area of ​​interest as an area in space or environment, the area of ​​interest may be defined based on which remote devices are closer to the user's line of sight. That is, a first remote device that is closest to the user's line of sight may be considered to be within the first area of ​​interest, while a second remote device that is less conspicuous (i.e., not close to the user's line of sight / has low prominence) may be determined to be outside the area of ​​interest or within a second area associated with a different quality level than the area of ​​interest. As another example, a first remote device that is closer to the user (i.e., based on depth data) may be considered more prominent than a second remote device that is further away from the user, even if both are within the same area of ​​interest. Therefore, the relative prominence score can be determined for each of the multiple transmitting devices.

[0036] If the current device is within the target area, the flowchart proceeds to block 325, where a higher quality transmission of avatar data (or other sender device data) from a remote device is requested. As described above, a higher quality transmission may include a transmission of the data stream at its original quality. Additionally or alternatively, a higher quality transmission may be a default quality from a set of quality available to the local device. According to one or more embodiments, a higher quality transmission may be requested by a local device that is subscribed to a higher quality transmission for a remote device from a central server.

[0037] Similarly, returning to block 320, if it is determined that the current remote device is not within the target area, the flowchart proceeds to block 330, where a lower-quality transmission of avatar data (or other sender device data) from the remote device is requested. As described above, a lower-quality transmission may include data that has been downsampled or otherwise reduced by the central server. For example, a lower-quality data stream may include a reduced frame rate and can be generated by the central server by dropping a predetermined number of frames from the data stream provided by the current remote device. According to one or more embodiments, a lower-quality transmission may be requested by a local device that is subscribed to a lower-quality transmission for the remote device from the central server.

[0038] According to some embodiments, different quality levels may be associated with different data types. For example, higher quality transmissions and lower quality transmissions may be associated with different types of data. For instance, a higher quality transmission may include 3D image data, while a lower quality transmission may include 2D image data. As another example, in some embodiments, a higher quality transmission may include additional data types than a lower quality data stream. For example, a higher quality data stream may include both image data and audio data, while a lower quality data stream may include only audio data.

[0039] In some embodiments, the quality level selected for transmission may be based on multiple factors or signals. For example, the selected quality of transmission may be based on a combination (or weighted combination) of signals such as presence within the area of ​​interest (or other identified area), relative prominence between remote devices, and / or available transmission quality at the remote devices.

[0040] Flowchart 300 proceeds to block 335, where a determination is made as to whether additional remote devices are present in the coexistence environment. If so, the flowchart returns to block 315, where the location of the next remote device in the coexistence environment is determined. This process is carried out for each remote device until all remote devices have been considered. In some embodiments, the additional remote devices in 335 may include only remote devices within the field of view, or only remote devices otherwise known to the local device.

[0041] With respect to Figure 3, only higher quality and lower quality transmissions are described, but it should be understood that other embodiments can provide a number of other substitutions for transmissions of different qualities. For example, in some embodiments, a central server can receive two transmissions from a sender device, and one of the transmissions is configured to be reduced if required by the requesting device. The two received transmissions may be of the same quality or different quality. Furthermore, in some embodiments, the central server can be configured to reduce the quality of the data stream in a number of ways, thereby providing multiple quality levels of the data stream that can be requested by the receiving device.

[0042] According to one or more embodiments, the process described with respect to flowchart 300 may be repeated when a change in the spatial configuration of a local device and / or one or more remote devices in a coexisting environment is detected. For example, if the position and / or orientation of a local device changes, the area of ​​interest in the coexisting environment may also change. Therefore, the remote devices represented within the new area of ​​interest may be different. Thus, a local device can join different quality transmissions for various remote devices based on its relative spatial configuration to one or more remote devices.

[0043] Figure 4 shows a flowchart of a technique for determining the target area according to one or more embodiments. According to one or more embodiments, the technique described in Figure 4 is performed by a receiving device. However, it should be understood that various functions may be performed by additional and / or alternative devices. Furthermore, various actions may be performed in different orders. Moreover, according to various embodiments, some actions may be performed simultaneously, some may not be necessary, or other actions may be added.

[0044] Flowchart 400 shows a set of processes that can be performed to determine the area of ​​interest in the coexistence environment, as described above with respect to block 310 in Figure 3. Flowchart 400 begins in block 405, where eye-tracking data is received by the local device. According to one or more embodiments, the local device may include one or more sensors that can acquire eye-tracking data. This may include, for example, a user-facing camera and / or other sensors that can determine the user's gaze.

[0045] Flowchart 400 proceeds to block 410, where the eye-tracking data is aligned to the shared coexistence environment. According to one or more embodiments, the eye-tracking data may be aligned to the shared coexistence environment by determining a gaze vector that starts from the user of the local device and projects into the coexistence environment in the direction according to the user's determined gaze.

[0046] Flowchart 400 proceeds to block 415, where contextual information is identified within the coexisting environment. According to some embodiments, the area of ​​interest may take into account the direction of gaze, as well as other factors such as active content within the user environment. One example may be a user interface projected into an environment containing content in which the user is interacting.

[0047] Flowchart 400 ends at block 420, where the area of ​​interest within the coexisting environment is determined. The area of ​​interest may be based on the line of sight vector, contextual information, etc. The area of ​​interest may be a portion of the user's field of view of the remote device within the coexisting environment. The area of ​​interest may be of a set size or may be dynamic based on the context of the activity within the coexisting environment. For example, if there are multiple active areas in the coexisting environment, the area of ​​interest may be larger than when there is a single active area or when the user is interacting with or paying attention to a single active component in the environment.

[0048] According to one or more embodiments, two or more regions of a target can be identified. For example, the first region may be a target region based on eye-tracking data, and the second region may be within the field of view but outside the target of the eye-tracking data, for example, surrounding the first region of the target. Additional regions may be determined, for example, by periphery. Different regions may be associated with different quality levels. For example, a different quality level is required depending on which region the remote device is located in. Furthermore, in some embodiments, when the remote device is in a particular region, such as when the remote device is in a peripheral region or not visible, transmission may not be required at all.

[0049] Figure 5 shows a network diagram of a system that can implement various embodiments of the present disclosure. Specifically, Figure 5 shows an electronic device 500 which is a computer system having XR capabilities. The electronic device 500 may be part of a multifunction device such as a mobile phone, tablet computer, personal digital assistant, portable music / video player, wearable device, head-mounted system, projection-based system, base station, laptop computer, desktop computer, network device, or any other electronic system described herein. The electronic device 500 may be connected via a network 502 to additional electronic devices (one or more) 504 and / or other devices such as accessory electronic devices, mobile devices, tablet devices, desktop devices, or remote sensing devices.

[0050] Referring to Figure 5, a simplified block diagram of an electronic device 500 is shown, in which the electronic device 500 is communicably connected to one or more additional electronic devices 504 according to one or more embodiments of the present disclosure. Each of the electronic device 500 and the additional electronic devices 504 may be part of a multifunction device such as a mobile phone, tablet computer, personal digital assistant, portable music / video player, wearable device, base station, laptop computer, desktop computer, network device, or any other electronic device. The electronic device 500 may be connected to one or more additional electronic devices 504 and one or more network devices 510 via a network 502. Exemplary networks include, but are not limited to, local networks such as a universal serial bus (USB) network, an organization's local area network, and a wide area network such as the Internet. According to one or more embodiments, the electronic device 500 and the one or more additional electronic devices 504 may participate in a communication session in which each device can render avatars of users of other client devices.

[0051] Each of the electronic devices 500 may include a processor 530, such as a central processing unit (CPU). The processor 530 may be a system-on-a-chip, such as those found in mobile devices, and may include one or more dedicated graphics processing units (GPUs). Furthermore, the processor 530 may include multiple processors of the same or different types. The electronic device 500 may also include memory, such as memory 540. Each memory may include one or more different types of memory that can be used in conjunction with one or more processors, such as the processor 530, to perform device functions. For example, each memory may include a cache, ROM, RAM, or any kind of temporary or non-temporary computer-readable storage medium capable of storing computer-readable code. Each memory may store various programming modules for execution by the processor, including avatar modules 585 and / or other applications (one or more) 575. The electronic device 500 may also include storage devices, such as storage devices 550. Each storage device may include, for example, optical media such as magnetic disks and tapes (fixed, floppy, and removable), CD-ROMs and digital video disks (DVDs), and another non-temporary computer-readable medium, including semiconductor memory devices such as electrically programmable read-only memory (EPROM) and electrically erasable programmable read-only memory (EEPROM). The storage device may include data for generating an avatar and for use when participating in a coexisting environment, including registration data 555. The registration data 555 can be used to determine eye-tracking data, etc.Furthermore, the registration data 555 can be used to generate avatar data for the user of the electronic device 500.

[0052] The electronic device 500 may also include one or more cameras, such as camera 518, or other sensors, such as eye-tracking sensor 560. In one or more embodiments, each of the one or more cameras 518 may be a conventional RGB camera, depth camera, infrared camera, etc. Furthermore, each of the one or more cameras 518 may include a stereo camera system or other multi-camera system, time-of-flight camera system, etc., that captures images capable of determining depth information of the scene. Each of the electronic device 500 and the additional electronic device(s) 504 can enable the user to interact with the Extended Reality (XR) environment. The existence of a wide variety of electronic systems allows a person to perceive and / or interact with a variety of XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mounted system may have one or more speakers and an integrated opaque display. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing sound of the physical environment. A head-mounted system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes.Display devices 580 and 508 can utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light source, or any combination thereof. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In one embodiment, the transparent or translucent display may be configured to be selectively opaque. The projection-based system may employ retinal projection technology to project a graphical image onto the retina of a person. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface.

[0053] The additional electronic device(s) 504 may include components that enable the device to generate avatar data or other shared content in the coexistence environment. Thus, the additional electronic device(s) 504 may include an avatar module 506 for generating avatar data. In one or more embodiments, each of the additional electronic device(s) may include one or more encoders that can generate one or more data streams of avatar data or other content to be shared with the coexistence environment. In some embodiments, a single encoder may be used to generate a video data stream and transmit it to a central server such as a network device(s) 510, along with instructions or other directives to the network device(s) 510 regarding how to reduce the data stream to a lower quality. In some embodiments, the additional electronic device(s) may include one encoder for encoding the data stream to prepare it for reduction, and another encoder for encoding the data stream at its original quality level.

[0054] One or more network devices 510 may include, for example, a content management module 512 that manages data stream requests from receiver devices in a coexisting session. The content management module 512 can retrieve the correct data stream for a given sender device and, if necessary, reduce the data stream before sending it to the requesting device. In some embodiments, one or more network devices 510 may include one or more encoders for packaging the requested data stream(s) and sending them to the requesting device(s).

[0055] Referring now to Figure 6, a simplified functional block diagram of an exemplary multifunctional electronic device 600 according to one embodiment is shown. The electronic device may be a multifunctional electronic device, or it may have some or all of the components of a multifunctional electronic device described herein. The multifunctional electronic device 600 may include any combination of a processor 605, a display 610, a user interface 615, graphics hardware 620, device sensors 625 (e.g., proximity sensor / ambient light sensor, accelerometer, and / or gyroscope), a microphone 630, an audio codec 635, one or more speakers 640, a communication circuit 645, a digital image capture circuit 650 (e.g., including a camera system), a memory 660, a storage device 665, and a communication bus 670. The multifunctional electronic device 600 may be, for example, a mobile phone, a personal music player, a wearable device, a tablet computer, and the like.

[0056] The processor 605 can execute instructions necessary to perform or control the operation of numerous functions performed by device 600. The processor 605 can, for example, drive the display 610 and receive user input from the user interface 615. The user interface 615 can enable the user to interact with device 600. For example, the user interface 615 can take various forms such as buttons, keypads, dials, click wheels, keyboards, display screens, and touchscreens. The processor 605 may also be a system-on-a-chip, such as those found in mobile devices, and may include a dedicated GPU. The processor 605 may be based on a reduced instruction-set computer (RISC) or complex instruction-set computer (CISC) architecture or any other suitable architecture, and may include one or more processing cores. The graphics hardware 620 may be dedicated computing hardware for processing graphics and / or assisting the processor 605 in processing graphics information. In one embodiment, the graphics hardware 620 may include a programmable GPU.

[0057] The image capture circuit 650 may include one or more lens assemblies, such as lenses 680A and 680B. The lens assemblies may have various combinations of characteristics, such as different focal lengths. For example, lens assembly 680A may have a shorter focal length than lens assembly 680B. Each lens assembly may have separate associated sensor elements 690A and 690B. Alternatively, two or more lens assemblies may share a common sensor element. The image capture circuit 650 can capture still images, video images, enhanced images, and the like. The output from the image capture circuit 650 can be processed, at least in part, by a dedicated image processing unit or pipeline incorporated within a video codec(single or multiple) 655, processor 605, graphics hardware 620, and / or communication circuit 645. The images thus captured can be stored in memory 660 and / or storage device 665.

[0058] Memory 660 may include one or more different types of media used by the processor 605 and graphics hardware 620 to perform the functions of the device. For example, memory 660 may include a memory cache, read-only memory (ROM), and / or random-access memory (RAM). Storage device 665 may store media (e.g., audio files, image files, and video files), computer program instructions or software, preference information, device profile information, and any other suitable data. Storage device 665 may include one or more non-temporary computer-readable storage media, including, for example, magnetic disks and tapes (fixed, floppy, and removable), optical media such as CD-ROMs and DVDs, and semiconductor memory devices such as EPROMs and EEPROMs. Memory 660 and storage device 665 may be organized into one or more modules and used to tangibly hold computer program instructions or computer-readable code written in any desired computer programming language. For example, when executed by the processor 605, such computer program code may perform one or more of the methods described herein.

[0059] It should be understood that the above description is illustrative and not limiting. The material is presented in the content of specific embodiments so that a person skilled in the art can manufacture and use the disclosed subject matter as claimed, and variations of those embodiments will be readily apparent to a person skilled in the art (for example, some of the disclosed embodiments may be used in combination with one another). Accordingly, the specific configurations of steps or actions shown in Figures 2-4, or the configurations of elements shown in Figures 1 and 5-6, should not be construed as limiting the scope of the disclosed subject matter. Accordingly, the scope of the invention should be determined by referring to the appended claims, together with the entire scope of equivalents given to such claims. In the appended claims, the terms “including” and “in which” are used as plain English equivalents of the terms “comprising” and “wherein,” respectively.

Claims

1. A non-temporary computer-readable medium containing computer-readable code, wherein the computer-readable code is In a shared coexistence environment, a virtual communication session is established between a local device and multiple remote devices, wherein each of the multiple transmitting devices transmits a transmission-quality data stream in the virtual communication session. The target area of ​​the local device is determined, and the target area includes a part of the shared coexistence environment. At least one remote device among the plurality of remote devices, wherein the representation of at least one remote device among the plurality of remote devices is included in the target region, and identifies at least one remote device among the plurality of remote devices, In accordance with the identification, a first quality data stream is joined for at least the first remote device among the plurality of remote devices, Join a second quality data stream for at least a second remote device among the plurality of remote devices not identified within the target region, It can be executed by one or more processors, Non-temporary computer-readable media.

2. The computer-readable code for determining the target area of ​​the local device is The aforementioned local device acquires eye-tracking data, Based on the aforementioned eye-tracking data, the gaze direction of the shared coexistence environment is determined. Includes computer-readable code for, The area of ​​the target is determined based on the line of sight. The non-temporary computer-readable medium according to claim 1.

3. Retrieve updated eye-tracking data, Based on the updated eye-tracking data, the updated gaze direction is determined. Based on the updated line of sight direction, the updated region of the target is identified. The join is modified to the first quality data stream for the second remote device among the plurality of remote devices, in accordance with the fact that the second remote device among the plurality of remote devices is included in the updated region of interest. Further includes computer-readable code for, The non-temporary computer-readable medium according to claim 2.

4. The join is modified to the second quality data stream for the first remote device among the plurality of remote devices, in accordance with the fact that the first remote device among the plurality of remote devices is not included in the updated area in question. Further includes computer-readable code for, The non-temporary computer-readable medium according to claim 3.

5. The non-temporary computer-readable medium according to any one of claims 1 to 4, wherein the first quality data stream includes the transmission quality data stream.

6. A non-temporary computer-readable medium according to any one of claims 1 to 4, wherein the data stream of the first quality has a greater number of frames per second than the data stream of the second quality.

7. The non-temporary computer-readable medium according to any one of claims 1 to 4, wherein the transmission quality data stream includes avatar data of the corresponding transmitting device.

8. The computer-readable code for determining the target area of ​​the local device is Includes a computer-readable code for determining the relative prominence score for each of the plurality of transmitting devices, The target region is determined based on the relative prominence score. The non-temporary computer-readable medium according to claim 1.

9. A non-temporary computer-readable medium according to any one of claims 1 to 8, wherein the first quality data stream includes data types that are excluded from the second quality data stream.

10. A non-temporary computer-readable medium according to any one of claims 1 to 8, wherein the first quality data stream includes a first data type, and the second quality data stream includes a second data type.

11. A virtual communication session between a local device and multiple remote devices in a shared coexistence environment, wherein each of the multiple transmitting devices transmits a transmission-quality data stream in the virtual communication session, and the process of initiating the virtual communication session, Determining the target area of ​​the local device, which includes a portion of the shared coexistence environment, At least one remote device among the plurality of remote devices, wherein the representation of at least one remote device among the plurality of remote devices is included in the target region, and identifies at least one remote device among the plurality of remote devices. In accordance with the identification, join a first quality data stream for at least the first remote device among the plurality of remote devices, Joining a second quality data stream for at least a second remote device among the plurality of remote devices not identified within the target area, Methods that include...

12. Determining the target area of ​​the local device is The local device acquires eye-tracking data, Based on the aforementioned eye-tracking data, the direction of gaze in the shared coexistence environment is determined, Includes, The area of ​​the target is determined based on the line of sight. The method according to claim 11.

13. To obtain updated eye-tracking data, Based on the updated eye-tracking data, the updated gaze direction is determined, Based on the updated line of sight direction, the updated region of the target is identified, Modify the join to the first quality data stream for the second remote device among the plurality of remote devices, in accordance with the fact that the second remote device among the plurality of remote devices is included in the updated area of ​​interest, The method according to claim 12, further comprising:

14. The further includes modifying the join to the second quality data stream for the first remote device among the plurality of remote devices, in accordance with the fact that the first remote device among the plurality of remote devices is not included in the updated region in question. The method according to claim 13.

15. The method according to any one of claims 11 to 14, wherein the first quality data stream includes the transmission quality data stream.

16. The method according to any one of claims 11 to 14, wherein the first quality data stream has more frames per second than the second quality data stream.

17. The method according to any one of claims 11 to 14, wherein the transmission quality data stream includes avatar data of the corresponding transmitting device.

18. Determining the target area of ​​the local device is This includes determining the relative prominence score for each of the plurality of transmitting devices, The target region is determined based on the relative prominence score. The method according to claim 11.

19. The method according to any one of claims 11 to 18, wherein the first quality data stream includes data types that are excluded from the second quality data stream.

20. The method according to any one of claims 11 to 18, wherein the first quality data stream includes a first data type, and the second quality data stream includes a second data type.

21. It is a system, One or more processors, One or more computer-readable media containing computer-readable code, The computer-readable code is provided, In a shared coexistence environment, a virtual communication session is established between a local device and multiple remote devices, wherein each of the multiple transmitting devices transmits a transmission-quality data stream in the virtual communication session. The target area of ​​the local device is determined, and the target area includes a part of the shared coexistence environment. At least one remote device among the plurality of remote devices, wherein the representation of at least one remote device among the plurality of remote devices is included in the target region, and identifies at least one remote device among the plurality of remote devices, In accordance with the identification, a first quality data stream is joined for at least the first remote device among the plurality of remote devices, Join a second quality data stream for at least a second remote device among the plurality of remote devices not identified within the target region, Executable by one or more processors as described above, system.

22. The computer-readable code for determining the target area of ​​the local device is The aforementioned local device acquires eye-tracking data, Based on the aforementioned eye-tracking data, the gaze direction of the shared coexistence environment is determined. Includes computer-readable code for, The area of ​​the target is determined based on the line of sight. The system according to claim 21.

23. Retrieve updated eye-tracking data, Based on the updated eye-tracking data, the updated gaze direction is determined. Based on the updated line of sight direction, the updated region of the target is identified. The join is modified to the first quality data stream for the second remote device among the plurality of remote devices, in accordance with the fact that the second remote device among the plurality of remote devices is included in the updated region of interest. Further includes computer-readable code for, The system according to claim 22.

24. The join is modified to the second quality data stream for the first remote device among the plurality of remote devices, in accordance with the fact that the first remote device among the plurality of remote devices is not included in the updated area in question. Further includes computer-readable code for, The system according to claim 23.

25. The system according to any one of claims 21 to 24, wherein the first quality data stream includes the transmission quality data stream.

26. The system according to any one of claims 21 to 24, wherein the first quality data stream has a greater number of frames per second than the second quality data stream.

27. The system according to any one of claims 21 to 24, wherein the transmission quality data stream includes avatar data of the corresponding transmitting device.

28. The computer-readable code for determining the target area of ​​the local device is Includes a computer-readable code for determining the relative prominence score for each of the plurality of transmitting devices, The target region is determined based on the relative prominence score. The system according to claim 21.

29. The system according to any one of claims 21 to 28, wherein the first quality data stream includes data types that are excluded from the second quality data stream.

30. The system according to any one of claims 21 to 28, wherein the first quality data stream includes a first data type, and the second quality data stream includes a second data type.