Presenting communication data based on environment
By obtaining communication data associated with the second device, determining the environmental relationship between devices, and deciding whether to present video or augmented reality representation based on the environmental relationship, the problem of poor communication experience between devices in the shared environment is solved, and a better user experience and communication effect is achieved.
Patent Information
- Application Number
- CN202510339624.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-31
- Filing Date
- 2020-05-28
- Publication Date
- 2025-06-20
AI Technical Summary
Existing devices often fail to effectively process the environmental relationships between devices when presenting communication data, resulting in unwise representations of video or augmented reality in a shared environment, affecting the user experience.
By acquiring communication data associated with the second device, it is determined whether the first device and the second device are in a shared environment, if not in a shared environment, a video or augmented reality representation of the second person is presented, and if in a shared environment, a transparent transmission of the environment is presented and a video or augmented reality representation is abandoned.
Improves the user experience of communication between devices in a shared environment, avoids unnecessary video or augmented reality representations, reduces interference between network audio and direct audio, and improves the communication experience.
Smart Images

Figure CN120179075A_ABST
Abstract
Description
[0001] This application is a divisional application of the patent application for invention with the application number 202080029986.9 and the invention name of "Presenting Communication Data Based on the Environment", which entered the Chinese national phase on December 3, 2020, with the international application number PCT / US2020 / 034771 and the international filing date of May 28, 2020.
[0002] Cross - Reference to Related Applications
[0003] This application claims the benefit of U.S. Provisional Patent Application No. 62 / 855155, filed on May 31, 2019, which is hereby incorporated by reference in its entirety. Technical Field
[0004] The present disclosure generally relates to presenting communication data based on the environment. Background Art
[0005] Some devices are capable of generating and presenting augmented reality (AR) scenes. Some AR scenes include virtual scenes, which are simulated replacements of physical scenes. Some AR scenes include enhanced scenes, which are modified versions of physical scenes. Some devices for presenting AR scenes include mobile communication devices, such as smart phones, head-mounted displays (HMDs), glasses, head-up displays (HUDs), and optical projection systems. Most of the previously available devices for presenting AR scenes are ineffective in presenting communication data. Brief Description of the Drawings
[0006] Accordingly, the present disclosure can be understood by those of ordinary skill in the art, and a more detailed description can refer to some exemplary aspects of specific implementations, some of which are shown in the drawings.
[0007] Figures 1A to 1G is an illustration of an exemplary operating environment according to some specific implementations.
[0008] Figures 2A to 2D is a flowchart representation of a method for presenting communication data according to some specific implementations.
[0009] Figures 3A to 3C is a flowchart representation of a method for masking communication data according to some specific implementations.
[0010] Figure 4 is a block diagram of a device for presenting communication data according to some specific implementations.
[0011] In accordance with common practice, the various feature portions shown in the drawings may not be drawn to scale. Thus, for clarity, the dimensions of the various feature portions may be arbitrarily enlarged or reduced. Additionally, some of the drawings may not depict all of the components of a given system, method, or apparatus. Finally, throughout the specification and the drawings, like reference numerals may be used to denote like feature portions. Summary of the Invention
[0012] The various embodiments disclosed herein include devices, systems, and methods for presenting communication data. In various embodiments, a first device associated with a first person includes a display, a non-transitory memory, and one or more processors coupled to the display and the non-transitory memory. In some embodiments, a method includes obtaining communication data associated with a second device corresponding to a second person. In some embodiments, the method includes determining whether the first device and the second device are in a shared environment. In some embodiments, the method includes, in response to determining that the first device and the second device are not in a shared environment, displaying an enhanced reality (ER) representation of the second person based on the communication data associated with the second device.
[0013] The various embodiments disclosed herein include devices, systems, and methods for masking communication data. In various embodiments, a first device includes an output device, a non-transitory memory, and one or more processors coupled to the output device and the non-transitory memory. In some embodiments, a method includes, when the first device is in a communication session with a second device, obtaining communication data associated with the second device. In some embodiments, the method includes determining that the first device and the second device are in a shared physical setting. In some embodiments, the method includes masking a portion of the communication data so as to prevent the output device from outputting the portion of the communication data.
[0014] According to some embodiments, a device includes one or more processors, a non-transitory memory, and one or more programs. In some embodiments, the one or more programs are stored in the non-transitory memory and executed by the one or more processors. In some embodiments, the one or more programs include instructions for performing or causing to be performed any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of a device, cause the device to perform or cause to be performed any of the methods described herein. According to some embodiments, a device includes one or more processors, a non-transitory memory, and means for performing or causing to be performed any of the methods described herein. Detailed Description
[0015] Numerous details are described to provide a thorough understanding of the example embodiments shown in the figures. However, the figures merely illustrate some example aspects of the present disclosure and should not be considered limiting. Those of ordinary skill in the art will understand that other effective aspects and / or variations do not include all of the specific details described herein. Additionally, well-known systems, methods, components, devices, and circuits have not been described in detail so as not to obscure more relevant aspects of the example embodiments described herein.
[0016] Various examples of electronic systems and techniques for using such systems in connection with various augmented reality technologies are described.
[0017] A physical setting refers to the world that individuals can sense and / or interact with without using an electronic system. Physical settings such as a physical park include physical elements such as physical wildlife, physical trees, and physical plants. People can directly sense and / or otherwise interact with the physical setting using, for example, one or more senses (including vision, smell, touch, taste, and hearing).
[0018] In contrast to a physical setting, an augmented reality (AR) setting refers to a wholly (or partially) computer-generated setting that various individuals can sense and / or otherwise interact with by using an electronic system. In AR, the movement of a person is partially monitored, and in response thereto, at least one attribute corresponding to at least one virtual object in the AR setting is changed in a manner consistent with one or more physical laws. For example, in response to the AR system detecting that a person is looking up, the AR system can adjust the various audio and graphics presented to the person in a manner consistent with how such sounds and appearances would change in a physical setting. Adjustments to the attributes of virtual objects in the AR setting can also be made, for example, in response to a representation of movement (e.g., a voice command).
[0019] A person can utilize one or more senses, such as vision, smell, taste, touch, and hearing, to sense and / or interact with AR objects. For example, a person can sense and / or interact with an object that creates a multi-dimensional or spatial acoustic setting. A multi-dimensional or spatial acoustic setting provides an individual with the perception of discrete sound sources in a multi-dimensional space. Such objects can also implement acoustic transparency, which can selectively incorporate audio from the physical setting with or without computer-generated audio. In some AR settings, a person can sense and / or interact with only audio objects.
[0020] Virtual reality (VR) is an example of ER. A VR scene refers to an enhanced scene configured to include only computer-generated sensory inputs for one or more senses. A VR scene includes multiple virtual objects that a person can sense and / or interact with. A person can sense and / or interact with the virtual objects in a VR scene by simulating at least some of the movements in a person's actions within the computer-generated scene and / or by simulating the person or their presence within the computer-generated scene.
[0021] Mixed reality (MR) is another example of ER. An MR scene refers to an enhanced scene configured to integrate computer-generated sensory inputs (e.g., virtual objects) with sensory inputs from a physical scene or a representation of sensory inputs from a physical scene. On the reality spectrum, an MR scene lies between a fully physical scene at one end and a VR scene at the other end and does not include these scenes.
[0022] In some MR scenes, the computer-generated sensory inputs can be adjusted based on changes in the sensory inputs from the physical scene. Additionally, some electronic systems for presenting an MR scene can detect the position and / or orientation relative to the physical scene to enable interaction between real objects (i.e., physical elements from the physical scene or their representations) and virtual objects. For example, the system can detect movement and adjust the computer-generated sensory inputs accordingly, such that, for example, a virtual tree appears fixed relative to a physical structure.
[0023] Augmented reality (AR) is an example of MR. An AR scene refers to an enhanced scene in which one or more virtual objects are superimposed on a physical scene (or its representation). For example, an electronic system can include an opaque display and one or more imaging sensors for capturing video and / or images of the physical scene. For example, such video and / or images can be a representation of the physical scene. The video and / or images are combined with the virtual objects, and the combination is then displayed on the opaque display. The physical scene can be viewed indirectly by a person via the images and / or video of the physical scene. Thus, a person can observe the virtual objects superimposed on the physical scene. When the system captures an image of the physical scene and uses the captured image to display the AR scene on the opaque display, the displayed image is referred to as video pass-through. Alternatively, a transparent or semi-transparent display can be included in the electronic system for displaying the AR scene, such that an individual can directly view the physical scene through the transparent or semi-transparent display. The virtual objects can be displayed on the semi-transparent or transparent display, such that an individual observes the virtual objects superimposed on the physical scene. In another example, a projection system can be utilized to project virtual objects onto the physical scene. For example, the virtual objects can be projected onto a physical surface or as a hologram, such that an individual observes the virtual objects superimposed on the physical scene.
[0024] An AR scene may also refer to an augmented scene in which the representation of the physical scene is modified by computer-generated sensory data. For example, at least a portion of the representation of the physical scene can be graphically modified (e.g., magnified) such that the modified portion still represents the initially captured image (but not an exact copy). Alternatively, when providing video passthrough, one or more sensor images can be modified to impose a particular viewpoint different from the viewpoint captured by the image sensor. As another example, a portion of the representation of the physical scene can be changed by graphically blurring or removing the portion.
[0025] Augmented virtual (AV) is another example of MR. An AV scene refers to an augmented scene in which a virtual or computer-generated scene is combined with one or more sensory inputs from a physical scene. Such sensory inputs can include the representation of one or more features of the physical scene. A virtual object can, for example, be combined with a color associated with a physical element captured by an imaging sensor. Alternatively, a virtual object can adopt features consistent with, for example, the current weather conditions of the physical scene, such as weather conditions identified via imaging, online weather information, and / or weather-related sensors. As another example, an AR park can include virtual structures, plants, and trees, although animals within the AR park scene can include features accurately replicated from images of physical animals.
[0026] Various systems allow people to sense and / or interact with an ER scene. For example, a head-mounted system can include one or more speakers and an opaque display. As another example, an external display (e.g., a smartphone) can be incorporated into the head-mounted system. The head-mounted system can include a microphone for capturing the audio of the physical scene and / or an image sensor for capturing the image / video of the physical scene. A transparent or translucent display can also be included in the head-mounted system. The translucent or transparent display can, for example, include a substrate through which light (representing an image) is directed to a person's eyes. The display can also include LEDs, OLEDs, liquid crystal on silicon, laser scanning light sources, digital light projectors, or any combination thereof. The substrate through which light is transmitted can be an optical reflector, a holographic substrate, an optical waveguide, a light combiner, or any combination thereof. The transparent or translucent display can, for example, selectively transition between a transparent / translucent state and an opaque state. As another example, an electronic system can be a projection-based system. In a projection-based system, retinal projection can be used to project an image onto a person's retina. Alternatively, a projection-based system can also project virtual objects into the physical scene, for example, such as projecting a virtual object as a hologram or onto a physical surface. Other examples of ER systems include windows configured to display graphics, head-mounted earphones, earphones, speaker arrangements, lenses configured to display graphics, head-up displays, automobile windshields configured to display graphics, input mechanisms (e.g., controllers with or without haptic capabilities), desktop or laptop computers, tablets, or smartphones.
[0027] When a first device associated with a first person communicates with a second device associated with a second person, it may sometimes be unwise to present a video or ER representation of the second person on the first device. For example, if the first device and the second device are in the same environment, it may be unwise to present a video or ER representation of the second person on the first device because the first person can see the second person. In some scenarios, it may be helpful to indicate the type of the second device so that the first person knows how to interact with the second person. For example, if the second device provides a limited view of the surrounding environment of the first device to the second person, the first person does not need to point to an area of the surrounding environment that is not visible to the second person.
[0028] The present disclosure provides methods, devices, and / or systems that allow a first device to present communication data associated with a second device based on the presence state of the second device. When the first device obtains communication data associated with the second device, if the second device is not in the same environment as the first device, the first device presents a video or ER representation of the second person. If the second device is in the same environment as the first device, the first device presents a pass-through of the environment and forgoes presenting the video or ER representation encoded in the communication data. In some scenarios, when the second device includes an HMD, the first device presents an ER representation of the second person, and when the second device includes a non-HMD device (e.g., a handheld device such as a tablet or smartphone, a laptop computer, and / or a desktop computer), the first device presents a video.
[0029] When a first device associated with a first person communicates with a second device associated with a second person, presenting network audio and video may result in a degraded experience due to inaudible speech and imperceptible video. For example, if the first device and the second device are in the same physical setting, the first person will likely hear network audio through the first device and direct audio from the second person. Interference between the network audio and the direct audio may result in inaudible speech. Similarly, if the first device displays an ER representation of the second person when the second device is in the same physical setting as the first device, the first person may view the ER representation of the second person instead of the second person, resulting in a degraded communication experience.
[0030] The present disclosure provides methods, apparatuses, and / or systems for masking communication data when a first device and a second device are in the same physical setting. If the second device is in the same physical setting as the first device, the first device masks network audio to reduce interference between the network audio and the direct audio. Masking the network audio when the second device is in the same physical setting as the first device allows a first person to listen to the direct audio. If the second device is in the same physical setting as the first device, the first device masks the video or ER representation of a second person indicated by the communication data. Forgoing the display of the video or ER representation of the second person improves the user experience of the first person by allowing the first person to view the second person. In some scenarios, the first device presents a passthrough of the physical setting, and the first person sees the second person via the passthrough.
[0031] Figure 1A is a block diagram of an exemplary operating environment 1 according to some specific implementations. Although relevant features are shown, those of ordinary skill in the art will recognize from the present disclosure that, for the sake of brevity and in order not to obscure more relevant aspects of the exemplary specific implementations disclosed herein, various other features are not shown. To that end, as a non-limiting example, the operating environment 1 includes a first environment 10 (e.g., a first setting) and a second environment 40 (e.g., a second setting). In some specific implementations, the first environment 10 includes a first physical setting (e.g., a first physical environment), and the second environment 40 includes a second physical setting (e.g., a second physical environment). In some specific implementations, the first environment 10 includes a first ER setting, and the second environment 40 includes a second ER setting.
[0032] As Figure 1A shown, the first environment 10 includes a first person 12 associated with (e.g., operating) a first electronic device 14, and the second environment 40 includes a second person 42 associated with (e.g., operating) a second electronic device 44. In Figure 1A the example, the first person 12 is holding the first electronic device 14, and the second person 42 is holding the second electronic device 44. In various specific implementations, the first electronic device 14 and the second electronic device 44 include handheld devices (e.g., tablets, smartphones, or laptops). In some specific implementations, the first electronic device 14 and the second electronic device 44 are not head-mounted devices (e.g., non-HMD devices, such as handheld devices, desktop computers, or watches).
[0033] In various embodiments, the first electronic device 14 and the second electronic device 44 communicate with each other via a network 70 (e.g., a part of the Internet, a wide area network (WAN), a local area network (LAN), etc.). When the first electronic device 14 and the second electronic device 44 communicate with each other, the first electronic device 14 transmits first communication data 16, and the second electronic device 44 transmits second communication data 46. The first communication data 16 includes first audio data 18 captured by a microphone of the first electronic device 14 and first video data 20 captured by an image sensor (e.g., a camera, such as a front-facing camera) of the first electronic device 14. Similarly, the second communication data 46 includes second audio data 48 captured by a microphone of the second electronic device 44 and second video data 50 captured by an image sensor of the second electronic device 44.
[0034] The first electronic device 14 receives the second communication data 46, and the second electronic device 44 receives the first communication data 16. As Figure 1A shown, the first electronic device 14 presents the second communication data 46, and the second electronic device 44 presents the first communication data 16. For example, the first electronic device 14 outputs (e.g., plays) the second audio data 48 via a speaker of the first electronic device 14 and displays the second video data 50 on a display of the first electronic device 14. Similarly, the second electronic device 44 outputs the first audio data 18 via a speaker of the second electronic device 44 and displays the first video data 20 on a display of the second electronic device 44. In various embodiments, the first audio data 18 encodes a voice input provided by the first person 12, and the first video data 20 encodes a video stream of the first person 12. Similarly, the second audio data 48 encodes a voice input provided by the second person 42, and the second video data 50 encodes a video stream of the second person 42.
[0035] In various embodiments, the first electronic device 14 determines whether the second electronic device 42 is in the first environment 10. When the first electronic device 14 determines that the second electronic device 44 is in the first environment 10, the first electronic device 14 considers the second electronic device 44 to be local. In various embodiments, when the first electronic device 14 determines that the second electronic device 44 is local, the first electronic device 14 changes the presentation of the second communication data 46. In various embodiments, in response to determining that the second electronic device 44 is local, the first electronic device 14 masks a part of the second communication data 46. For example, in some embodiments, in response to determining that the second electronic device 44 is local, the first electronic device 14 abandons outputting the second audio data 48 and / or abandons displaying the second video data 50.
[0036] In Figure 1AIn the example, the first electronic device 14 determines that the second electronic device 44 is not in the first environment 10. When the first electronic device 14 determines that the second electronic device 44 is not in the first environment 10, the first electronic device 14 considers the second electronic device 44 to be remote. As Figure 1A shown, in response to determining that the second electronic device 44 is remote, the first electronic device 14 presents the second communication data 46. For example, in response to the second electronic device 44 being remote, the first electronic device 14 outputs the second audio data 48 and / or displays the second video data 50. In some specific implementations, in response to determining that the second electronic device 44 is remote, the first electronic device 14 abandons masking the second communication data 46.
[0037] See Figure 1B , the first head-mounted device (HMD) 24 worn by the first person 12 presents (e.g., displays) the first ER scene 26 according to various specific implementations. Although Figure 1B it is shown that the first person 12 holds the first HMD 24, in various specific implementations, the first person 12 wears the first HMD 24 on the head of the first person 12. In some specific implementations, the first HMD 24 includes an integrated display (e.g., a built-in display) that displays the first ER scene 26. In some specific implementations, the first HMD 24 includes a head-mounted housing. In various specific implementations, the head-mounted housing includes an attachment area to which another device having a display can be attached. For example, in some specific implementations, the first electronic device 14 can be attached to the head-mounted housing. In various specific implementations, the head-mounted housing is shaped to form a container for receiving another device (e.g., the first electronic device 14) including a display. For example, in some specific implementations, the first electronic device 14 slides / snaps into the head-mounted housing or is otherwise attached to the head-mounted housing. In some specific implementations, the display of the device attached to the head-mounted housing presents (e.g., displays) the first ER scene 26. In various specific implementations, examples of the first electronic device 14 include a smart phone, a tablet computer, a media player, a laptop computer, etc.
[0038] In Figure 1BIn the example of, the first ER scene 26 presents (e.g., displays) the second video data 50 within the card graphical user interface (GUI) element 28. In some specific implementations, the card GUI element 28 is within the similarity threshold of the card (e.g., the visual appearance of the card GUI element 28 is similar to the visual appearance of the card). In some specific implementations, the card GUI element 28 is referred to as a video card. In some specific implementations, the first HMD 24 presents the second video data 50 in the card GUI element 28 based on the type of the second electronic device 44. For example, in response to determining that the second electronic device 44 is a non-HMD (e.g., a handheld device such as a tablet, a smart phone, a laptop computer, a desktop computer, or a watch), the first HMD 24 presents the second video data 50 in the card GUI element 28. In some specific implementations, the first HMD 24 outputs (e.g., plays) the second audio data 48 via the speaker of the first HMD 24. In some specific implementations, the first HMD 24 spatializes the second audio data 48 to provide the appearance that the second audio data 48 originates from the card GUI element 28. In some specific implementations, the first HMD 24 changes the position of the card GUI element 28 within the first ER scene 26 based on the movement of the second person 42 and / or the second electronic device 44 within the second environment 40. For example, the first HMD 24 moves the card GUI element 28 in the same direction as the second person 42 and / or the second electronic device 44.
[0039] See Figure 1C , the second HMD 54 worn by the second person 42 presents (e.g., displays) the second ER scene 56 according to various specific implementations. Although Figure 1C it is shown that the second person 42 holds the second HMD 54, in various specific implementations, the second person 42 wears the second HMD 54 on the head of the second person 42. In some specific implementations, the second HMD 54 includes an integrated display (e.g., a built-in display) for displaying the second ER scene 56. In some specific implementations, the second HMD 54 includes a head-mounted housing. In various specific implementations, the head-mounted housing includes an attachment area to which another device having a display can be attached. For example, in some specific implementations, the second electronic device 44 can be attached to the head-mounted housing. In various specific implementations, the head-mounted housing is shaped to form a container for receiving another device (e.g., the second electronic device 44) including a display. For example, in some specific implementations, the second electronic device 44 slides / snaps into the head-mounted housing or is otherwise attached to the head-mounted housing. In some specific implementations, the display of the device attached to the head-mounted housing presents (e.g., displays) the second ER scene 56. In various specific implementations, examples of the second electronic device 44 include smart phones, tablets, media players, laptop computers, etc.
[0040] AsFigure 1C As shown, in some specific embodiments, the first communication data 16 includes first environmental data 22, and the second communication data 46 includes second environmental data 52. In some specific embodiments, the first environmental data 22 indicates various characteristics and / or features of the first environment 10, and the second environmental data 52 indicates various characteristics and / or features of the second environment 40. In some specific embodiments, the first environmental data 22 indicates physical elements located in the first environment 10, and the second environmental data 52 indicates physical elements allocated in the second environment 40. In some specific embodiments, the first environmental data 22 includes a grid map of the first environment 10, and the second environmental data 52 includes a grid map of the second environment 40. In some specific embodiments, the first environmental data 22 indicates the body posture of the first person 12, and the second environmental data 52 indicates the body posture of the second person 42. In some specific embodiments, the first environmental data 22 indicates the facial expression of the first person 12, and the second environmental data 52 indicates the facial expression of the second person 42. In some specific embodiments, the first environmental data 22 includes a grid map of the face of the first person 12, and the second environmental data 52 includes a grid map of the face of the second person 42.
[0041] In various specific embodiments, the first HMD 24 presents an ER object 30 representing the second person 42. In some specific embodiments, the first HMD 24 generates the ER object 30 based on the second environmental data 52. For example, in some specific embodiments, the second environmental data 52 encodes the ER object 30 representing the second person 42. In various specific embodiments, the ER object 30 includes an ER representation of the second person 42. For example, in some specific embodiments, the ER object 30 includes an avatar of the second person 42. In some specific embodiments, the second environmental data 52 indicates the body posture of the second person 42, and the posture of the ER object 30 is within a similarity to the body posture of the second person 42. In some specific embodiments, the second environmental data 52 indicates the physical facial expression of the second person 42, and the ER expression of the ER face of the ER object 30 is within a similarity to the physical facial expression of the second person 42. In various specific embodiments, the second environmental data 52 indicates the movement of the second person 42, and the ER object 30 mimics the movement of the second person 42.
[0042] In various embodiments, the second HMD 54 presents an ER object 60 representing the first person 12. In some embodiments, the second HMD 54 generates the ER object 60 based on the first environmental data 22. For example, in some embodiments, the first environmental data 22 encodes the ER object 60 representing the first person 12. In various embodiments, the ER object 60 includes an ER representation of the first person 12. For example, in some embodiments, the ER object 60 includes a head image of the first person 12. In some embodiments, the first environmental data 22 indicates the body posture of the first person 12, and the posture of the ER object 60 is within a similarity to the body posture of the first person 12 (e.g., within a similarity threshold of the body posture of the first person 12). In some embodiments, the first environmental data 22 indicates the physical facial expression of the first person 12, and the ER expression of the ER face of the ER object 60 is within a similarity to the physical facial expression of the first person 12 (e.g., within a similarity threshold of the physical facial expression of the first person 12). In various embodiments, the first environmental data 22 indicates the movement of the first person 12, and the ER object 60 mimics the movement of the first person 12.
[0043] In various embodiments, the first HMD 24 presents an ER object 30 based on the type of device associated with the second person 42. In some embodiments, the first HMD 24 generates and presents the ER object 30 in response to determining that the second person 42 is using an HMD rather than a non-HMD. In some embodiments, in response to obtaining second communication data 46 including second environmental data 52, the first HMD 24 determines that the second person 42 is using the second HMD 54. In some embodiments, the first HMD 24 spatializes the second audio data 48 to provide an appearance that the second audio data 48 originates from the ER object 30. Presenting the ER object 30 enhances the user experience of the first HMD 24. For example, presenting the ER object 30 provides an appearance that the second person 42 is in the first environment 10, even if the second person 42 is actually remote.
[0044] In Figure 1D the example, the first person 12 and the second person 42 are in the first environment 10. Accordingly, the first electronic device 14 and the second electronic device 44 are in the first environment 10. In Figure 1D the example, the first electronic device 14 determines that the second electronic device 44 is in the first environment 10. When the first electronic device 14 and the second electronic device 44 are in the same environment, the environment can be referred to as a shared environment (e.g., a shared physical set). When the first electronic device 14 detects that the second electronic device 44 is in the same environment as the first electronic device 14, the first electronic device 14 considers the second electronic device 44 to be local (e.g., rather than remote).
[0045] In various embodiments, in response to determining that the second electronic device 44 is local, the first electronic device 14 masks a portion of the second communication data 46. For example, in some embodiments, the first electronic device 14 forgoes displaying the second video data 50 on the display of the first electronic device 14. Since the first electronic device 14 does not display the second video data 50, the first electronic device 14 allows the first person 12 to view the second person 42 without being distracted by the presentation of the second video data 50. In some embodiments, the first electronic device 14 forgoes playing the second audio data 48 through the speakers of the first electronic device 14. In some embodiments, not playing the second audio data 48 allows the first person 12 to hear the voice 80 of the second person 42. Since the first electronic device 14 does not play the second audio data 48, the second audio data 48 does not interfere with the voice 80 of the second person 42, thus allowing the first person 12 to hear the voice 80. Similarly, in some embodiments, in response to determining that the first electronic device 14 is local, the second electronic device 44 masks a portion of the first communication data 16 (e.g., the second electronic device 44 forgoes displaying the first video data 20 and / or forgoes playing the first audio data 18). As described herein, in some embodiments, the first communication data 16 includes first environmental data 22 (as Figure 1C shown), and the second communication data 46 includes second environmental data 52 (as Figure 1C shown). In such embodiments, when the first electronic device 14 and the second electronic device 44 are local, the first electronic device 14 masks a portion of the second environmental data 52, and the second electronic device 44 masks a portion of the first environmental data 22.
[0046] In some embodiments, the first electronic device 14 displays a first video passthrough 74 of the first environment 10. In some embodiments, the first electronic device 14 includes an image sensor (e.g., a rear camera) having a first field of view 72. The first video passthrough 74 represents a video feed captured by the image sensor of the first electronic device 14. Since the second person 42 is within the first field of view 72, the first video passthrough 74 includes a representation of the second person 42 (e.g., a video feed of the second person 42). Similarly, in some embodiments, the second electronic device 44 displays a second video passthrough 78 of the first environment 10. In some embodiments, the second electronic device 44 includes an image sensor having a second field of view 76. The second video passthrough 78 represents a video feed captured by the image sensor of the second electronic device 44. Since the first person 12 is within the first field of view 72, the second video passthrough 78 includes a representation of the first person 12 (e.g., a video feed of the first person 12).
[0047] See Figure 1E, in response to determining that the second electronic device 44 is local, the first HMD 24 masks a portion of the second communication data 46. For example, in some embodiments, the first HMD 24 forgoes displaying the second video data 50 and / or the ER object 30 representing the second person 42 (as shown in Figure 1C ) within the first ER scene 26. In some embodiments, the first HMD 24 forgoes playing the second audio data 48 through the speakers of the first HMD 24. In some embodiments, not playing the second audio data 48 allows the first person 12 to hear the voice 80 of the second person 42. Since the first HMD 24 does not play the second audio data 48, the second audio data 48 does not interfere with the voice 80 of the second person 42, thus allowing the first person 12 to hear the voice 80.
[0048] In some embodiments, the first HMD 24 presents a first passthrough 84 of the first environment 10. In some embodiments, the first HMD 24 includes an environmental sensor having a first detection field 82 (e.g., a depth sensor such as a depth camera, and / or an image sensor such as a rear camera). In some embodiments, the first passthrough 84 includes a video passthrough similar to the first video passthrough 74 shown in Figure 1D . In some embodiments, the first passthrough 84 includes an optical passthrough, where light from the first environment 10 (e.g., natural light or artificial light) is allowed to enter the eyes of the first person 12. Since the second person 42 is within the first detection field 82, the first passthrough 84 includes a representation of the second person 42 (e.g., a video feed of the second person 42, or light reflected from the second person 42). Presenting the first passthrough 84 enhances the user experience provided by the first HMD 24 because the first passthrough 84 typically has lower latency than displaying the second video data 50 or generating the ER object 30 based on the second environmental data 52.
[0049] Refer to Figure 1F , in response to determining that the first HMD 24 is local, the second HMD 54 masks a portion of the first communication data 16. For example, in some embodiments, the second HMD 54 forgoes displaying the first video data 20 or the ER object 60 representing the first person 12 (as shown in Figure 1C ) within the second ER scene 56. In some embodiments, the second HMD 54 forgoes playing the first audio data 18 through the speakers of the second HMD 54. In some embodiments, not playing the first audio data 18 allows the second person 42 to hear the voice of the first person 12 (e.g., direct audio from the first person 12). Since the second HMD 54 does not play the first audio data 18, the first audio data 18 does not interfere with the voice of the first person 12, thus allowing the second person 42 to hear the voice.
[0050] In some specific implementations, the second HMD 54 presents a second passthrough 88 of the first environment 10. In some specific implementations, the second HMD 54 includes an environmental sensor (e.g., a depth sensor such as a depth camera, and / or an image sensor such as a rear camera) having a second detection field 86. In some specific implementations, the second passthrough 88 includes a video passthrough similar to Figure 1D the second video passthrough 78 shown. In some specific implementations, the second passthrough 88 includes an optical passthrough that allows light (e.g., natural light or artificial light) from the first environment 10 to enter the eyes of the second person 42. Since the first person 12 is within the second detection field 86, the second passthrough 88 includes a representation of the first person 12 (e.g., a video feed of the first person 12, or light reflected from the first person 12). Presenting the second passthrough 88 enhances the user experience provided by the second HMD 54 because the second passthrough 88 tends to have lower latency than displaying the first video data 20 or generating an ER object 60 based on the first environmental data 22.
[0051] As Figure 1G shown, in some specific implementations, the first HMD 24 communicates with multiple devices. Figure 1G Shown is a third environment 90 including a third person 92 associated with a third electronic device 94. Figure 1G Also shown is a fourth environment 100 including a fourth person 102 (e.g., the fourth person 102 is wearing the third HMD 104 on the head of the fourth person 102) associated with a third HMD 104. In Figure 1G the example, the first HMD 24 communicates with the second HMD 54, the third electronic device 94, and the third HMD 104. In some specific implementations, the first HMD 24, the second HMD 54, the third electronic device 94, and the third HMD 104 are in a conference call (e.g., in a conference call session).
[0052] In some specific implementations, the first HMD 24 determines that the second HMD 54 is local because the second HMD 54 is in the same environment as the first HMD 24. Thus, as described herein, in some specific implementations, the first HMD 24 masks a portion of the communication data associated with the second HMD 24. Additionally, as described herein, in some specific implementations, the first HMD 24 presents a first passthrough 84 of the first environment 10. As Figure 1G shown, the first HMD 24 presents the first passthrough 84 within the first ER set 26.
[0053] In some specific implementations, the first HMD 24 determines that the third electronic device 94 is remote because the third electronic device 94 is not in the same environment as the first HMD 24. In some specific implementations, the first HMD 24 determines that the third electronic device 94 is a non-HMD (e.g., a tablet computer, a smart phone, a media player, a laptop computer, or a desktop computer). In Figure 1G an example, the first HMD 24 displays a card GUI element 96 that includes video data 98 associated with (e.g., originating from) the third electronic device 94.
[0054] In some specific implementations, the first HMD 24 determines that the third HMD 104 is remote because the third HMD 104 is not in the same environment as the first HMD 24. In some specific implementations, the first HMD 24 determines that a fourth person 102 is using an HMD-type device. Thus, as Figure 1G shown, the first HMD 24 displays an ER object 106 representing the fourth person 102. In some specific implementations, the ER object 106 includes an ER representation of the fourth person 102. In some specific implementations, the ER object 106 includes an avatar of the fourth person 102.
[0055] Figure 2A is a flowchart representation of a method 200 for presenting communication data. In various specific implementations, the method 200 is performed by a first device associated with a first person. In some specific implementations, the first device includes a display, a non-transitory memory, and one or more processors coupled to the display and the non-transitory memory. In some specific implementations, the method 200 is performed by the first electronic device 14, the second electronic device 44, the first HMD 24, and / or the second HMD 54. In some specific implementations, the method 200 is performed by processing logic (including hardware, firmware, software, or a combination thereof). In some specific implementations, the method 200 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory).
[0056] As shown in block 210, in various specific implementations, the method 200 includes obtaining communication data associated with (e.g., originating from or generated by) a second device corresponding to a second person. For example, as Figure 1A shown, the first electronic device 14 obtains second communication data 46 associated with the second electronic device 44. In some specific implementations, the method 200 includes receiving the communication data via a network. In some specific implementations, the communication data includes audio data (e.g., network audio, such as Figures 1A to 1F shown second audio data 48). In some specific implementations, the communication data includes video data (e.g., Figures 1A to 1FThe second video data shown (50). In some specific implementations, the communication data includes environmental data (e.g., Figure 1C The second environmental data shown (52). In some specific implementations, the communication data (e.g., video data and / or environmental data) encodes an ER object (e.g., an avatar) representing a second person.
[0057] As shown in block 220, in various specific implementations, method 200 includes determining whether the first device and the second device are in a shared environment. In some specific implementations, method 200 includes the first device determining whether the second device is in the same environment as the first device. For example, the first electronic device 14 determines whether the second electronic device 44 is in the first environment 10. In some specific implementations, method 200 includes the first device determining whether the second device is local or remote.
[0058] As shown in block 230, in various specific implementations, method 200 includes, in response to determining that the first device and the second device are not in a shared environment, displaying an ER representation of the second person based on communication data associated with the second device. In some specific implementations, displaying the ER representation includes displaying video data included in the communication data. For example, as Figure 1A shown, the first electronic device 14 displays the second video data 50 in response to determining that the second electronic device 44 is not in the first environment 10. In some specific implementations, displaying the ER representation includes displaying an ER scene, displaying card GUI elements within the ER scene, and displaying video data associated with the second device within the card GUI elements. For example, as Figure 1B shown, the first HMD 24 displays the first ER scene 26, the card GUI elements 28 within the first ER scene 26, and the second video data 50 within the card GUI elements 28. In some specific implementations, displaying the ER representation includes generating an ER object based on video data and / or environmental data associated with the second device and displaying the ER object in the ER scene. For example, as Figure 1C shown, the first HMD 24 generates and displays an ER object 30 representing the second person 42.
[0059] Referring to Figure 2B , as shown in block 240, in some specific implementations, method 200 includes determining the device type of the second device and generating an ER representation of the second person based on the device type. For example, as Figures 1B to 1C shown, the first HMD 24 determines the device type of the device associated with the second person 42 and generates an ER representation of the second person based on the device type. In Figure 1B the example, the first HMD 24 determines that the second person 42 is using a non-HMD, and the first HMD 24 decodes and displays the second video data 50 within the card GUI elements 28. InFigure 1C In the example of, the first HMD 24 determines that the second person 42 is using an HMD, and the first HMD 24 generates and displays an ER object 30 representing the second person 42.
[0060] As shown in block 242, in some specific embodiments, method 200 includes, in response to the device type being a first device type, generating a first type of ER representation of the second person based on communication data associated with the second device. As shown in block 242a, in some specific embodiments, generating a first type of ER representation of the second person includes generating a three-dimensional (3D) ER object (e.g., an avatar) representing the second person. As shown in block 242b, in some specific embodiments, the first device type includes an HMD. For example, as Figure 1C shown, the first HMD 24 generates an ER object 30 representing the second person 42 in response to determining that the second person 42 is using an HMD.
[0061] As shown in block 244, in some specific embodiments, method 200 includes, in response to the device type being a second device type, generating a second type of ER representation of the second person based on communication data associated with the second device. As shown in block 244a, in some specific embodiments, generating a second type of ER representation of the second person includes generating a two-dimensional (2D) ER object representing the second person. As shown in block 244b, in some specific embodiments, the second type of ER representation includes a video of the second person. In some specific embodiments, the video is encoded in the communication data associated with the second device. As shown in block 244c, in some specific embodiments, the second device type includes a handheld device (e.g., a smart phone, a tablet computer, a laptop computer, a media player, and / or a watch). As shown in block 244d, in some specific embodiments, the second device type includes a device that is not an HMD (e.g., a non-HMD, such as a handheld device, a desktop computer, a television, and / or a projector). As shown in block 244e, in some specific embodiments, method 200 includes displaying a second type of ER representation of the second person within a GUI element within a similarity to a card (e.g., a card GUI element, such as Figure 1B the card GUI element 28 shown). For example, as Figure 1B shown, the first HMD 24 displays second video data 50 in response to determining that the second person 42 is associated with a non-HMD. In some specific embodiments, method 200 includes changing the position of a GUI element (e.g., the card GUI element 28) within the first ER scene based on the movement of the second person and / or the second electronic device within the second environment. For example, method 200 includes moving the GUI element in the same direction as the second person and / or the second electronic device.
[0062] Refer to Figure 2C, as shown in block 250, in some embodiments, the first device includes one or more speakers, and method 200 includes outputting audio corresponding to the second person via the one or more speakers. In some embodiments, the audio is spatialized to provide the appearance that the audio originates from the ER representation of the second person. For example, as described with respect to Figure 1B above, the first HMD 24 spatializes the second audio data 48 to provide the appearance that the second audio data 48 originates from the card GUI element 28. Similarly, as described with respect to Figure 1C above, the first HMD 24 spatializes the second audio data 48 to provide the appearance that the second audio data 48 originates from the ER object 30 representing the second person 42.
[0063] As shown in block 252, in some embodiments, method 200 includes generating audio based on communication data associated with the second device. For example, as shown in Figure 1A above, the first electronic device 14 extracts the second audio data 48 from the second communication data 46. In some embodiments, method 200 includes decoding the communication data to identify the audio (e.g., the first electronic device 14 decodes the second communication data 46 to identify the second audio data 48).
[0064] As shown in block 254, in some embodiments, method 200 includes generating early reflections to provide the appearance that the audio is reflecting off a surface. For example, the first HMD 24 generates early reflections for the second audio data 48 to provide the appearance that the sound corresponding to the second audio data 48 is reflecting off the surface of the first environment 10. In some embodiments, method 200 includes outputting the early reflections before outputting the audio (e.g., the first HMD 24 outputs the early reflections before playing the second audio data 48). In some embodiments, method 200 includes outputting the early reflections and the audio simultaneously (e.g., the first HMD 24 plays the early reflections of the second audio data 48 and the second audio data 48 simultaneously). In some embodiments, method 200 includes generating early reflections based on the type of the first environment. In some embodiments, the first environment is a physical set, and the early reflections provide the appearance that the audio is reflecting off the physical surface of the physical set. In some embodiments, the first environment is an ER set (e.g., a virtual environment), and the early reflections provide the appearance that the audio is reflecting off the ER surface (e.g., a virtual surface) of the ER set.
[0065] As shown in block 256, in some embodiments, method 200 includes generating a late echo to provide an appearance of audio echoing. For example, the first HMD 24 generates a late echo for the second audio data 48 to provide an appearance that the sound corresponding to the second audio data 48 is echoing in the first environment 10. In some embodiments, method 200 includes outputting the late echo after outputting the audio (e.g., the first HMD 24 outputs the late echo after playing the second audio data 48). In some embodiments, method 200 includes generating the late echo based on the type of the first environment. In some embodiments, the first environment is a physical set, and the late echo provides an appearance that the audio is echoing in the physical set. In some embodiments, the first environment is an ER set (e.g., a virtual environment), and the early reflections provide an appearance that the audio is reflecting from the ER surface (e.g., a virtual surface) of the ER set.
[0066] As shown in block 260, in some embodiments, method 200 includes, in response to determining that the first device and the second device are in a shared environment, forgoing displaying an ER representation of the second person. In some embodiments, method 200 includes forgoing displaying video data included in communication data associated with the second device. For example, as Figure 1D shown, in response to determining that the second electronic device 44 is in the same environment as the first electronic device 14, the first electronic device 14 forgoes displaying the second video data 50. Similarly, as Figure 1E shown, in response to determining that the second electronic device 44 is in the same environment as the first electronic device 14, the first HMD 24 forgoes displaying the second video data 50. In some embodiments, method 200 includes forgoing displaying an ER representation of the second person generated based on communication data associated with the second person (e.g., forgoing displaying the avatar of the second person when the second person is local). For example, as Figure 1F shown, in response to determining that the second HMD 54 is in the same environment as the first HMD 24, the first HMD 24 forgoes displaying the ER object 30 representing the second person.
[0067] As shown in block 262, in some embodiments, method 200 includes presenting a passthrough of the shared environment. For example, as Figures 1E to 1F shown, the first HMD 24 presents a first passthrough 84 of the first environment 10. As shown in block 262a, in some embodiments, method 200 includes presenting an optical passthrough of the shared environment. For example, as described with respect to Figures 1E to 1F above, in some embodiments, the first passthrough 84 includes an optical passthrough. As shown in block 262b, in some embodiments, method 200 includes displaying a video passthrough of the shared environment. For example, displaying Figure 1D the first video passthrough 74 shown in. For example, as described with respect to Figures 1E to 1FAs described, in some specific embodiments, the first passthrough 84 includes video passthrough. As shown in block 262c, in some specific embodiments, method 200 includes displaying a window and presenting the passthrough within the window. For example, as Figures 1E to 1F shown, within a card GUI element similar to the card GUI element 28 shown in Figure 1B the first passthrough 84 is shown.
[0068] Referring Figure 2D , as shown in block 220a, in some specific embodiments, the shared environment includes a shared physical set. For example, as with respect to Figure 1A described, in some specific embodiments, the first environment 10 includes a first physical set. As shown in block 220b, in some specific embodiments, the shared environment includes a shared ER set. For example, as with respect to Figure 1A described, in some specific embodiments, the first environment 10 includes a first ER set.
[0069] As shown in block 220c, in some specific embodiments, determining whether a first device and a second device are in a shared environment includes determining whether an identifier (ID) associated with the second device can be detected via short-range communication. Exemplary short-range communications include Bluetooth, Wi-Fi, Near Field Communication (NFC), ZigBee, etc. For example, with respect to Figure 1A , the first electronic device 14 determines whether the first electronic device 14 can detect the ID associated with the second electronic device 44. In Figure 1A the example of, the first electronic device 14 does not detect, via short-range communication, the ID associated with the second electronic device 44. Accordingly, the first electronic device 14 determines that the second electronic device 44 is remote. Again, with respect to Figure 1C , in some specific embodiments, the first HMD 24 determines that the second HMD 54 is remote because the first HMD 24 cannot detect the ID of the second HMD 54 via short-range communication. Again, with respect to Figure 1F , the first HMD 24 determines that the second HMD 54 is local because the first HMD 24 detects the ID of the second HMD 54 via short-range communication.
[0070] As shown in block 220d, in some specific embodiments, determining whether a first device and a second device are in a shared environment includes determining whether the audio received via the microphone of the first device is within a similarity of the audio encoded in the communication data associated with the second device. For example, in some specific embodiments, method 200 includes determining whether the direct audio from a second person is within a similarity of the network audio associated with the second device. For example, with respect to Figure 1D, in some specific implementations, the first electronic device 14 determines that the second electronic device 44 is local because the voice 80 of the second person 42 is within the similarity of the second audio data 48 encoded in the second communication data 46.
[0071] As shown in block 220e, in some specific implementations, determining whether the first device and the second device are in a shared environment includes determining whether the second person is in the shared environment based on an image captured via the image sensor of the first device. For example, with respect to Figure 1D , in some specific implementations, the first electronic device 14 captures an image corresponding to the first field of view 72 and performs face detection / recognition on the image to determine whether the second person 42 is in the first environment 10. Since the second person 42 is in the first field of view 72, after performing face detection / recognition on the image, the first electronic device 14 determines that the second person 42 is local.
[0072] Figure 3A is a flowchart representation of a method 300 for masking communication data. In various specific implementations, the method 300 is performed by a first device that includes an output device, a non-transitory memory, and one or more processors coupled to the output device and the non-transitory memory. In some specific implementations, the method 300 is performed by the first electronic device 14, the second electronic device 44, the first HMD 24, and / or the second HMD 54. In some specific implementations, the method 300 is performed by processing logic (including hardware, firmware, software, or a combination thereof). In some specific implementations, the method 300 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., a memory).
[0073] As shown in block 310, in various specific implementations, the method 300 includes obtaining communication data associated with the second device (e.g., originating from or generated by the second device) when the first device is in a communication session with the second device. For example, as Figure 1D shown, the first electronic device 14 obtains the second communication data 46 associated with the second electronic device 44. In some specific implementations, the method 300 includes receiving communication data via a network. In some specific implementations, the communication data includes audio data (e.g., network audio, such as Figures 1A to 1F the second audio data 48 shown). In some specific implementations, the communication data includes video data (e.g., Figures 1A to 1F the second video data 50 shown). In some specific implementations, the communication data includes environmental data (e.g., Figure 1C the second environmental data 52 shown). In some specific implementations, the communication data (e.g., video data and / or environmental data) encodes an ER object (e.g., an avatar) representing the second person.
[0074] As shown in block 320, in various embodiments, method 300 includes determining that a first device and a second device are in a shared environment. In some embodiments, method 300 includes the first device determining that the second device is in the same environment as the first device. For example, with respect to Figure 1D , the first electronic device 14 determines that the second electronic device 44 is in the first environment 10. In some embodiments, method 300 includes the first device determining that the second device is local.
[0075] As shown in block 330, in various embodiments, method 300 includes masking a portion of the communication data to prevent the output device from outputting that portion of the communication data. In some embodiments, masking a portion of the communication data includes forgoing rendering that portion of the communication data. For example, with respect to Figure 1D , the first electronic device 14 forgoes rendering the second communication data 46 in response to determining that the second electronic device 44 is local. In various embodiments, when the second device is local, masking a portion of the communication data enhances the user experience of the first device by allowing a first person to directly hear and / or see a second person without the latency associated with communicating over a network.
[0076] Referring to Figure 3B , as shown in block 320a, in some embodiments, determining whether the first device and the second device are in a shared physical setting includes detecting an identifier associated with the second device via short-range communication. Exemplary short-range communication includes Bluetooth, Wi-Fi, Near Field Communication (NFC), ZigBee, and the like. For example, with respect to Figure 1D , in some embodiments, the first electronic device 14 determines that the second electronic device 44 is local in response to detecting an ID associated with the second electronic device 44 via short-range communication. As another example, with respect to Figure 1F , in some embodiments, the first HMD 24 determines that the second HMD 54 is local in response to detecting the ID of the second HMD 54 via short-range communication.
[0077] As shown in block 320b, in some embodiments, determining that the first device and the second device are in a shared physical setting includes the first device's microphone detecting first audio that is within a similarity to second audio encoded in the communication data (e.g., within a similarity threshold of the second audio). In some embodiments, method 300 includes determining that the second device is local in response to detecting direct audio that is within a similarity to network audio encoded in the communication data. For example, with respect to Figure 1E, in some specific implementations, the first HMD 24 determines that the second electronic device 44 is local in response to detecting speech 80 via the microphone of the first HMD 24 and determining that the speech 80 is within a similarity to the second audio data 48 encoded in the second communication data 46.
[0078] As shown in block 320c, in some specific implementations, determining that the first device and the second device are in a shared physical setting includes detecting a person associated with the second device via the image sensor of the first device. For example, with respect to Figure 1D , in some specific implementations, the first electronic device 14 captures an image corresponding to the first field of view 72 and performs face detection / recognition on the image to determine whether the second person 42 is in the first environment 10. Since the second person 42 is in the first field of view 72, after performing face detection / recognition on the image, the first electronic device 14 determines that the second device 44 is local.
[0079] As shown in block 330a, in some specific implementations, the output device includes a speaker, and masking a portion of the communication data includes masking the audio portion of the communication data to prevent the speaker from playing that audio portion of the communication data. As with respect to Figure 1D described, in some specific implementations, the first electronic device 14 masks the second audio data 48 to prevent the speaker of the first electronic device 14 from playing the second audio data 48. As described herein, forgoing playing the second audio data 48 at the first electronic device 14 enhances the user experience of the first electronic device 14 by allowing the first person 12 to listen to the speech 80 spoken by the second person 42 without the interference of the second audio data 48.
[0080] As shown in block 330b, in some specific implementations, the output device includes a display, and masking a portion of the communication data includes masking the video portion of the communication data to prevent the display from showing that video portion of the communication data. As with respect to Figure 1E described, in some specific implementations, the first HMD 24 masks the second video data 50 to prevent the display of the first HMD 24 from showing the second video data 50. As described herein, forgoing showing the second video data 50 at the first HMD 24 enhances the user experience of the first HMD 24 by not distracting the first person 12 with the second video data 50 and allowing the first person 12 to see the first passthrough 84 of the first environment 10.
[0081] As shown in block 330c, in some specific implementations, the communication data encodes an ER representation of a person associated with the second device, and masking a portion of the communication data includes forgoing showing the ER representation of that person. For example, as with respect to Figure 1FAs shown, in some specific implementations, the first HMD 24 abandons displaying the ER object 30 representing the second person 42 in response to determining that the second person 42 is in the same environment as the first HMD 24. As described herein, abandoning the display of the ER object 30 at the first HMD 24 enhances the user experience of the first HMD 24 by allowing the first person 12 to see the first passthrough 84 associated with a lower latency compared to generating and displaying the ER object 30.
[0082] As shown in block 340, in some specific implementations, the method 300 includes presenting a passthrough of a shared physical set. For example, as Figures 1E to 1F shown, the first HMD 24 presents a first passthrough 84 of the first environment 10. As shown in block 340a, in some specific implementations, presenting the passthrough includes presenting an optical passthrough of the shared physical set. For example, as with respect to Figures 1E to 1F described, in some specific implementations, the first passthrough 84 includes an optical passthrough, where the first HMD 24 allows light from the first environment 10 to reach the eyes of the first person 12. As shown in block 340b, in some specific implementations, presenting the passthrough includes displaying a video passthrough of the shared physical set. For example, as with respect to Figures 1E to 1F described, in some specific implementations, the first passthrough 84 includes a video passthrough, where the display of the first HMD 24 displays a video stream corresponding to the first detection field 82.
[0083] See Figure 3C , as shown in block 350, in some specific implementations, the method 300 includes detecting movement of a second device away from the shared physical set and abandoning masking a portion of communication data to allow an output device to output that portion of the communication data. For example, in some specific implementations, the first electronic device 14 detects that the second electronic device 44 has left the first environment 10, and the first electronic device 14 abandons masking the second communication data 46 in response to detecting that the second electronic device 44 has left the first environment 10.
[0084] As shown in block 350a, in some specific implementations, detecting movement of a second device away from the shared physical set includes determining that an identifier associated with the second device cannot be detected via short-range communication. For example, in some specific implementations, the first electronic device 14 and / or the first HMD 24 determine that an ID associated with the second electronic device 44 and / or the second HMD 54 cannot be detected via short-range communication.
[0085] As shown in block 350b, in some embodiments, detecting movement of the second device away from the shared physical set includes determining that a first audio detected via a microphone of the first device is not within a similarity with a second audio encoded in communication data. For example, in some embodiments, the first electronic device 14 and / or the first HMD 24 determine that the audio detected via a microphone of the first electronic device 14 and / or the first HMD 24 does not match the second audio data 48.
[0086] As shown in block 350c, in some embodiments, detecting movement of the second device away from the shared physical set includes determining that environmental data captured by an environmental sensor of the first device indicates that a person associated with the second device has moved away from the shared physical set. For example, in some embodiments, the first electronic device 14 and / or the first HMD 24 determine that environmental data (e.g., images captured by a camera and / or depth data captured by a depth sensor) captured by an environmental sensor of the first electronic device 14 and / or the first HMD 24 indicates that the second person 42 is not in the first environment 10.
[0087] As shown in block 350d, in some embodiments, the output device includes a speaker, and forgoing masking a portion of the communication data includes outputting an audio portion of the communication data via the speaker. For example, in response to determining that the second electronic device 44 and / or the second HMD 54 has left the first environment 10, the first electronic device 14 and / or the first HMD 24 outputs the second audio data 48.
[0088] As shown in block 350e, in some embodiments, outputting the audio portion includes spatializing the audio portion to provide an appearance that the audio portion is originating from an ER representation of a person associated with the second device. For example, as described with respect to Figures 1B to 1C the first electronic device 14 and / or the first HMD 24 spatialize the second audio data 48 to provide an appearance that the second audio data 48 is originating from the card GUI element 28 or the ER object 30 representing the second person 42.
[0089] As shown in block 350f, in some embodiments, the output device includes a display, and forgoing masking a portion of the communication data includes displaying a video portion of the communication data on the display. For example, in some embodiments, in response to detecting that the second electronic device 44 has left the first environment 10, the first electronic device 14 and / or the first HMD 24 displays the second video data 50.
[0090] As shown in block 350g, in some embodiments, communication data encodes an ER representation of a person associated with a second device, and forgoing masking a portion of the communication data includes displaying the person's ER representation. For example, in some embodiments, in response to detecting that the second HMD 54 has left the first environment 10, the first HMD 24 displays an ER object 30 representing a second person 42.
[0091] Figure 4 FIG. is a block diagram of a device 400 for presenting / masking communication data according to some embodiments. Although some specific features are shown, those of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure more relevant aspects of the embodiments disclosed herein. To that end, as a non-limiting example, in some embodiments, the device 400 includes one or more processing units (CPUs) 401, a network interface 402, a programming interface 403, a memory 404, an environmental sensor 407, one or more input / output (I / O) devices 409, and one or more communication buses 405 for interconnecting these and various other components.
[0092] In some embodiments, the network interface 402 is provided to establish and maintain a metadata tunnel, among other things, between a cloud-hosted network management system and at least one private network including one or more compatible devices. In some embodiments, one or more communication buses 405 include circuitry for interconnecting and controlling communication between system components. The memory 404 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices, and may include non-volatile memory, such as one or more disk storage devices, optical storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 404 optionally includes one or more storage devices that are remotely located from one or more CPUs 401. The memory 404 includes non-transitory computer-readable storage medium.
[0093] In various embodiments, the environmental sensor 407 includes an image sensor. For example, in some embodiments, the environmental sensor 407 includes a camera (e.g., a scene-facing camera, an outward-facing camera, or a rear-facing camera). In some embodiments, the environmental sensor 407 includes a depth sensor. For example, in some embodiments, the environmental sensor 407 includes a depth camera.
[0094] In some embodiments, the memory 404 or the non-transitory computer-readable storage medium of the memory 404 stores the following programs, modules, and data structures or subsets thereof, including an optional operating system 406, a data acquirer 410, an environmental analyzer 420, and an ER experience generator 430. In various embodiments, the device 400 performsFigures 2A to 2D The method 200 shown. In various specific embodiments, the device 400 executes Figures 3A to 3C The method 300 shown. In various specific embodiments, the device 400 implements the first electronic device 14, the second electronic device 44, the first HMD 24, the second HMD 54, the third electronic device 94, and / or the third HMD 104.
[0095] In some specific embodiments, the data acquirer 410 acquires data. In some specific embodiments, the data acquirer 410 acquires communication data associated with another device (e.g., Figures 1A to 1F The second communication data 46 shown). In some specific embodiments, the data acquirer 410 executes at least a part of the method 200. For example, in some specific embodiments, the data acquirer 410 executes the operation represented by Figure 2A The block 210 shown. In some specific embodiments, the data acquirer 410 executes at least a part of the method 300. For example, in some specific embodiments, the data acquirer 410 executes the operation represented by Figure 3A The block 310 shown. For this purpose, the data acquirer 410 includes instructions 410a and heuristics and metadata 410b.
[0096] As described herein, in some specific embodiments, the environment analyzer 420 determines whether the device 400 and another device are in a shared environment. For example, the environment analyzer 420 determines whether the first electronic device 14 and the second electronic device 44 are in the first environment 10. In some specific embodiments, the environment analyzer 420 executes at least a part of the method 200. For example, in some specific embodiments, the environment analyzer 420 executes the operations represented by Figure 2A and Figure 2D The block 220 in. In some specific embodiments, the environment analyzer 420 executes at least a part of the method 300. For example, in some specific embodiments, the environment analyzer 420 executes the operations represented by Figure 3A and Figure 3B The block 320 in. For this purpose, the environment analyzer 420 includes instructions 420a and heuristics and metadata 420b.
[0097] In some specific embodiments, in response to the environment analyzer 420 determining that another device is remote (e.g., not in the same environment as the device 400), the ER experience generator 430 displays an ER representation of the person associated with that another device. In some specific embodiments, the ER experience generator 430 executes at least a part of the method 200. For example, in some specific embodiments, the ER experience generator 430 executes the operation represented by Figures 2A to 2CThe operations represented by boxes 230, 240, 250, and 260 in. In some specific implementations, in response to the environment analyzer 420 determining that another device is local (e.g., in the same environment as device 400), the ER experience generator 430 masks a portion of the communication data. In some specific implementations, the ER experience generator 430 performs at least a portion of method 300. For example, in some specific implementations, the ER experience generator 430 performs the operations represented by boxes 330, 340, and 350 shown by Figures 3A to 3C . To this end, the ER experience generator 430 includes instructions 430a as well as heuristics and metadata 430b.
[0098] In some specific implementations, one or more I / O devices 409 include one or more sensors for capturing environmental data associated with the environment (e.g., Figures 1A to 1G the first environment 10 shown). For example, in some specific implementations, one or more I / O devices 409 include an image sensor (e.g., a camera), an ambient light sensor (ALS), a microphone, and / or a position sensor. In some specific implementations, one or more I / O devices 409 include a display (e.g., an opaque display or an optical see-through display) and / or a speaker for presenting the communication data.
[0099] While the various aspects of the specific implementations within the scope of the appended claims have been described above, it should be apparent that the various features of the above-described specific implementations can be embodied in a wide variety of forms, and any specific structure and / or function described above is merely illustrative. Based on the present disclosure, those skilled in the art should understand that the aspects described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, a device can be implemented using any number of the aspects described herein and / or a method can be practiced. Additionally, other structures and / or functions can be used to implement such a device and / or to practice such a method in addition to or different from one or more of the aspects described herein.
[0100] It will also be understood that although terms such as "first," "second," etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first node can be referred to as a second node, and similarly, a second node can be referred to as a first node, which changes the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. The first node and the second node are both nodes, but they are not the same node.
[0101] The terms used in this specification are merely for the purpose of describing particular embodiments and are not intended to limit the claims. As used in the description of the present embodiments and the appended claims, the singular forms "a", "an", and "the" are intended to also include the plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms "comprises" and / or "comprising", when used in this specification, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0102] As used herein, the term "if" can be interpreted to mean "when the precondition is true" or "while the precondition is true" or "in response to determining" or "in accordance with determining" or "in response to detecting" that the precondition is true, depending on the context. Similarly, the phrases "if it is determined [that the precondition is true]" or "if [the precondition is true]" or "when [the precondition is true]" are interpreted to mean "when it is determined that the precondition is true" or "in response to determining" or "in accordance with determining" that the precondition is true or "when it is detected that the precondition is true" or "in response to detecting" that the precondition is true, depending on the context.
Claims
1. A method, the method comprising: At a first device including an output device, a non-transitory memory, and one or more processors coupled to the output device and the non-transitory memory: When the first device is in a communication session with a second device: Obtain communication data associated with the second device; Determine that the first device and the second device are in a shared physical setting; and Mask a portion of the communication data to prevent the output device from outputting the portion of the communication data.
2. The method according to claim 1, wherein the output device comprises a speaker, and wherein masking the portion of the communication data comprises masking the audio portion of the communication data so as to prevent the speaker from playing the audio portion of the communication data.
3. The method according to any one of claims 1 and 2, wherein the output device comprises a display, and wherein masking the portion of the communication data comprises masking the video portion of the communication data so as to prevent the display from displaying the video portion of the communication data.
4. The method according to any one of claims 1 to 3, wherein the communication data encodes a virtual representation of a person associated with the second device, and wherein masking the portion of the communication data comprises refraining from displaying the virtual representation of the person.
5. The method according to any one of claims 1 to 4, further comprising: Present a pass-through of the shared physical setting.
6. The method according to claim 5, wherein presenting the passthrough comprises presenting an optical passthrough of the shared physical set.
7. The method according to claim 5, wherein presenting the passthrough comprises displaying a video passthrough of the shared physical set.
8. The method according to any one of claims 1 to 7, wherein determining that the first device and the second device are in the shared physical set comprises detecting an identifier associated with the second device via short-range communication.
9. The method according to any one of claims 1 to 8, wherein determining that the first device and the second device are in the shared physical set comprises detecting, via a microphone of the first device, a first audio within a similarity threshold of a second audio encoded in the communication data.
10. The method according to any one of claims 1 to 9, wherein determining that the first device and the second device are in the shared physical set comprises detecting, via an image sensor of the first device, a person associated with the second device.
11. The method according to any one of claims 1 to 10, further comprising: Detect movement of the second device away from the shared physical setting; And Cease masking the portion of the communication data to allow the output device to output the portion of the communication data.
12. The method according to claim 11, wherein detecting the movement of the second device away from the shared physical setting includes determining that an identifier associated with the second device cannot be detected via short-range communication.
13. The method according to any one of claims 11 and 12, wherein detecting the movement of the second device away from the shared physical setting includes determining that a first audio detected via a microphone of the first device is not within a similarity threshold of a second audio encoded in the communication data.
14. The method according to any one of claims 11 to 13, wherein detecting the movement of the second device away from the shared physical setting includes determining that environmental data captured by an environmental sensor of the first device indicates that a person associated with the second device has moved away from the shared physical setting.
15. The method according to any one of claims 11 to 14, wherein the output device includes a speaker, and wherein forgoing masking of the portion of the communication data includes outputting an audio portion of the communication data via the speaker.
16. The method according to claim 15, wherein outputting the audio portion includes spatializing the audio portion so as to provide an appearance that the audio portion is originating from a virtual representation of a person associated with the second device.
17. The method according to any one of claims 11 to 16, wherein the output device includes a display, and wherein forgoing masking of the portion of the communication data includes displaying a video portion of the communication data on the display.
18. The method according to any one of claims 11 to 17, wherein the communication data encodes a virtual representation of a person associated with the second device, and wherein forgoing masking of the portion of the communication data includes displaying the virtual representation of the person.
19. An apparatus, comprising: One or more processors; A display; A non-transitory memory; And One or more programs stored in the non-transitory memory, the one or more programs, when executed by the one or more processors, cause the device to perform any of the methods according to claims 1 to 18.
20. A non-transitory memory storing one or more programs, the one or more programs, when executed by one or more processors of a device having a display, cause the device to perform any one of the methods recited in claims 1 to 18.
21. An apparatus, comprising: One or more processors; A display; A non-transitory memory; And Means for causing the device to perform any of the methods according to claims 1 to 18.