Presenting views and / or representations of objects in three-dimensional environment
By synchronously presenting the physical object model of the 3D environment on a first electronic device and a second electronic device, the problem of unstable view and virtual representation of real-world objects between users in different physical locations is solved, achieving stable and smooth real-time collaboration and interaction, and improving user experience and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2026-03-31
AI Technical Summary
In existing computer graphics environments, there are challenges in the virtual representation of real-world objects and the synchronization and stable rendering of views, especially in real-time collaboration between users in different physical locations, resulting in unstable images and inefficient interaction.
The system identifies a region within the 3D environment using a first electronic device, captures a portion of that region and sends it to a second electronic device, while simultaneously presenting a 3D model of the physical object on both devices. It also tracks the user's gaze and movement using internal and external image sensors, and provides a stable view and virtual representation in conjunction with a display, thus enabling enhanced real-time guidance.
It enables synchronized and stable real-world object views and virtual representations between users in different physical locations, improving the smoothness and efficiency of real-time collaboration, reducing image motion artifacts and user interaction errors, and optimizing bandwidth usage and power consumption.
Smart Images

Figure CN121764366A_ABST
Abstract
Description
Cross-reference to related applications
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 752,492, filed January 31, 2025; U.S. Provisional Application No. 63 / 700,656, filed September 28, 2024; and U.S. Patent Application No. 19 / 295,372, filed August 8, 2025, the contents of which are incorporated herein by reference in their entirety for all purposes. Technical Field
[0002] This disclosure relates in its entirety to methods and apparatus for providing views and / or virtual representations of real-world objects. Background Technology
[0003] Some computer graphics environments provide two-dimensional and / or three-dimensional environments in which at least some of the real-world objects displayed for the user to view are virtual and computer-generated. Summary of the Invention
[0004] This disclosure relates throughout to methods and apparatus for providing views and / or virtual representations of real-world objects (also more generally referred to herein as objects). In some examples, a first electronic device communicates with one or more input devices and a second electronic device. In some examples, the first electronic device identifies a region within a three-dimensional environment; captures a portion of the three-dimensional environment corresponding to the identified region via the one or more input devices; and transmits the portion of the three-dimensional environment corresponding to the identified region to the second electronic device.
[0005] A full description of these examples is provided in the accompanying drawings and detailed embodiments, and it should be understood that the content of this invention does not limit the scope of this disclosure in any way. Attached Figure Description
[0006] To better understand the various examples described herein, reference should be made to the following detailed embodiments and accompanying drawings. Throughout the drawings, similar reference numerals generally refer to corresponding parts.
[0007] Figure 1 Examples of electronic devices that present extended real-world environments according to some examples of this disclosure are illustrated.
[0008] Figure 2 A block diagram illustrating an example architecture of a device according to some examples of this disclosure is shown.
[0009] Figure 3 Example procedures for generating views of objects are illustrated according to some examples of this disclosure.
[0010] Figures 4A to 4LExamples of views and / or virtual representations of real-world objects are illustrated according to some examples of this disclosure.
[0011] Figure 5 A flowchart illustrating an example process for sending a portion of a three-dimensional environment to an electronic device, according to some examples of this disclosure, is shown.
[0012] Figure 6 A flowchart illustrating an example process for presenting a virtual representation of a real-world object, according to some examples of this disclosure, is shown. Detailed Implementation
[0013] Some examples of this disclosure relate to methods and apparatus for providing views and / or virtual representations of real-world objects (also more generally referred to herein as objects). In some examples, a first electronic device communicates with one or more input devices and a second electronic device. In some examples, the first electronic device identifies a region within a three-dimensional environment; captures a portion of the three-dimensional environment corresponding to the identified region via the one or more input devices; and sends the portion of the three-dimensional environment corresponding to the region to the second electronic device. In some examples, when presenting a user interface element that includes a portion of the three-dimensional environment corresponding to the three-dimensional environment of the second electronic device, the first electronic device determines a physical object within that portion of the three-dimensional environment of the second electronic device and presents a three-dimensional model corresponding to the physical object within the three-dimensional environment of the first electronic device. Presenting a portion of the three-dimensional environment of the first electronic device to the second electronic device and presenting a portion of the three-dimensional environment of the second electronic device to the first electronic device is particularly useful for collaboration and provides enhanced real-time guidance by simultaneously presenting the same portion of the three-dimensional environment to users located in different physical locations.
[0014] Figure 1 An electronic device 101 is illustrated according to some examples of this disclosure, which presents an extended reality (XR) environment (e.g., a computer-generated environment that optionally includes representations of physical and / or virtual objects). In some examples, such as Figure 1 As shown, electronic device 101 is a head-mounted display or other head-mountable device configured to be worn on the head of a user of electronic device 101. An example of electronic device 101 is referenced below. Figure 2 Describe the architecture using a block diagram. For example... Figure 1 As shown, electronic device 101 and table 106 are located in a physical environment. The physical environment may include physical features such as physical surfaces (e.g., floor, wall) or physical objects (e.g., table, lamp, etc.). In some examples, electronic device 101 may be configured to detect and / or capture images of the physical environment including table 106 (exemplified in the field of view of electronic device 101).
[0015] In some examples, such as Figure 1 As shown, the electronic device 101 includes one or more internal image sensors 114a oriented toward the user's face (e.g., referred to below). Figure 2 (The described eye-tracking camera). In some examples, an internal image sensor 114a is used for eye tracking (e.g., detecting the user's gaze). The internal image sensor 114a is optionally arranged on the left and right portions of the display 120 to enable eye tracking of the user's left and right eyes. In some examples, the electronic device 101 also includes external image sensors 114b and 114c facing outwards from the user to detect and / or capture the physical environment of the electronic device 101 and / or movement of the user's hands or other body parts.
[0016] In some examples, display 120 has a field of view visible to the user (e.g., it may or may not correspond to the field of view of external image sensors 114b and 114c). Because display 120 is optionally part of a head-mounted device, the field of view of display 120 may be the same as or similar to the field of view of the user's eyes. In other examples, the field of view of display 120 may be smaller than the field of view of the user's eyes. In some examples, electronics 101 may be an optical pass-through device, through which display 120 is a transparent or translucent display through which parts of the physical environment can be directly viewed. In some examples, display 120 may be included within a transparent lens and may overlap with all or only a portion of the transparent lens. In other examples, electronics may be a video pass-through device, through which display 120 is an opaque display configured to display images of the physical environment captured by external image sensors 114b and 114c. Although a single display 120 is shown, it should be understood that display 120 may include a stereoscopic display pair.
[0017] In some examples, in response to a trigger, electronic device 101 can be configured to display in an XR environment... Figure 1 The illustrated cube represents a virtual object 104 that does not exist in the physical environment but is displayed in an XR environment positioned on top of a real-world table 106 (or a representation thereof). Optionally, in response to detecting a flat surface of the table 106 in the physical environment, the virtual object 104 may be displayed on the surface of the table 106 in the XR environment as displayed via a display 120 of the electronic device 101.
[0018] In some examples, display 120 is provided as a passive component (e.g., rather than an active component) within electronic device 101. For example, display 120 may be a transparent or semi-transparent display, as described above, and may not be configured to display virtual content (e.g., images of the physical environment captured by external image sensors 114b and 114c and / or virtual object 104). Alternatively, in some examples, electronic device 101 does not include display 120. In some such examples where display 120 is provided as a passive component or not included in electronic device 101, electronic device 101 may still include sensors (e.g., internal image sensors 114a and / or external image sensors 114b and 114c) and / or other input devices, such as those referenced below. Figure 2 One or more of the components described.
[0019] It should be understood that virtual object 104 is a representative virtual object and may include and render one or more different virtual objects (e.g., virtual objects with various dimensions, such as two-dimensional or other three-dimensional virtual objects) in a three-dimensional XR environment. For example, a virtual object may represent an application or user interface displayed in an XR environment. In some examples, a virtual object may represent content corresponding to an application and / or displayed via a user interface in an XR environment. In some examples, virtual object 104 may optionally be configured to be interactive and responsive to user input (e.g., air gestures, such as air pinch gestures, air tap gestures, and / or air touch gestures), allowing the user to virtually touch, tap, move, rotate, or otherwise interact with virtual object 104.
[0020] In some examples, displaying an object in a 3D environment may include interaction with one or more user interface objects in the 3D environment. For example, initiating the display of an object in a 3D environment may include interaction with one or more virtual option / power representations displayed in the 3D environment. In some examples, when initiating the display of an object in a 3D environment, the electronic device may track the user's gaze as input for identifying one or more virtual option / power representations as a target for selection. For example, a gaze may be used to identify one or more virtual option / power representations as a target for selection using another selection input. In some examples, a virtual option / power representation may be selected using hand-tracking input detected via an input device communicating with the electronic device. In some examples, an object displayed in a 3D environment may move and / or reorient itself in the 3D environment based on movement input detected via an input device.
[0021] In the following discussion, an electronic device communicating with a display generation component and one or more input devices is described. It should be understood that the electronic device may optionally communicate with one or more other physical user interface devices, such as a touch-sensitive surface, physical keyboard, mouse, joystick, hand-tracking device, eye-tracking device, stylus, etc. Furthermore, as described above, it should be understood that the described electronic device, display, and touch-sensitive surface may optionally be distributed among two or more devices. Therefore, as used in this disclosure, information on or displayed by an electronic device may optionally be used to describe information output by the electronic device for display on a separate display device (touch-sensitive or non-touch-sensitive). Similarly, as used in this disclosure, input received on an electronic device (e.g., touch input received on a touch-sensitive surface of the electronic device, or touch input received on the surface of a stylus) may optionally be used to describe input received on a separate input device from which the electronic device receives input information.
[0022] The device typically supports a variety of applications, such as one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, game applications, telephone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camera applications, web browsing applications, digital music player applications, TV channel browsing applications, and / or digital video player applications.
[0023] Figure 2 Block diagrams illustrating example architectures of electronic devices 201 according to some examples of this disclosure are provided. In some examples, electronic device 201 includes one or more electronic devices. For example, electronic device 201 may be a portable device, an auxiliary device for communicating with another device, a head-mounted display, etc. In some examples, electronic device 201 corresponds to the above reference. Figure 1 The described electronic device 101.
[0024] like Figure 2 As illustrated, electronic device 201 optionally includes various sensors, such as one or more hand tracking sensors 202, one or more position sensors 204, and one or more image sensors 206 (optionally corresponding to...). Figure 1 The internal image sensor 114a and / or external image sensors 114b and 114c, one or more touch-sensitive surfaces 209, one or more motion and / or orientation sensors 210, one or more eye-tracking sensors 212, one or more microphones 213 or other audio sensors, one or more body tracking sensors (e.g., torso and / or head tracking sensors), and one or more display generation components 214 (optionally corresponding to...) Figure 1The electronic device 201 includes a display 120, one or more speakers 216, one or more processors 218, one or more memories 220, and / or communication circuitry 222. One or more communication buses 208 are optionally used for communication between the aforementioned components of the electronic device 201.
[0025] Communication circuit 222 optionally includes circuitry for communicating with electronic devices, networks such as the Internet, intranets, wired and / or wireless networks, cellular networks, and wireless local area networks (LANs). Communication circuit 222 optionally includes circuitry for communicating using near field communication (NFC) and / or short-range communication such as Bluetooth®.
[0026] Processor 218 includes one or more general-purpose processors, one or more graphics processors, and / or one or more digital signal processors. In some examples, memory 220 is a non-transitory computer-readable storage medium (e.g., flash memory, random access memory, or other volatile or non-volatile memory or storage device) storing computer-readable instructions configured to be executed by processor 218 to perform the techniques, processes, and / or methods described below. In some examples, memory 220 may include more than one non-transitory computer-readable storage medium. A non-transitory computer-readable storage medium can be any medium (e.g., excluding signals) that can tangibly contain or store computer-executable instructions for use by or in connection with instruction execution systems, apparatuses, and devices. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage media may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include hard disks, optical discs based on compressed disc (CD), digital versatile optical disc (DVD), or Blu-ray technology, as well as persistent solid-state storage (such as flash memory, solid-state drives, etc.).
[0027] In some examples, the display generating component 214 includes a single display (e.g., a liquid crystal display (LCD), an organic light-emitting diode (OLED), or other types of display). In some examples, the display generating component 214 includes multiple displays. In some examples, the display generating component 214 may include a touch-enabled display (e.g., a touchscreen), a projector, a holographic projector, a retinal projector, a transparent or translucent display, etc. In some examples, the electronic device 201 includes touch-sensitive surfaces 209 for receiving user input, such as tap input and swipe input or other gestures. In some examples, the display generating component 214 and the touch-sensitive surface 209 form a touch-sensitive display (e.g., a touchscreen integrated with the electronic device 201 or a touchscreen external to the electronic device 201 and communicating with the electronic device 201).
[0028] Electronic device 201 optionally includes an image sensor 206. Image sensor 206 optionally includes one or more visible light image sensors (such as charge-coupled device (CCD) sensors) and / or complementary metal-oxide-semiconductor (CMOS) sensors operable to acquire images of physical objects from a real-world environment. Image sensor 206 may also optionally include one or more infrared (IR) sensors, such as passive or active IR sensors, for detecting infrared light from the real-world environment. For example, an active IR sensor includes an IR emitter for emitting infrared light into the real-world environment. Image sensor 206 may also optionally include one or more cameras configured to capture movement of physical objects in the real-world environment. Image sensor 206 may also optionally include one or more depth sensors configured to detect the distance between the physical object and electronic device 201. In some examples, information from one or more depth sensors may allow the device to identify objects in the real-world environment and distinguish them from other objects in the real-world environment. In some examples, one or more depth sensors may allow the device to determine the texture and / or shape of objects in the real-world environment.
[0029] In some examples, electronic device 201 uses a combination of a CCD sensor, an event camera, and a depth sensor to detect the physical environment surrounding electronic device 201. In some examples, image sensor 206 includes a first image sensor and a second image sensor. The first and second image sensors work cooperatively and are optionally configured to capture different information about physical objects in the real-world environment. In some examples, the first image sensor is a visible light image sensor, and the second image sensor is a depth sensor. In some examples, electronic device 201 uses image sensor 206 to detect the location and orientation of electronic device 201 and / or display generation component 214 in the real-world environment. For example, electronic device 201 uses image sensor 206 to track the location and orientation of display generation component 214 relative to one or more stationary objects in the real-world environment.
[0030] In some examples, electronic device 201 includes microphone 213 or other audio sensors. Electronic device 201 optionally uses microphone 213 to detect sound from the user and / or the user's real-world environment. In some examples, microphone 213 includes a microphone array (multiple microphones) that optionally operates in cooperation, such as to identify ambient noise or locate sound sources in the space of the real-world environment.
[0031] Electronic device 201 includes a position sensor 204 for detecting the position of electronic device 201 and / or display generation component 214. For example, position sensor 204 may include a Global Positioning System (GPS) receiver that receives data from one or more satellites and allows electronic device 201 to determine the absolute position of the device in the physical world.
[0032] Electronic device 201 includes an orientation sensor 210 for detecting the orientation and / or movement of electronic device 201 and / or display generation component 214. For example, electronic device 201 uses orientation sensor 210 to track changes in the positioning and / or orientation of electronic device 201 and / or display generation component 214, such as relative to physical objects in a real-world environment. Orientation sensor 210 optionally includes one or more gyroscopes and / or one or more accelerometers.
[0033] In some examples, electronic device 201 includes a hand tracking sensor 202 and / or an eye tracking sensor 212 and / or other body tracking sensors, such as a leg tracking sensor, a torso tracking sensor, and / or a head tracking sensor. The hand tracking sensor 202 is configured to track the localization / position of one or more portions of a user's hand, and / or the movement of one or more portions of the user's hand relative to the extended reality environment, relative to the display generation component 214, and / or relative to another defined coordinate system. The eye tracking sensor 212 is configured to track the localization and movement of the user's gaze (more generally, the eyes, face, or head) relative to the real world or extended reality environment and / or relative to the display generation component 214. In some examples, the hand tracking sensor 202 and / or the eye tracking sensor 212 are implemented together with the display generation component 214. In some examples, the hand tracking sensor 202 and / or the eye tracking sensor 212 are implemented separately from the display generation component 214.
[0034] In some examples, the hand-tracking sensor 202 (and / or other body-tracking sensors, such as leg, torso, and / or head-tracking sensors) may use an image sensor 206 (e.g., one or more IR cameras, 3D cameras, depth cameras, etc.) that captures 3D information from the real world, including one or more body parts (e.g., a human user's hand, leg, torso). In some examples, sufficient resolution is available to distinguish the hand to differentiate the fingers and their corresponding positions. In some examples, one or more image sensors 206 are positioned relative to the user to define the field of view and interaction space of the image sensors 206, in which the finger / hand positions, orientations, and / or movements captured by the image sensors are used as input (e.g., to differentiate from the user's resting hand or other hands of other people in the real-world environment). Tracking the fingers / hands used for input (e.g., gestures, touches, taps, etc.) may be advantageous because it does not require the user to touch, hold, or wear any type of beacon, sensor, or other marker.
[0035] In some examples, the eye-tracking sensor 212 includes at least one eye-tracking camera (e.g., an infrared (IR) camera) and / or an illumination source (e.g., an IR light source, such as an LED) that emits light toward the user's eyes. The eye-tracking camera may be pointed at the user's eyes to receive reflected IR light from the light source directly or indirectly from the eyes. In some examples, both eyes are tracked separately by the respective eye-tracking cameras and illumination sources, and focus / gaze can be determined by tracking both eyes. In some examples, one eye (e.g., the dominant eye) is tracked by one or more respective eye-tracking cameras / illumination sources.
[0036] Electronic devices 201 are not limited to Figure 2 The electronic device 201 may include, but is not limited to, a few components, other components, or additional components in various configurations. In some examples, electronic device 201 may be implemented between two electronic devices (e.g., as a system). In some such examples, each (or more) electronic device may each include one or more of the same components discussed above, such as various sensors, one or more display generation components, one or more speakers, one or more processors, one or more memories, and / or communication circuitry. One or more persons using electronic device 201 are optionally referred to herein as one or more users of the device.
[0037] As described herein, electronic device 101 (e.g., a first electronic device) may optionally send a request to a second electronic device for knowledge and / or guidance relating to the configuration and / or maintenance of a physical object (or optionally, a virtual object) such as a machine, computing system, consumer electronics device, software program, etc. In some examples, the request is sent from a first user of electronic device 101 to a second user of the second electronic device. In some examples, the first user of electronic device 101 and the second user of the second electronic device are participants in a real-time or near-real-time communication session (e.g., a telephone or video conference) involving the transmission of captured video and / or audio content from one or more corresponding input devices and / or one or more corresponding cameras of electronic device 101 and / or the second electronic device.
[0038] In some examples, the communication session includes displaying and / or otherwise conveying questions related to a physical object via electronic device 101 and / or a second electronic device. In some examples, a second user of the second electronic device is troubleshooting and / or providing instructions to a first user of the first electronic device.
[0039] In some examples, when communicating with a second electronic device, electronic device 101 sends to the second electronic device a portion of its three-dimensional environment, including physical objects, as will be described in more detail below. In some examples, sending this portion of the three-dimensional environment of electronic device 101 does not include sending the entire view of the three-dimensional environment of electronic device 101. In some examples, the computer system does not send the entire view of the three-dimensional environment of electronic device 101 because the user of electronic device 101 has chosen not to share the entire view and / or has chosen to share only a portion of the three-dimensional environment of electronic device 101 (e.g., portions other than the sent portion are private to the first electronic device 101). In some examples, sending this portion of the three-dimensional environment includes an initiation process to cause the second electronic device to display a view of that portion of the three-dimensional environment of electronic device 101, including physical objects. In some examples, the second electronic device presents a view overlaid on a portion of its three-dimensional environment (e.g., the physical environment of the second electronic device), as will be described in more detail below.
[0040] In some examples, when and / or in response to the rendering of a view, the second electronic device renders a representation of a physical object within the three-dimensional environment of the second electronic device (e.g., a virtual three-dimensional model corresponding to the physical object), as will be described in more detail below. In some examples, the view and / or representation of the physical object is presented to a second user of the second electronic device, enabling the second user to view and interact with the representation of the physical object in their three-dimensional environment (e.g., similar to an object in the physical environment of the second user of the second electronic device).
[0041] Additionally or alternatively, electronic device 101 provides a stable view of the physical object to a second user of the second electronic device. In some examples, the stable view may show the physical object as substantially stationary, thereby eliminating the effects of any movement that might cause unstable images in the video. For example, the user of electronic device 101 may move while they are engaged in communication with the second user of the second electronic device, and the stable view presented to the second user of the second electronic device may show the physical object remaining substantially stationary. Thus, in some examples, providing a stable view of the physical object improves the image quality in the video captured and transmitted to the second electronic device. Additionally or alternatively, electronic device 101 and / or the second electronic device present annotations and / or indications of annotations to the physical object, which are presented within the view of the physical object, the representation of the physical object, and / or overlaid on the physical object itself, as will be described in more detail below.
[0042] Figure 3 Example procedures for generating views of objects are illustrated according to some examples of this disclosure. In some examples, the above references Figure 1 and Figure 2 The described electronic device 101 (e.g., a first electronic device) can perform method 300. For example, method 300 includes electronic device 101 defining a bounding box (302) or a two-dimensional container or a three-dimensional volume. In some examples, the bounding box optionally serves as a container for a region within the three-dimensional environment of electronic device 101. This region includes physical objects (e.g., real-world objects in the physical environment of electronic device 101) for the purpose of sharing representations of physical objects. In some examples, electronic device 101 sends a portion of its three-dimensional environment corresponding to this region (e.g., based on the bounding box) to a second electronic device. In some examples, and as described in method 300, electronic device 101 initiates a process to cause the second electronic device to display this portion within the three-dimensional environment of the second electronic device. In some examples, electronic device 101 shares this portion during a communication session with the second electronic device.
[0043] In some examples, defining the bounding box includes via one or more displays (e.g., Figure 1The display 120 presents a representation of a two-dimensional boundary region or a three-dimensional boundary volume to capture a target region including a target physical object. In some examples, the representation of the two-dimensional boundary region includes user interface window elements (e.g., in x and y coordinates) specifying the boundary around the region of the three-dimensional environment of the electronic device 101. In some examples, the representation of the three-dimensional boundary volume includes user interface volume elements (e.g., in x, y, and z coordinates) specifying the boundary around the region of the three-dimensional environment of the electronic device 101. In some examples, the electronic device 101 identifies the region within the three-dimensional environment based on one or more dimensions of the three-dimensional boundary region defined by the two-dimensional boundary region or the three-dimensional boundary volume (e.g., a bounding box or volume). In some examples, the electronic device 101 displays a bounding box having a first area or volume at a first location within the three-dimensional environment via the one or more displays. In some examples, the electronic device 101 detects user input pointing to the bounding box via the one or more input devices to move and / or change the size of the bounding box or volume in one or more dimensions. For example, the user input corresponds to moving the bounding box (e.g., a representation of a two-dimensional boundary region or a three-dimensional boundary volume) within the three-dimensional environment from a first location within the three-dimensional environment to a second location. In some examples, the second location includes a second portion of a physical object that is different from the first portion of the physical object associated with the first location of the bounding box. In some examples, the user input corresponds to a request to increase or decrease the size or volume of the bounding box, rotate the bounding box, or perform other suitable transformations of the bounding box. In some examples, the user input is attention-based gesture input or voice input, or any other input described herein.
[0044] In some examples, electronic device 101 automatically displays a bounding box to include / capture physical objects without detecting user input pointing to the bounding box to include a target area containing the physical object. For example, electronic device 101 uses Object Detection and Tracking (ODT) or other object recognition methods to detect physical objects. In some examples, detecting physical objects includes automatically moving the bounding box and / or adjusting its size to include the detected physical object (e.g., without detecting explicit user input to move and / or adjust the bounding box size). In some examples, electronic device 101 requests user confirmation that the detected object is a target object. For example, electronic device 101 detects user input confirming that the bounding box automatically defined by the electronic device to include the detected object is a target object to be shared with a second electronic device.
[0045] In some examples, after defining the bounding box (and / or simultaneously), the electronic device 101 reprojects (e.g., the corners of the bounding box) onto the camera plane (304). For example, the electronic device 101 transforms the two-dimensional coordinates (or optionally, three-dimensional coordinates) of the bounding box onto the camera plane of the electronic device 101 (e.g., defined by one or more image sensors 206). In some examples, and as referenced above... Figure 2 As described, one or more image sensors 206 are configured to face outwards from a first user to obtain information corresponding to the scene of the electronic device 101 (e.g., a three-dimensional environment including the physical environment). In some examples, the electronic device 101 derives the distance between the one or more image sensors 206 and physical objects based on the size of the bounding box and / or intrinsic and / or extrinsic image sensor parameters. In some examples, the intrinsic parameters of the one or more image sensors include field of view / focal length, sensor size, sensor height, and / or other intrinsic sensor parameters. In some examples, the extrinsic image sensor parameters include the positioning and / or orientation of the one or more image sensors 206 relative to the three-dimensional environment.
[0046] In some examples, the electronic device creates a bounding rectangle (306), or a box, volume, or any other shape. For example, the electronic device 101 captures a portion of a three-dimensional environment corresponding to a target region, which includes target physical objects identified via bounding boxes as described above. In some examples, the bounding rectangle is two-dimensional or three-dimensional. In some examples, the electronic device 101 sets one or more desired parameters for the bounding rectangle, such as one or more edges, alignment, and / or other desired parameters. In some examples, capturing the portion of the three-dimensional environment corresponding to the target region includes generating a two-dimensional bounding region based on the current viewpoint of a first user of the electronic device 101. In some examples, the electronic device 101 defines the rectangle within the image boundary (308) to ensure stable visual output. For example, the electronic device 101 clamps the bounding rectangle to one or more edges of the image boundary. Clamping the bounding rectangle to the one or more edges of the image boundary includes restricting the bounding rectangle within the image boundary so as not to extend beyond the image boundary, which in turn reduces unnecessary computation associated with portions outside the boundary.
[0047] In some examples, electronic device 101 crops the scene camera stream (310) (e.g., one or more image sensor 206 streams). For example, capturing a portion of the 3D environment corresponding to the target region includes cropping a portion of the camera stream captured by the one or more input devices (e.g., one or more image sensors 206) within a two-dimensional boundary region (e.g., the boundary rectangle described above). In some examples, electronic device 101 generates at least two outputs (312): two-dimensional window positioning and orientation information; and a cropped scene camera stream showing only the desired region (314). In some examples, the two-dimensional window positioning and orientation information is based on the aforementioned boundary rectangle. In some examples, the two-dimensional window positioning and orientation information is used to generate an enhanced cropped camera scene, as described in more detail below. In some examples, outputs such as 312 and / or 314 are sent to a second electronic device. In some examples, and as described in more detail below, sending outputs 312 and / or 314 includes initiating a process to cause the second electronic device to display a view of the enhanced camera scene including physical objects and / or a view of the target region including physical objects. In some examples, sending outputs 312 and / or 314 includes initiating a process that causes a second electronic device to display a virtual representation of the physical object. This example process, which provides a stable and / or enhanced view of a region of the physical object and / or (e.g., of the first electronic device) its corresponding three-dimensional environment, offers an efficient way of presenting real-time video (e.g., to the second electronic device) that is consistent and free of motion artifacts (or with reduced amounts of motion artifacts). This provides a smooth and seamless viewing experience for the user, enhances the operability of the electronic device, reduces the power consumption of the electronic device, optimizes bandwidth, reduces video transmission errors, reduces errors in user-electronic device interactions, and reduces the input required to correct such errors.
[0048] Figures 4A to 4L Examples of views and / or virtual representations of real-world objects are illustrated according to some examples of this disclosure. Figure 4A An electronic device 101, or optionally referred to as a first electronic device 101, is illustrated according to some examples of the present disclosure, presenting a computer-generated environment 400 or alternatively referred to as a three-dimensional environment 400 (e.g., extended reality (XR) environment, three-dimensional environment, etc.). Figure 1 Electronic devices 101; and Figure 2The first electronic device 101 (electronic device 201). A computer-generated environment 400 (or optionally, referred to as a three-dimensional environment) is visible from the viewpoint of a first user of the first electronic device 101 (e.g., facing a rear wall and between two walls of the physical environment in which the first electronic device 101 is located). In some examples, the first electronic device 101 is a handheld or mobile device, such as a tablet computer, laptop computer, smartphone, wearable device, or head-mounted display. (See above reference) Figure 2 The architecture block diagram illustrates an example of the first electronic device 101. For example... Figure 4A As shown, the first electronic device 101, window 402, and machine 404 are located in the physical environment of the computer-generated environment 400. In some examples, the first electronic device 101 may be configured to capture areas of the physical environment, including window 402 and machine 404 (e.g., physical objects).
[0049] In some examples, the viewpoint of a first user of the first electronic device 101 determines what is visible in the viewport (e.g., a view of a three-dimensional environment visible to the user via one or more displays such as one or more image sensors 206 or a pair of display modules providing stereoscopic content to different eyes of the same user). In some examples, the (virtual) viewport has a viewport boundary that defines the first user's view via the one or more displays (e.g., Figures 4A to 4LThe viewport (120) defines the extent of the three-dimensional environment visible to the first user. In some examples, the area defined by the viewport boundary is smaller than the first user's visual range in one or more dimensions (e.g., the size, optical properties, or other physical characteristics of the one or more displays and / or the position and / or orientation of the one or more displays relative to the user's eyes, based on the user's visual range). In some examples, the area defined by the viewport boundary is larger than the first user's visual range in one or more dimensions (e.g., the size, optical properties, or other physical characteristics of the one or more displays and / or the position and / or orientation of the one or more displays relative to the first user's eyes, based on the user's visual range). The viewport and viewport boundary typically move with the one or more displays (e.g., with the first user's head moving in a head-mounted device, or with the first user's hand moving in a handheld device such as a tablet or smartphone). The first user's viewpoint determines what is visible in the viewport, typically specifying the position and orientation relative to the three-dimensional environment, and as the viewpoint shifts, the view of the three-dimensional environment also shifts within the viewport. For head-mounted devices, the viewpoint is typically based on the position and orientation of the first user's head, face, and / or eyes to provide a perceptually accurate view of the three-dimensional environment and to provide an immersive experience while the first user is using the head-mounted device. For handheld or fixed devices, the viewpoint shifts as the handheld or fixed device moves and / or as the first user's positioning relative to the handheld or fixed device changes (e.g., the user moves toward the device, away from the device, up, down, to the right and / or left of the device). For a device including one or more displays having video passthrough (or optionally, referred to as visual passthrough), a portion of the physical environment visible (e.g., displayed and / or projected) via the one or more displays is based on the field of view of one or more cameras communicating with the one or more displays, which typically move with the one or more displays (e.g., with the head of a first user in the case of a head-mounted device, or with the hand of a first user in the case of a handheld device such as a tablet or smartphone), because the first user’s viewpoint moves with the field of view of the one or more cameras (and the appearance of one or more virtual objects displayed via the one or more displays is updated based on the first user’s viewpoint (e.g., the display position and pose of the virtual objects are updated based on the movement of the first user’s viewpoint)).For the one or more displays with optical transparency, the portion of the physical environment visible via the one or more display generating components (e.g., optically visible through one or more portions of the display generating components or a completely transparent portion) is based on the field of view of the first user through the portion of the display generating components or the completely transparent portion (e.g., for a head-mounted device, as the first user's head moves, or for a handheld device such as a tablet or smartphone, as the first user's hand moves), because the first user's viewpoint moves with the user's field of view through the portion of the display generating components or the completely transparent portion (and the appearance of one or more virtual objects is updated based on the first user's viewpoint).
[0050] Figures 4A to 4L An example use case is illustrated, in which a first user of a first electronic device 101 initiates a communication session with a customer service representative to obtain customer support related to machine 404 (e.g., referred to as an object or physical object). In some examples, the first electronic device 101 captures the machine for viewing during the communication session with the customer service representative, optionally when the first user initiates the communication session. In some examples, the first electronic device 101 identifies machine 404 and sends the identity of machine 404 to the customer service representative. For example, in Figure 4A In this embodiment, machine 404 includes machine-readable code 406 (e.g., barcode, quick-response (QR) code, displayed characters, or another type of visual pattern including machine-readable information). In some examples, first electronic device 101 detects user input (e.g., air pinch gesture 414) while the attention 412 (e.g., gaze) of the first user of first electronic device 101 is directed to a location corresponding to the machine-readable code 406 of machine 404. In some examples, first electronic device 101 infers that the first user's attention 412 directed to machine-readable code 406 indicates that the first user intends to activate the QR code (e.g., initiate a QR code-related process). In some examples, upon detecting the first user's attention 412 directed to machine-readable code 406 and / or in response to detecting the first user's attention 412 directed to machine-readable code 406, and optionally before detecting air pinch gesture 414, electronic device 101 presents via display 120 an indication 408 that the user's attention 412 is focused on machine-readable code 406. Figure 4AIn this context, indication 408 includes a dashed container or box surrounding machine-readable code 406. In some examples, indication 408 provides a direction for the user's attention to be directed to machine-readable code 406. In some examples, indication 408 informs a first user that machine-readable code 406 can be selected to perform an action associated with machine 404 (e.g., displaying information about machine 404, initiating a communication session with a customer service representative, as described in more detail below, and / or any other action described below).
[0051] Additionally or alternatively, the first electronic device 101 uses optical character recognition (OCR), ODT methods (e.g., as described above), computer vision, and / or other scanning techniques to identify machine 404 and present machine-readable code 406. For example, the first electronic device 101 uses one or more input devices to capture one or more images of machine 404 and uses these images to identify machine 404 in order to retrieve and present machine-readable code 406. Therefore, in some examples, machine 404 does not provide machine-readable code 406 until it detects that a user's attention is directed at machine 404 (e.g., not within the physical environment of the first electronic device 101). In some examples, the first electronic device 101 sends information associated with machine-readable code 406 for lookup in a remote server / database and / or local database (e.g., by the first electronic device 101 through an application running on the first electronic device 101 and / or maintained by a third party communicating with the first electronic device 101 to retrieve customer service information about machine 404). In some examples, the first electronic device 101 retrieves other information about machine 404 (e.g., content, graphics, and / or metadata).
[0052] In some examples, in response to the detection of an air pinch gesture 414 when the first user's attention 412 is directed at machine-readable code 406, such as Figure 4A As shown, the first electronic device 101 displays user interface elements 410a via a display 120, such as... Figure 4B As shown. Additionally or alternatively, electronic device 101 displays user interface elements 410a without detecting user input including air pinch gestures 414 and attention 412. For example, the first electronic device 101 uses one or more input devices (e.g., one or more image sensors 206) to capture one or more images of machine 404 and uses these images to identify machine 404 to retrieve customer service information. Figure 4BIn the interface element 410a, a representation of a first customer service representative 410b associated with machine 404 (e.g., identified via machine-readable code 406) and an option 410c for initiating a communication session (e.g., an audio and / or video communication session) with the first customer service representative. The interface element 410a also includes a representation of a second customer service representative 410d associated with machine 404 and an option 410e for initiating a communication session with the second customer service representative. Figure 4B In this process, the first electronic device 101 detects user input (e.g., a pinch gesture 414 in the air), while the attention 412 (e.g., a gaze) of the first user of the first electronic device 101 is directed to option 410c. In some examples, in response to detecting including Figure 4B User input via air pinch gesture 414 and attention 412 initiates a communication session between the first electronic device 101 and a first customer service representative or a second user, referred to herein as the second electronic device.
[0053] In some examples, initiating a communication session with a second user of a second electronic device involves sharing a stable view (e.g., a substantially static view) of a portion of a three-dimensional environment (e.g., as in reference). Figure 3 (A more detailed description). For example, such as Figure 4C As shown, when initiating a session with a second user of a second electronic device, electronic device 101 displays a representation of the second user 414a (or, optionally, a representation of a first customer service presenter, such as an avatar or 3D avatar). In some examples, first electronic device 101 displays a user interface element 414b including a first option 414c, which, when selected, causes the first electronic device to unshare a view of the 3D environment 400 of first electronic device 101. In some examples, user interface element 414b includes a second option 414d and a third option 414e, which, when selected, causes the first electronic device 101 to initiate sharing of a portion of machine 404; the third option, when selected, causes the first electronic device 101 to display the entire machine 404 (or, optionally, a view of the 3D environment 400 of first electronic device 101 from the viewpoint of the first user of first electronic device 101).
[0054] In some examples, such as Figure 4C As shown, the first electronic device 101 detects user input (e.g., a pinch gesture 414 in the air), while the user's attention 412 (e.g., gazing) is directed to a location corresponding to the second option 414d, indicating a portion of the shared machine 404 (or optionally, a portion of the three-dimensional environment 400 of the first electronic device 101). In some examples, in response to detection Figure 4CUser input is displayed on the first electronic device 101 via the display 120, such as... Figure 4D The user interface element 418a and control user interface element 418b shown are interactive to select and / or define a portion of the three-dimensional environment 400 of the first electronic device 101 for sharing with a second electronic device of a second user. For example, the electronic device 101 may use the control user interface element 418b to increase or decrease the size or volume of a boundary area as indicated by the user interface element 418a. In some examples, the first electronic device 101 detects user input (e.g., including a moving air pinch gesture 414 while the first user's attention 412 (e.g., gaze) is directed at the control user interface element 418), and in response, the first electronic device 101 adjusts the size of the control user interface element 418b (e.g., the boundary area) according to the movement of the air pinch gesture 414.
[0055] In some examples, and such as Figure 4D As shown, a first electronic device 101 displays a user interface element 416a via a display 120. This user interface element includes a first option 416b and a second option 416c. When the first option is selected, the first electronic device 101 cancels sharing a view of its 3D environment 400. When the second option is selected, the first electronic device 101 confirms that the portion indicated by the user interface element 418a is the portion of the 3D environment 400 to be shared with the second electronic device. For example, in... Figure 4D In this process, the first electronic device 101 detects user input (e.g., a pinch gesture 414 in the air, while the attention 412 of the first user (e.g., a gaze) of the first electronic device 101 is directed to the location corresponding to the second option 416c), and in response, the first electronic device 101 shares the selected portion of the three-dimensional environment 400 with the second electronic device 101z (e.g., the electronic device of the first customer service representative), such as Figure 4EAs shown. In some examples, the first electronic device 101 applies visual processing (e.g., blurring effects or other effects described herein) to a second portion of the camera stream captured by the one or more input devices (e.g., a portion other than the portion of the 3D environment corresponding to the area described herein). In some examples, the first electronic device 101 applies visual processing in such a way as to focus on the area and / or prevent accidental display of the second portion. In some examples, the computer system applies visual processing in the manner described herein because the user of the first electronic device 101 chooses not to share the entire view and / or chooses to share only a region of the 3D environment of the first electronic device 101 (e.g., a region other than the selected region is private to the first electronic device 101 and is not shared with or viewable by the second electronic device). In some examples, the first electronic device 101 applies visual processing to the second portion, which is outside the portion of the 3D environment corresponding to the area, before sending the portion of the area to the second electronic device. In some examples, the first electronic device 101 sends the second portion of the 3D environment to the second electronic device after applying visual processing to the second portion.
[0056] Figure 4E The three-dimensional environment 400z of the second user (e.g., the first customer service representative) of the second electronic device 101z is illustrated. Figure 4E As shown, the environment 400z of the second electronic device 101z includes physical objects, such as lamp 436, but in some examples, the three-dimensional environment 400z may be an XR environment without physical objects. Figure 4EIn this context, a portion of the three-dimensional environment 400 of the first electronic device 101 (e.g., selected by the first user of the first electronic device 101 as described above) is presented within the three-dimensional environment 400z of the second electronic device 101z via a window or user interface element 428. In some examples, the first electronic device 101 enhances that portion of the three-dimensional environment before sending it to the second electronic device 101z. For example, the first electronic device 101 may change the brightness level, increase the size of the content, increase the sharpness level, and / or apply other visual processing to increase the readability of the content. In some examples, during a communication session between the second electronic device 101z and the first electronic device 101, the second electronic device 101z displays a representation of the first user 424a of the first electronic device 101 via a display 120z, such as an avatar, a three-dimensional character, or other representation of the first user of the first electronic device 101. In some examples, the second electronic device 101z displays a user interface element 426a, which includes the name or identifier of the first user and / or one or more options 426b, which, when selected, cause the second electronic device 101z to perform an operation associated with a communication session, such as enabling or disabling video during the communication session; enabling or disabling the microphone; ending the communication session; or other operations as described herein. In some examples, when the first electronic device 101 and the second electronic device 101z are in a communication session, the first electronic device 101 displays a representation of the second user 414a of the second electronic device 101z (or optionally referred to as user interface element 414a) via a display 120, as described above. In some examples, the first electronic device 101 displays a user interface element 420a including one or more options 420b via a display 120. In some examples, the user interface element 420a is similar to and / or includes one or more characteristics of the user interface element 426a described above. In some examples, one or more options 420b are similar to and / or include one or more characteristics of the one or more options 426b described above. In some examples, the first electronic device 101 displays via a display 120 a user interface element 422a indicating that the first electronic device 101 is sharing a portion of the three-dimensional environment 400, and an option 422b that, when selected, causes the first electronic device 101 to end or terminate the communication session.
[0057] In some examples, and such as Figure 4E As shown, the first electronic device 101 sends a 3D model corresponding to the machine 404 in the 3D environment 400 of the first electronic device 101 to the second electronic device 101z, so that the second electronic device 101z simultaneously presents the corresponding portion of the 3D environment via the area, such as via user interface element 428. For example, in Figure 4EIn this embodiment, the second electronic device 101z displays a representation 430a (e.g., a 3D model) of the machine 404 via a display 120z. In some examples, the second electronic device 101z displays a representation 430a that receives a 3D model and / or instructions on the 3D model from the first electronic device 101. For example, the second electronic device 101z uses an object recognition technique described above and / or uses machine-readable code 406 to retrieve a computer-aided design (CAD) model of the machine 404 to generate representation 430a, thereby determining the machine 404 within the user interface element 428. In some examples, a second user (e.g., a first customer service representative) of the second electronic device 101z can interact with representation 430a. For example, the second electronic device 101z detects user input 432 (e.g., a pinching gesture as described above) corresponding to a request to zoom in or zoom out on a specific area of representation 430a, and in response, the second electronic device 101z zooms in on representation 430a, such as... Figure 4F As shown, its size is larger than the corresponding size representing 430a before user input 432 is detected, such as Figure 4E As shown in representation 430a. In some examples, the second electronic device 101z displays a second user interface element 430b at the location of representation 430a, which indicates that portion of the machine 404 presented by the first electronic device 101. In some examples, when displaying representation 430a, the second electronic device 101z presents an indication of the position of a first user of the first electronic device 101 relative to the machine 404 (e.g., representation 430a).
[0058] In some examples, the first electronic device 101 and / or the second electronic device 101z present one or more annotations made by the first electronic device 101 and / or the second electronic device 101z to the machine 404 and / or representation 430a. For example, the first electronic device 101 receives from the second electronic device 101z an instruction for input received at the second electronic device 101z, such as input corresponding to, for example, a request to add annotation 434a to representation 430a, as... Figure 4F As shown. In some examples, the first electronic device 101 presents annotation 434c in the three-dimensional environment 400, which corresponds to the portion of the three-dimensional environment corresponding to that area, such as the area shown via user interface element 418a or a physical object (e.g., machine 404) within that portion of the three-dimensional environment 400 corresponding to that area. In some examples, annotation 434c corresponds to input received at the second electronic device, such as for adding... Figure 4F The input of annotation 434a. In some examples, annotations are presented in the corresponding part of the 3D environment via the second electronic device 101z, such as Figure 4FAs shown in note 434b.
[0059] In some examples, the first electronic device 101 and / or the second electronic device 101z present supplementary information associated with machine 404 (e.g., internal wiring and / or circuit content). For example, in Figure 4G In the process of presenting representation 430a, the second electronic device 101z detects user input (e.g., similar to user input 432 or user input 414 as described above) corresponding to a request to view supplementary information associated with machine 404. In some examples, the user input includes moving the second user interface element 430b to capture different portions of representation 430a, such as... Figure 4G As shown, and voice input requesting the presentation of supplementary information. In some examples, in response to detecting user input, the second electronic device 101z presents a user interface element 428 including a representation of supplementary information 440. In some examples, the representation of supplementary information 440 is presented as an overlay on the corresponding portion of representation 430a. In some examples, when presenting the representation of supplementary information, the first electronic device 101 automatically (e.g., without user input) moves user interface element 418a to a position corresponding to the position of the second user interface element 430b. Therefore, in some examples, the first user of the first electronic device 101 knows the specific part of machine 404 that the second user of the second electronic device 101z is viewing.
[0060] In some examples, the first electronic device 101 initially (e.g., at the start of a communication session with a second user of the second electronic device 101z) presents user interface elements 414a (e.g., a representation of the second user 414a), 420a, and / or 422a with orientations oriented towards the user's viewpoint. For example, as Figure 4H As shown, top view 446 includes a first position of machine 404, a first position of first user of first electronic device 101, and a first positioning and / or orientation of user interface element 414a such that the front surface of user interface element 414a faces the user's viewpoint. It should be understood that although the examples described herein relate to user interface element 414a with a first positioning and / or orientation, such functionality and / or characteristics may optionally be applied to other user interface elements, such as... Figure 4H User interface elements 420a and / or 422a.
[0061] In some examples, the first electronic device 101 detects movement of the viewpoint of the first user. For example, in Figure 4IIn this process, the first electronic device 101 detects movement of a first user from a first position 442 to a second position 444. In some examples, in response to detecting movement, and based on determining that the first electronic device 101 is transmitting that portion of the 3D environment according to a first mode (e.g., world-locked mode), the first electronic device 101 maintains the corresponding orientation of the user interface element 414a, such as... Figure 4I As shown in the top view 446. For example, the first electronic device 101 does not change the corresponding orientation of the user interface elements 414a, 420a, and 422a, such that the corresponding orientation of the user interface elements 414a, 420a, and 422a continues to face the corresponding viewpoint of the user at the first position 442. In some examples, in response to detecting a movement of the first user of the first electronic device 101 from the first position 442 to the second position 444, and based on determining that the first electronic device 101 is transmitting this part of the three-dimensional environment according to a second mode different from this mode (e.g., lazy follow mode), the first electronic device 101 presents the user interface elements 414a, 420a, and 422a with the corresponding orientation based on the movement of the first user's viewpoint, as shown in the top view 446. Figure 4J As shown. For example, the first electronic device 101 changes the orientation and / or positioning of the user interface element 414a such that the front surface of the user interface element 414a faces the user's viewpoint, as shown in top view 446. Thus, in some examples, the first electronic device 101 changes the respective orientations of user interface elements 412a, 420a, and 422a to face the first user's viewpoint.
[0062] In some examples, the portion of the three-dimensional environment corresponding to the area sent to the second electronic device 101z for rendering by the second electronic device 101z (such as...) Figure 4EThe user interface element 428 (illustrated in the example) may change or remain unchanged based on the movement of the first user of the first electronic device 101 to ensure a stable presentation. For example, when the first electronic device 101 sends a portion of the three-dimensional environment corresponding to the region to the second electronic device 101z, and in response to detecting movement of the first user of the first electronic device 101 from a first position 442 to a second position 444, and based on determining that the movement meets a movement difference threshold (e.g., 30 degrees, 40 degrees, 50 degrees, 60 degrees, 70 degrees, or 80 degrees), the first electronic device 101 abandons sending a view of the movement of that portion of the three-dimensional environment corresponding to the region to the second electronic device 101z, and presents a notification via one or more displays (e.g., display 120) to recenter the field of view of the first user of the first electronic device 101. In some examples, in response to detecting movement, and based on determining that the movement does not meet a movement difference threshold, the first electronic device 101 sends a view of the movement of that portion of the three-dimensional environment corresponding to the region to the second electronic device 101z, and abandons presenting a notification to recenter the field of view of the first user. In some examples, in response to motion detection, and based on determining that the motion meets a motion difference threshold, the first electronic device 101 presents a previously transmitted portion of the three-dimensional environment corresponding to that area to the second electronic device 101z. In this example, presenting a previously transmitted portion of the three-dimensional environment provides the user with a consistent and seamless viewing experience. In some examples, in response to motion detection, the first electronic device 101 sends a view of the three-dimensional environment corresponding to that area to the second electronic device 101z, based on viewpoint movement.
[0063] In some examples, the first electronic device 101 presents an indication of the location of a second user to the second electronic device 101z. For example, in Figure 4K In this embodiment, the first electronic device 101 presents a second representation of the second electronic device 101z to a second user 448 relative to a physical object (e.g., machine 404) via one or more displays (e.g., display 120). In some examples, the first electronic device 101 detects user input via one or more input devices. In some examples, the user input is similar to user input 432 or user input 414 as described above (or optionally referred to as air pinch gesture 414) and corresponds to a request to present a portion of the three-dimensional environment 400z of the second electronic device 101z from the viewpoint of the second user of the second electronic device 101z. In some examples, in response to detecting user input, the first electronic device 101 presents a user interface element 450a that includes that portion of the three-dimensional environment 400z of the second electronic device 101z as viewed from the viewpoint of the second user of the second electronic device 101z, such as... Figure 4LAs shown. For example, this part includes the back face 450b of the 3D model (e.g., Figure 4G The representation in 430a) and with Figure 4G User interface element 428 corresponds to user interface element 450c. In Figure 4L In the middle, user interface element 450c includes and Figure 4G Supplementary information 440 corresponds to supplementary information 450d.
[0064] In some examples, the first electronic device 101 receives from the second electronic device 101z the movement of a second user relative to a three-dimensional model (e.g., denoted as 430a) (e.g., similar to the movement described above in...). Figure 4H and Figure 4I The first electronic device 101 described herein is used to indicate the movement of a first user. In some examples, in response to receiving an indication of movement, the first electronic device 101 presents a representation of the position of a second user relative to a physical object in a three-dimensional environment 400 via one or more displays of the first electronic device (e.g., display 12), such as presenting a representation of a second user 448 of the second electronic device 101z.
[0065] Figure 5Flowcharts illustrating example processes for sending a portion of a three-dimensional environment to an electronic device, according to some examples of this disclosure, are provided. The devices, methods, and / or computer-readable storage media described below enhance device operability and make user-device interfaces more efficient (e.g., by assisting users in providing correct input and reducing user errors when operating / interacting with the device). This also reduces power consumption and / or extends device battery life by enabling users to use the device faster and more efficiently. Performing operations (such as sending a portion of a three-dimensional environment corresponding to a certain area to another electronic device) without further user input when a set of conditions are met reduces device power consumption, thereby enhancing device operability. Thus, method 500 provides a technical improvement that results in minimizing the amount of data sent between devices while ensuring that the most important and / or most relevant data is prioritized. Furthermore, when sending this portion of the three-dimensional environment over, for example, a bandwidth-limited network, transmission time can be minimized due to the reduced data volume. Therefore, process 500 provides savings in memory, bandwidth, processing, and time. Additionally, Method 500 enhances the AR / VR environment by improving the stability and visibility of real-world content during navigation within the environment's physical space. Method 500 facilitates easier interaction with the environment, provides dynamic content enhancement, incorporates real-time adjustments to maintain the integrity of the AR / VR environment, and ensures a seamless user experience as the user moves within the environment's physical space. In some examples, process 500 begins with interaction with one or more input devices and a second electronic device (e.g., Figure 4E The first electronic device (e.g., the second electronic device 101z) communicates with the second electronic device 101z. Figure 4E The first electronic device 101 in the environment. In some examples, the first electronic component identifies (502) a region within the three-dimensional environment, such as, for example, by... Figure 4D The user interface element 418a captures a region within the three-dimensional environment 400. In some examples, the first electronic device captures (504) a portion of the three-dimensional environment corresponding to an area identified within that three-dimensional environment via one or more input devices, such as... Figure 3 The method discussed in method 300. In some examples, the first electronic device sends (506) the portion of the three-dimensional environment corresponding to that region to the second electronic device, such as via, for example, through Figure 4E The portion illustrated by the user interface element 428 presented by the second electronic device 101z.
[0066] It should be understood that process 500 is an example, and more, fewer, or different operations may be performed in the same or different order. Furthermore, the operations in process 500 described above may optionally be performed by running an information processing device such as a general-purpose processor (e.g., as per [reference to...]). Figure 2 (as described) or one or more functional modules in a dedicated chip and / or by Figure 2 It is achieved through other components.
[0067] Therefore, based on the foregoing, some examples of this disclosure relate to a method comprising, at a first electronic device communicating with one or more input devices and a second electronic device: identifying a region within a three-dimensional environment; capturing a portion of the three-dimensional environment corresponding to the identified region within the three-dimensional environment via the one or more input devices; and transmitting the portion of the three-dimensional environment corresponding to the region to the second electronic device. Additionally or alternatively, the region includes a physical object. Additionally or alternatively, identifying the region within the three-dimensional environment includes presenting a representation of a two-dimensional boundary region or a three-dimensional boundary volume. Additionally or alternatively, identifying the region within the three-dimensional environment includes input for moving the representation of the two-dimensional boundary region or the three-dimensional boundary volume within the three-dimensional environment. Additionally or alternatively, identifying the region within the three-dimensional environment is based on one or more dimensions of the three-dimensional boundary region. Additionally or alternatively, capturing the portion of the three-dimensional environment corresponding to the region includes generating a two-dimensional boundary region based on the current viewpoint of a first user of the first electronic device. Additionally or alternatively, capturing the portion of the three-dimensional environment corresponding to the region includes cropping a portion of a camera stream captured by the one or more input devices within the two-dimensional boundary region.
[0068] Additionally or alternatively, in some examples, the method further includes: detecting movement of the viewpoint of a first user of the first electronic device via the one or more input devices while sending the portion of the three-dimensional environment corresponding to the region to the second electronic device. In some examples, in response to detecting movement, and based on determining that the movement meets a movement difference threshold, the method further includes: abandoning the transmission of a view of the portion of the three-dimensional environment corresponding to the region, based on viewpoint movement, to the second electronic device; and presenting a notification via one or more displays to recenter the first user's field of view. In some examples, in response to detecting movement, and based on determining that the movement does not meet a movement difference threshold, the method further includes: sending a view of the portion of the three-dimensional environment corresponding to the region, based on viewpoint movement, to the second electronic device; and abandoning the presentation of a notification to recenter the first user's field of view.
[0069] Additionally or alternatively, in some examples, the method further includes: transmitting a previously transmitted portion of the three-dimensional environment corresponding to the region to a second electronic device based on determining that the movement satisfies a movement difference threshold. Additionally or alternatively, in some examples, the method further includes: detecting movement of a first user's viewpoint via the one or more input devices while transmitting the portion of the three-dimensional environment corresponding to the region to the second electronic device; and, in response to detecting the movement, transmitting a view of the portion of the three-dimensional environment corresponding to the region, based on the viewpoint movement, to the second electronic device. Additionally or alternatively, in some examples, the method further includes: enhancing the region before transmitting the portion of the three-dimensional environment corresponding to the region to the second electronic device. Additionally or alternatively, in some examples, identifying a region within the three-dimensional environment includes capturing decoded images via the one or more input devices to identify physical objects. Additionally or alternatively, in some examples, identifying a region within the three-dimensional environment is based on a full view of the environment or a partial view of the environment.
[0070] Additionally or alternatively, in some examples, the method further includes: applying visual processing to a second portion of a camera stream captured by the one or more input devices, the second portion being outside the portion of the three-dimensional environment corresponding to the region, before sending the portion of the region to the second electronic device; and sending the second portion of the three-dimensional environment to the second electronic device. Additionally or alternatively, in some examples, the method further includes: presenting a user interface element comprising a representation of a second user of the second electronic device via one or more displays of the first electronic device, wherein the representation has a first orientation in the three-dimensional environment based on the viewpoint of the first user of the first electronic device. In some examples, when presenting the user interface element, the method further includes: detecting movement of the first user's viewpoint via the one or more input devices; and, in response to detecting the movement and based on determining that the first electronic device is sending the portion of the three-dimensional environment according to a first mode, presenting the user interface element in a second orientation based on the movement of the first user's viewpoint. In some examples, in response to detecting the movement and based on determining that the first electronic device is sending the portion of the three-dimensional environment according to a second mode different from the first mode, maintaining the first orientation of the user interface element.
[0071] Additionally or alternatively, in some examples, the method further includes: sending a three-dimensional model corresponding to a physical object in the three-dimensional environment of the first electronic device to a second electronic device for simultaneous presentation via the second electronic device and the portion of the three-dimensional environment corresponding to the region. Additionally or alternatively, in some examples, the method further includes: receiving from the second electronic device an indication of input received at the second electronic device; and presenting an annotation in the three-dimensional environment corresponding to the portion of the three-dimensional environment corresponding to the region or a physical object within the portion of the three-dimensional environment corresponding to the region, wherein the annotation corresponds to input received at the second electronic device. Additionally or alternatively, in some examples, the method further includes: receiving from the second electronic device an indication of movement of a second user of the second electronic device relative to the three-dimensional model; and presenting in the three-dimensional environment a representation of the position of the second user relative to a physical object corresponding to the movement received at the second electronic device.
[0072] Figure 6 Flowcharts illustrating example processes for presenting virtual representations of real-world objects according to some examples of this disclosure are shown. The devices, methods, and / or computer-readable storage media described below enhance device operability and make user-device interfaces more efficient (e.g., by assisting users in providing correct input and reducing user errors when operating / interacting with the device), which in turn reduces power consumption and / or extends device battery life by enabling users to use the device faster and more efficiently. Performing operations (such as by presenting a 3D model corresponding to a physical object in a 3D environment) without further user input when a set of conditions are met reduces device power consumption, thereby enhancing device operability. Thus, method 600 provides a technical improvement that results in providing additional control options (such as by presenting a 3D model) without cluttering the UI with additional display controls, thereby reducing device power consumption and enhancing device operability by reducing unnecessary input and / or navigating between different user interfaces or control groups. In some examples, process 600 begins with one or more input devices and a second electronic device (e.g., Figure 4E The first electronic device (e.g., the second electronic device 101z) communicates with the second electronic device 101z. Figure 4EThe first electronic device 101 in the system. In some examples, the first electronic component identifies (502) a region within the three-dimensional environment, such as a region presented via user interface element 428. In some examples, when presenting the user interface element, the user interface element includes a portion (602) of the three-dimensional environment corresponding to the three-dimensional environment of the second electronic device, and the first electronic device identifies (604) physical objects within that portion of the three-dimensional environment (e.g., via...). Figure 4E The user interface element 428 in the machine 404 is presented; and a three-dimensional model corresponding to the physical object is presented (606) within the three-dimensional environment of the first electronic device, such as Figure 4E The three-dimensional model in (e.g., 430a).
[0073] It should be understood that process 600 is an example, and more, fewer, or different operations may be performed in the same or different order. Additionally, the operations in process 600 described above may optionally be performed by running a general-purpose processor (e.g., as per [reference to...]). Figure 2 (as described) one or more functional modules in an information processing device or dedicated chip and / or by Figure 2 It is implemented using other components.
[0074] Therefore, based on the foregoing, some examples of this disclosure relate to a method comprising, at a first electronic device communicating with one or more input devices and a second electronic device: when presenting a user interface element comprising a portion of a three-dimensional environment corresponding to the three-dimensional environment of the second electronic device: determining a physical object within that portion of the three-dimensional environment; and presenting a three-dimensional model corresponding to the physical object within the three-dimensional environment of the first electronic device. Additionally or alternatively, in some examples, the method further comprises: receiving from the second electronic device an indication of input received at the second electronic device; and applying an annotation to the three-dimensional model, wherein the annotation corresponds to the input received at the second electronic device. Additionally or alternatively, in some examples, the method further comprises: presenting the three-dimensional model from a first viewpoint; detecting input via the one or more input devices while presenting the three-dimensional model from the first viewpoint; and, in response to detecting the input, presenting the three-dimensional model from a second viewpoint different from the first viewpoint. Additionally or alternatively, in some examples, the method further comprises: presenting an indication of the position of the second electronic device relative to the three-dimensional model while presenting the three-dimensional model. Additionally or alternatively, in some examples, the method further includes: detecting input via the one or more input devices when presenting the 3D model; in response to detecting the input: applying an annotation to the 3D model, wherein the annotation corresponds to the input; presenting the annotation via the user interface element in an area corresponding to the physical object; and initiating a process of causing a second electronic device to display the annotation.
[0075] Some examples of this disclosure relate to an electronic device comprising: one or more processors; a memory; and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing any of the methods described above.
[0076] Some examples of this disclosure relate to a non-transitory computer-readable storage medium that stores one or more programs, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform any of the methods described above.
[0077] Some examples of this disclosure relate to an electronic device that includes one or more processors, a memory, and components for performing any of the methods described above.
[0078] Some examples of this disclosure relate to an information processing apparatus used in an electronic device, the information processing apparatus including components for performing any of the methods described above.
[0079] This disclosure envisions that, in some examples, the data utilized may include personal information data that can uniquely identify or be used to contact or locate specific individuals. Such personal information data may include demographic data, content consumption activity, location-based data, telephone numbers, email addresses, Twitter IDs, home addresses, data or records related to a user's health or fitness level (e.g., vital sign measurements, medication information, exercise information), date of birth, or any other identifying information or personal information. Specifically, as described herein, one aspect of this disclosure is the tracking of users' biometric data.
[0080] This disclosure recognizes that the use of such personal information data in the present invention can benefit users. For example, personal information data can be used to display suggested text that changes based on changes in a user's biometric data. For example, suggested text can be updated based on changes in a user's age, height, weight, and / or medical history.
[0081] This disclosure anticipates that entities responsible for the collection, analysis, disclosure, transmission, storage, or other use of such personal information data will comply with robust privacy policies and / or privacy measures. Specifically, such entities should implement and adhere to privacy policies and measures that are recognized as meeting or exceeding industry or governmental requirements for maintaining the privacy and security of personal information data. These policies should be readily accessible to users and should be updated as data collection and / or use change. Personal information from users should be collected for legitimate and reasonable entity purposes and should not be shared or sold outside of these legitimate purposes. Furthermore, such collection / sharing should be conducted only after receiving informed consent from users. Additionally, such entities should consider taking any necessary steps to protect and safeguard the right to access such personal information data and ensure that other entities with access to such personal information data comply with the privacy policies and procedures of those other entities. Additionally, such entities may be subject to third-party assessments to demonstrate their compliance with widely accepted privacy policies and privacy measures. Moreover, policies and measures should be adapted to the specific types of personal information data collected and / or accessed, and to applicable laws and standards, including considerations of specific jurisdictions. For example, in the United States, the collection or acquisition of certain health data may be governed by federal and / or state laws, such as the Health Insurance Portability and Accountability Act (HIPAA); while in other countries, health data may be subject to other regulations and policies and should be handled accordingly. Therefore, different privacy measures should be advocated for different types of personal data in each country.
[0082] Regardless of the foregoing, this disclosure also envisions examples of users selectively blocking the use or access to personal information data. That is, this disclosure contemplates providing hardware and / or software components to prevent or block access to such personal information data. For example, the inventive technology can be configured to allow a user to opt-in or opt-out during or at any time after registering for the service. In another example, a user can choose not to enable the recording of personal information data in a specific application (e.g., a first application and / or a second application). In addition to providing "opt-in" and "opt-out" options, this disclosure also contemplates providing notifications related to access to or use of personal information. For example, a user can be notified that their personal information data will be accessed after collection is initiated, and then reminded again just before the device accesses the personal information data.
[0083] Furthermore, the intent of this disclosure is that personal information data should be managed and processed in a manner that minimizes the risk of unintentional or unauthorized access or use. Once data is no longer needed, this risk can be minimized by restricting data collection and deleting data. Additionally, and where applicable, including in certain health-related applications, data deidentification can be used to protect user privacy. Deidentification can be facilitated, where appropriate, by removing specific identifiers (e.g., date of birth, etc.), controlling the amount or specificity of stored data (e.g., collecting location data at the city level rather than the address level), controlling how data is stored (e.g., aggregating data among users), and / or other methods.
[0084] For purposes of explanation, the foregoing description has been presented with reference to specific examples. However, the illustrative discussion above is not intended to be exhaustive or to limit this disclosure to the precise form disclosed. Many modifications and variations are possible based on the teachings above. The examples were chosen and described to best elucidate the principles of this disclosure and its practical application, thereby enabling others skilled in the art to make optimal use of this disclosure with various modifications suitable for the particular intended purpose, as well as the various described examples.
Claims
1. A method comprising: at a first electronic device in communication with one or more input devices and a second electronic device: identifying an area within a three-dimensional environment; capturing, via the one or more input devices, a portion of the three-dimensional environment corresponding to the area identified within the three-dimensional environment; and sending, to the second electronic device, the portion of the three-dimensional environment corresponding to the area.
2. The method of claim 1, wherein the area comprises a physical object.
3. The method of claim 1, wherein identifying the area within the three-dimensional environment comprises presenting a representation of a two-dimensional boundary area or a three- dimensional boundary volume.
4. The method of claim 3, wherein identifying the area within the three-dimensional environment comprises an input that moves the representation of the two-dimensional boundary area or the three-dimensional boundary volume within the three-dimensional environment.
5. The method of claim 1, wherein identifying the area within the three-dimensional environment is based on one or more dimensions of a three-dimensional boundary area.
6. The method of claim 1, wherein capturing the portion of the three-dimensional environment corresponding to the area comprises generating a two-dimensional boundary area based on a current viewpoint of a first user of the first electronic device.
7. The method of claim 6, wherein capturing the portion of the three-dimensional environment corresponding to the area comprises cropping a portion of a camera stream captured by the one or more input devices that is within the two-dimensional boundary area.
8. The method of claim 1, further comprising: while sending, to the second electronic device, the portion of the three-dimensional environment corresponding to the area, detecting, via the one or more input devices, a movement of a viewpoint of a first user of the first electronic device; and in response to detecting the movement: in accordance with a determination that the movement satisfies a movement difference threshold: forfeit sending, to the second electronic device, a view of the portion of the three- dimensional environment corresponding to the area that is based on the movement of the viewpoint; and present, via one or more displays, a notification to re-center a field of view of the first user; and in accordance with a determination that the movement does not satisfy the movement difference threshold: send, to the second electronic device, the view of the portion of the three- dimensional environment corresponding to the area that is based on the movement of the viewpoint; and forfeit presenting the notification to re-center the field of view of the first user.
9. The method of claim 8, further comprising: in accordance with a determination that the movement satisfies the movement difference threshold, sending, to the second electronic device, a previously sent portion of the three- dimensional environment corresponding to the area.
10. The method of claim 1, further comprising: while sending, to the second electronic device, the portion of the three-dimensional environment corresponding to the area, detecting, via the one or more input devices, a movement of a viewpoint of a first user of the first electronic device; and in response to detecting the movement, sending, to the second electronic device, the view of the portion of the three-dimensional environment corresponding to the region based on the movement of the viewpoint.
11. The method of claim 1, further comprising: enhancing the portion of the region prior to sending, to the second electronic device, the portion of the three-dimensional environment corresponding to the region.
12. The method of claim 1, wherein identifying the region within the three- dimensional environment comprises capturing a transliteration image via the one or more input devices to identify a physical object.
13. The method of claim 1, wherein identifying the region within the three- dimensional environment is based on a full view of the environment or a partial view of the environment.
14. The method of claim 1, further comprising: prior to sending, to the second electronic device, the portion of the region, applying visual processing to a second portion of a camera stream captured by the one or more input devices, the second portion being outside of the portion of the three- dimensional environment corresponding to the region; and sending, to the second electronic device, the second portion of the three- dimensional environment with the visual processing applied.
15. The method of claim 14, further comprising: presenting, via one or more displays of the first electronic device, a user interface element including a representation of a second user of the second electronic device, wherein the representation has a first orientation in the three-dimensional environment based on a viewpoint of a first user of the first electronic device; while presenting the user interface element, detecting, via the one or more input devices, a movement of the viewpoint of the first user; and in response to detecting the movement: in accordance with a determination that the first electronic device is sending the portion of the three-dimensional environment in accordance with a first mode, presenting the user interface element with a second orientation based on the movement of the viewpoint of the first user; and in accordance with a determination that the first electronic device is sending the portion of the three-dimensional environment in accordance with a second mode different from the first mode, maintaining the first orientation of the user interface element.
16. The method of claim 1, further comprising: sending, to the second electronic device, a three-dimensional model corresponding to a physical object in the three-dimensional environment of the first electronic device for presentation via the second electronic device concurrently with the portion of the three-dimensional environment corresponding to the region.
17. The method of claim 1, further comprising: receiving, from the second electronic device, an indication of an input received at the second electronic device; and presenting, in the three-dimensional environment, an annotation corresponding to the portion of the three-dimensional environment corresponding to the region or a physical object within the portion of the three-dimensional environment corresponding to the region, wherein the annotation corresponds to the input received at the second electronic device.
18. The method of claim 1, further comprising: receive, from the second electronic device, an indication of movement of a second user of the second electronic device relative to a three-dimensional model; and present, via one or more displays of the first electronic device, a representation of a location of the second user relative to a physical object in the three-dimensional environment, the physical object corresponding to the movement received at the second electronic device.
19. An electronic device, comprising: one or more processors; memory; and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for performing the method of any of claims 1-18.
20. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by one or more processors of an electronic device, cause the electronic device to perform the method of any of claims 1-18.