Indicate the position of the occluding physical object
Acquisition of data through environmental sensors and generating grids, the problem of computer-generated content that blocks physical objects on the display leads to degradation of user experience, and effective object indication and computing resource savings are achieved.
Patent Information
- Application Number
- CN202210284260.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-26
- Filing Date
- 2022-03-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-22
AI Technical Summary
When displaying computer generated content that obscures physical objects on the display, it is difficult for users to detect the existence of physical objects, resulting in degradation of user experience.
Acquire environmental data through environmental sensors, determine the position value associated with the physical object, identify a part of the computer-generated content that meets the occlusion criteria, and generate a grid associated with the physical object, displaying the grid as an object indicator.
Effectively indicate the obstructed part of the physical object, improve user experience, save computing resources, and prevent users from completely losing awareness of the physical object.
Smart Images

Figure CN115113839B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to display content, and more particularly to displaying indicators associated with physical objects. Background Art
[0002] In some cases, a device displays computer-generated content on a display that obscures at least a first portion of a physical object. The physical object may be of interest to a user of the device. In the case where the computer-generated content does not obscure a second portion of the physical object, the user can view the second portion on the display and thus notice the obscuration. Based on the user viewing the second portion, the user can reposition the computer-generated content or reposition the device so that more of the physical object is viewable. However, the device expends computing resources to reposition the computer-generated content or reposition the device. Additionally, when the computer-generated content obscures the entire physical object, the user experience degrades because the user is completely unaware of the existence of the physical object. Summary of the Invention
[0003] According to some embodiments, a method is performed at an electronic device having one or more processors, non-transitory memory, one or more environmental sensors, and a display. The method includes displaying computer-generated content on the display. The method includes determining a first position value associated with a physical object based on environmental data from the one or more environmental sensors. The method includes identifying a portion of the computer-generated content that satisfies an occlusion criterion relative to a corresponding portion of the physical object based on the first position value. The method includes, in response to identifying that the occlusion criterion is satisfied, generating a grid associated with the physical object based on the first position value, and displaying the grid on the display. The grid overlaps a portion of the computer-generated content.
[0004] According to some embodiments, an electronic device includes one or more processors, non-transitory memory, one or more environmental sensors, and a display. One or more programs are stored in the non-transitory memory and are configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing the performance of the operations of any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of an electronic device, cause the device to perform or cause the performance of the operations of any of the methods described herein. According to some embodiments, an electronic device includes means for performing or causing the performance of the operations of any of the methods described herein. According to some embodiments, an information processing device for use in an electronic device includes means for performing or causing the performance of the operations of any of the methods described herein. Brief Description of the Drawings
[0005] To better understand the various described specific implementations, reference should be made to the following detailed description in conjunction with the accompanying drawings, in which like reference numerals indicate corresponding parts in all the drawings.
[0006] Figure 1 is a block diagram of an example of a portable multifunctional device according to some specific implementations.
[0007] Figures 2A to 2M is an example of an electronic device that displays an object indicator indicating an occluded portion of a physical object according to some specific implementations.
[0008] Figure 3 is an example of a block diagram of a system that generates and displays a mesh corresponding to an occluded portion of a physical object according to some specific implementations.
[0009] Figure 4 is an example of a flowchart of a method for displaying an object indicator indicating an occluded portion of a physical object according to some specific implementations. Detailed Description
[0010] In some cases, the device displays computer-generated content on the display that occludes (e.g., blocks from view) at least a first portion of a physical (e.g., real-world) object. The physical object may be of interest to the user of the device. For example, the physical object is a physical agent, such as an individual walking through the user's physical environment. In the case where the computer-generated content does not occlude a second portion of the physical object, the user can view the second portion on the display and thus notice the occlusion. Accordingly, the user can reposition the computer-generated content so that more of the physical object is viewable. For example, the user can dismiss (e.g., close) a menu interface or move the virtual display screen to a different location within the operating environment. The device utilizes computing resources to reposition the computer-generated content. Additionally or alternatively, the user can reposition the device (e.g., reposition or reorient the device relative to the physical environment), and thus the device has an updated viewing area that includes more portions of the physical object. The device utilizes additional computing resources associated with obtaining and processing position sensor data based on the repositioning of the device. Further, in the case where the computer-generated content occludes the entire physical object for the user, the user experience degrades because the user is completely unaware of the existence of the physical object.
[0011] In contrast, various specific implementations include methods, systems, and electronic devices for displaying an object indicator that indicates a portion of a physical agent occluded by computer-generated content. To this end, the electronic device determines a plurality of position values associated with the physical object based on a function of environmental data (e.g., a combination of image data, depth data, and ambient light data). When displaying the computer-generated content, the electronic device identifies a portion of the computer-generated content that satisfies an occlusion criterion with respect to the corresponding portion of the physical object based on the plurality of position values. For example, the computer-generated content at least partially overlaps the physical object on the display. As another example, the physical object is associated with one or more depth values greater than a first depth value associated with the computer-generated content. The electronic device displays an object indicator that indicates the corresponding portion of the physical object. The object indicator overlaps the portion of the computer-generated content. In some specific implementations, the electronic device displays the object indicator based on a semantic value associated with the physical object, such as when the semantic value indicates a physical agent (e.g., a person, an animal, or a robot) or a predefined (e.g., user-defined) object. For example, the electronic device includes an image sensor that captures image data representing the physical object, and the electronic device performs semantic segmentation with respect to the image data to semantically identify a "person".
[0012] In some specific implementations, the electronic device generates a mesh associated with the physical object based on the plurality of position values, and the electronic device displays the mesh as the object indicator. For example, the mesh is a volume (e.g., three-dimensional (3D)) mesh based on depth values associated with the physical object, where the depth values are indicated in depth data from a depth sensor. In some specific implementations, the electronic device stores the mesh in a non-transitory memory (e.g., a buffer) of the electronic device. The electronic device can retrieve the mesh from the non-transitory memory to synthesize the mesh with a corresponding portion of the environmental data. In other words, the electronic device uses a common memory to store the mesh during mesh generation and for retrieving the mesh during synthesis. Thus, compared to storing the mesh in a first memory during mesh generation, the electronic device synthesizes the mesh with less latency and, at the same time, uses fewer computing resources by copying the mesh from the first memory to a second memory and retrieving the mesh from the second memory during synthesis. Detailed Description
[0014] Reference will now be made in detail to specific implementations, examples of which are illustrated in the accompanying drawings. Many specific details are set forth in the following detailed description in order to provide a thorough understanding of the various described specific implementations. However, it will be apparent to one of ordinary skill in the art that the various described specific implementations may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the specific implementations.
[0015] It will also be understood that, although in some instances the terms "first", "second", etc. are used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first contact could be named a second contact, and similarly, a second contact could be named a first contact, without departing from the scope of the various described specific implementations. The first contact and the second contact are both contacts, but they are not the same contact unless the context clearly indicates otherwise.
[0016] The terms used in the description of the various specific implementations herein are for the purpose of describing particular specific implementations only and are not intended to be limiting. As used in the description of the various specific implementations and in the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. It will also be understood that the term "comprises" ("includes", "including", "comprises", and / or "comprising") when used in this specification is specifying the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0017] As used herein, depending on the context, the term "if" is optionally interpreted to mean "when", "upon", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if it is determined that..." or "if [the stated condition or event] is detected" is optionally interpreted to mean "when it is determined that...", "in response to determining...", "when [the stated condition or event] is detected", or "in response to detecting [the stated condition or event]".
[0018] The physical environment refers to the physical world that people can sense and / or interact with without the help of electronic devices. The physical environment can include physical features such as physical surfaces or physical objects. For example, the physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic device. For example, an XR environment can include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, etc. In the case of an XR system, a subset of a person's physical movements or their representations is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR system are adjusted in a manner that conforms to at least one physical law. For example, an XR system can detect head movements and, in response, adjust the graphical content and sound field presented to a person in a manner similar to how such views and sounds would change in the physical environment. As another example, an XR system can detect the movement of an electronic device (e.g., a mobile phone, a tablet computer, a laptop computer, etc.) presenting the XR environment and, in response, adjust the graphical content and sound field presented to a person in a manner similar to how such views and sounds would change in the physical environment. In some cases (e.g., for accessibility reasons), an XR system can adjust the characteristics of graphical content in the XR environment in response to a representation of physical movement (e.g., a voice command).
[0019] There are many different types of electronic systems that enable a person to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, head-up displays (HUDs), vehicle windshields integrated with display capabilities, windows integrated with display capabilities, displays formed as lenses designed to be placed on a person's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smart phones, tablets, and desktop / laptop computers. A head-mounted system may have an integrated opaque display and one or more speakers. Alternatively, the head-mounted system may be configured to receive an external opaque display (e.g., a smart phone). The head-mounted system may incorporate one or more imaging sensors for capturing images or video of the physical environment and / or one or more microphones for capturing audio of the physical environment. The head-mounted system may have a transparent or translucent display instead of an opaque display. The transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, uLED, liquid crystal on silicon, laser scanning light sources, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical combiner, an optical reflector, or any combination thereof. In some embodiments, the transparent or translucent display may be configured to selectively become opaque. A projection-based system may employ retinal projection techniques that project graphical images onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, such as as a hologram or on a physical surface.
[0020] Figure 1FIG. 0 is a block diagram of an example of a portable multifunctional device 100 (for simplicity, sometimes also referred to herein as “electronic device 100”) in accordance with some embodiments. Electronic device 100 includes a memory 102 (e.g., one or more non-transitory computer-readable storage media), a memory controller 122, one or more processing units (CPUs) 120, a peripheral device interface 118, an input / output (I / O) subsystem 106, a display system 112, an inertial measurement unit (IMU) 130, an image sensor 143 (e.g., a camera), a contact intensity sensor 165, an audio sensor 113 (e.g., a microphone), an eye tracking sensor 164 (e.g., included within a head-mounted device (HMD)), a body tracking sensor 150, and other input or control devices 116. In some embodiments, electronic device 100 corresponds to one of a mobile phone, a tablet computer, a laptop computer, a wearable computing device, a head-mounted device (HMD), a head-mounted housing (e.g., electronic device 100 slides into or otherwise attaches to a head-mounted housing), and the like. In some embodiments, the head-mounted housing is shaped to form a receptacle for receiving an electronic device 100 having a display.
[0021] In some embodiments, peripheral device interface 118, one or more processing units 120, and memory controller 122 are optionally implemented on a single chip such as chip 103. In some other embodiments, they are optionally implemented on separate chips.
[0022] The I / O subsystem 106 couples input / output peripheral devices on the electronic device 100, such as the display system 112 and other input or control devices 116, to the peripheral device interface 118. The I / O subsystem 106 optionally includes a display controller 156, an image sensor controller 158, an intensity sensor controller 159, an audio controller 157, an eye tracking controller 160, one or more input controllers 152 for other input or control devices, an IMU controller 132, a limb tracking controller 180, and a privacy subsystem 170. One or more input controllers 152 receive electrical signals from other input or control devices 116 / send electrical signals to the other input or control devices. The other input control devices 116 optionally include physical buttons (e.g., push buttons, rocker buttons, etc.), dials, slide switches, joysticks, click wheels, etc. In some alternative embodiments, one or more input controllers 152 are optionally coupled (or not coupled) to any of the following: a keyboard, an infrared port, a universal serial bus (USB) port, a stylus, a finger wearable device, and / or a pointing device such as a mouse. One or more buttons optionally include push buttons. In some embodiments, the other input or control devices 116 include a positioning system (e.g., GPS) that obtains information about the position and / or orientation of the electronic device 100 relative to a particular object. In some embodiments, the other input or control devices 116 include a depth sensor and / or a time-of-flight sensor that obtains depth information characterizing physical objects within a physical environment. In some embodiments, the other input or control devices 116 include an ambient light sensor that senses ambient light from the physical environment and outputs corresponding ambient light data.
[0023] The display system 112 provides an input interface and an output interface between the electronic device 100 and the user. The display controller 156 receives electrical signals from the display system 112 and / or sends electrical signals to the display system. The display system 112 displays a visual output to the user. The visual output optionally includes graphics, text, icons, video, and any combination thereof (sometimes referred to herein as "computer-generated content"). In some embodiments, some or all of the visual outputs correspond to user interface objects. As used herein, the term "indicative" refers to a user-interactive graphical user interface object (e.g., a graphical user interface object configured to respond to an input directed to the graphical user interface object). Examples of user-interactive graphical user interface objects include, but are not limited to, buttons, sliders, icons, selectable menu items, switches, hyperlinks, or other user interface controls.
[0024] The display system 112 may have a touch-sensitive surface, sensors, or a group of sensors that accept input from a user based on tactile and / or haptic contact. The display system 112 and the display controller 156 (along with any associated modules and / or instruction sets in the memory 102) detect contact (and any movement or interruption of that contact) on the display system 112 and convert the detected contact into an interaction with user interface objects (e.g., one or more soft keys, icons, web pages, or images) displayed on the display system 112. In an exemplary embodiment, the point of contact between the display system 112 and the user corresponds to the user's finger or a finger-worn device.
[0025] The display system 112 optionally uses LCD (Liquid Crystal Display) technology, LPD (Light-Emitting Polymer Display) technology, or LED (Light-Emitting Diode) technology, but uses other display technologies in other embodiments. The display system 112 and the display controller 156 optionally use any of a variety of touch-sensing technologies now known or later developed, as well as other proximity sensor arrays or other elements for determining one or more points of contact with the display system 112, the variety of touch-sensing technologies including but not limited to capacitive technology, resistive technology, infrared technology, and surface acoustic wave technology.
[0026] The user optionally uses any suitable object or attachment such as a stylus, a finger-worn device, a finger, etc. to contact the display system 112. In some embodiments, the user interface is designed to work with finger-based contact and gestures, which may not be as precise as stylus-based input due to the larger contact area of the finger on the touch screen. In some embodiments, the electronic device 100 converts finger-based rough input into an accurate pointer / cursor position or command for performing the action desired by the user.
[0027] The audio circuit also receives an electrical signal converted from sound waves by the audio sensor 113 (e.g., a microphone). The audio circuit converts the electrical signal into audio data and transmits the audio data to the peripheral device interface 118 for processing. The audio data is optionally retrieved from and / or transmitted to the memory 102 and / or the RF circuit by the peripheral device interface 118. In some embodiments, the audio circuit also includes a headphone jack. The headphone jack provides an interface between the audio circuit and a removable audio input / output peripheral device, the removable audio input / output peripheral device being such as an output-only headset or a headset having both an output (e.g., a mono or stereo headset) and an input (e.g., a microphone).
[0028] The inertial measurement unit (IMU) 130 includes an accelerometer, a gyroscope, and / or a magnetometer to measure various forces, angular rates, and / or magnetic field information relative to the electronic device 100. Accordingly, in various embodiments, the IMU 130 detects one or more position change inputs of the electronic device 100, such as the electronic device 100 being shaken, rotated, moved in a particular direction, etc.
[0029] The image sensor 143 captures still images and / or videos. In some embodiments, the image sensor 143 is located on the back surface of the electronic device 100, opposite the touch screen on the front surface of the electronic device 100, such that the touch screen can be used as a viewfinder for still image and / or video image capture. In some embodiments, another image sensor 143 is located on the front surface of the electronic device 100 such that an image of the user is obtained (e.g., for selfies, for video conferencing when the user is viewing other video conferencing participants on the touch screen, etc.). In some embodiments, the image sensor is integrated within the HMD. For example, the image sensor 143 outputs image data representing physical objects (e.g., physical agents) within the physical environment.
[0030] The contact intensity sensor 165 detects the intensity of contacts on the electronic device 100 (e.g., touch inputs on the touch-sensitive surface of the electronic device 100). The contact intensity sensor 165 is coupled to the intensity sensor controller 159 in the I / O subsystem 106. The contact intensity sensor 165 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-mechanical force sensors, piezoelectric force sensors, optical force sensors, capacitive touch-sensitive surfaces, or other intensity sensors (e.g., sensors for measuring the force (or pressure) of a contact on a touch-sensitive surface). The contact intensity sensor 165 receives contact intensity information (e.g., pressure information or a surrogate for pressure information) from the physical environment. In some embodiments, at least one contact intensity sensor 165 is juxtaposed or adjacent to the touch-sensitive surface of the electronic device 100. In some embodiments, at least one contact intensity sensor 165 is located on the side surface of the electronic device 100.
[0031] The eye tracking sensor 164 detects the eye gaze of a user of the electronic device 100 and generates eye tracking data indicative of the user's gaze position. In various embodiments, the eye tracking data includes data indicative of a fixation point (e.g., a point of interest) of the user on a display panel, such as a display panel within a head-mounted device (HMD), a head-mounted housing, or a heads-up display.
[0032] The limb tracking sensor 150 obtains limb tracking data indicating the position of a user's limb. For example, in some embodiments, the limb tracking sensor 150 corresponds to a hand tracking sensor that obtains hand tracking data indicating the position of a user's hand or finger within a particular object. In some embodiments, the limb tracking sensor 150 utilizes computer vision techniques to estimate the pose of a limb based on camera images.
[0033] In various embodiments, the electronic device 100 includes a privacy subsystem 170 that includes one or more privacy setting filters associated with user information, such as user information included in limb tracking data, eye gaze data, and / or body position data associated with the user. In some embodiments, the privacy subsystem 170 selectively prevents and / or restricts the electronic device 100 or portions thereof from obtaining and / or transmitting user information. To this end, the privacy subsystem 170 receives user preferences and / or selections from the user in response to prompting the user for user preferences and / or selections. In some embodiments, the privacy subsystem 170 prevents the electronic device 100 from obtaining and / or transmitting user information unless and until the privacy subsystem 170 obtains informed consent from the user. In some embodiments, the privacy subsystem 170 anonymizes (e.g., scrambles or obfuscates) certain types of user information. For example, the privacy subsystem 170 receives user input specifying which types of user information the privacy subsystem 170 anonymizes. As another example, the privacy subsystem 170 anonymizes (e.g., automatically) certain types of user information that may include sensitive and / or identifying information independent of user specification.
[0034] Figures 2A to 2M is an example of an electronic device 210 that displays an object indicator indicating an occluded portion of a physical object according to some embodiments. In various embodiments, the electronic device 210 is similar to and modified from the Figure 1 illustrated electronic device 100. The electronic device 210 is associated with (e.g., located within) a physical environment 200. The physical environment 200 includes a first wall 202, a second wall 204, and a physical cabinet 220 disposed against the first wall 202. As Figure 2A illustrated, a user 50 holds the electronic device 210 facing a portion of the physical environment 200.
[0035] The electronic device 210 includes a display 212. The display 212 is associated with a visual region 214 that includes a portion of the physical environment 200. As Figure 2A illustrated, the visual region 214 includes the physical cabinet 220, a portion of the first wall 202, and a portion of the second wall 204. The visual region 214 is a function of the position or orientation of the electronic device 210 relative to the physical environment 200.
[0036] In some specific implementations, the electronic device 210 corresponds to a mobile device including a display. For example, the electronic device 210 corresponds to a smart phone, a tablet computer, or a laptop computer. In some specific implementations, the electronic device 210 corresponds to a wearable device (such as a smart watch or a head-mounted device (HMD)) including an integrated display (e.g., a built-in display) for displaying a representation of the physical environment 200. In some specific implementations, the electronic device 210 includes a head-mounted housing. The head-mounted housing may include an attachment area to which another device having a display may be attached. The head-mounted housing may be shaped to form a receiver for receiving another device (e.g., the electronic device 210) including a display. For example, in some specific implementations, the electronic device 210 slides / snaps into the head-mounted housing or is otherwise attached to the head-mounted housing. In some specific implementations, the display of the device attached to the head-mounted housing presents (e.g., displays) a representation of the physical environment 200. For example, in some specific implementations, the electronic device 210 corresponds to a mobile phone attachable to the head-mounted housing.
[0037] The electronic device 210 includes one or more environmental sensors whose outputs characterize corresponding sensor data of the physical environment 200. For example, in some specific implementations, the environmental sensors include image sensors, such as scene cameras. In some specific implementations, the image sensors obtain image data characterizing the physical environment 200, and the electronic device 210 synthesizes the image data with computer-generated content to generate display data for display on the display 212. The display data may be characterized by an XR environment.
[0038] In some specific implementations, the display 212 corresponds to a see-through display. The see-through display allows ambient light from the physical environment 200 to enter and pass through the see-through display for display. For example, the see-through display is a semi-transparent display, such as glass with optical see-through. In some specific implementations, the see-through display is an additional display allowing optical see-through of the physical environment 200, such as an optical head-mounted display (OHMD). Different from the specific implementations including synthetic image data, the additional display is capable of reflecting a projected image from the display 212 while enabling the user's vision to pass through the display 212. The HMD may add computer-generated content to the ambient light entering the see-through display to enable display of the physical environment 200. In some specific implementations, the see-through display includes electrochromic lenses.
[0039] As Figure 2BAs shown, the electronic device 210 displays computer-generated content on the display 212. That is, the computer-generated content includes a computer-generated view screen 230, which itself includes a computer-generated dog 232. Those of ordinary skill in the art will understand that the computer-generated content can correspond to any type of content, such as a menu (e.g., a head-up display (HUD)), three-dimensional (3D) content (e.g., a virtual couch), etc. In some specific implementations, the electronic device 210 stores and retrieves computer-generated content from a local memory (e.g., a cache or RAM). In some specific implementations, the electronic device 210 obtains computer-generated content from another system, such as by downloading computer-generated content from a server.
[0040] In some specific implementations, as Figure 2C shown, the electronic device 210 includes an eye tracker 216 that tracks the eye gaze of the user 50. To this end, the eye tracker 216 outputs eye tracking data 218 indicating the gaze position of the user 50. The gaze position is directed to the computer-generated dog 232, as indicated by the first marking line 219. Determining whether and how to display an object indicator using the gaze position is described below.
[0041] As Figure 2D shown, a physical agent 240 (e.g., an individual) moves within the physical environment 200, such as when the physical agent 240 enters the room where the user 50 is located. Notably, a portion of the physical agent 240 is within the visible area 214 such that the electronic device 210 can detect the physical agent 240.
[0042] Based on the environmental data 244, the electronic device 210 determines a plurality of position values 248 associated with a physical object via the object tracker 242. The electronic device 210 includes one or more environmental sensors that output the environmental data 244. For example, one or more environmental sensors include an image sensor that senses environmental light from the physical environment 200 and outputs corresponding image data representing the physical environment 200. The electronic device 210 can identify the plurality of position values 248 based on the image data, such as through computer vision techniques. For example, the electronic device 210 identifies a first pixel value corresponding to a first part of the physical object and a second pixel value corresponding to a second part of the physical object. In some specific implementations, the environmental data 244 includes a combination of image data and depth sensor data.
[0043] As another example, referring to Figure 2D , a plurality of position values 248 are associated with the physical agent 240. The plurality of position values 248 indicate corresponding multiple positions of the physical agent 240. For example, the first position value among the plurality of position values 248 indicates the first position of the right hand of the physical agent 240, where the first position is determined by Figure 2Dis indicated by the second marking line 246 therein. Other position values indicate other parts of the physical agent 240, such as the left hand, abdomen, legs, head, face, etc. For clarity and brevity, illustrations of the marking lines associated with other position values are omitted from Figures 2D to 2M In some embodiments, the electronic device 210 includes an image sensor that outputs image data representing the physical environment 200, and the object tracker 242 identifies (e.g., via computer vision techniques) a subset of pixels of the image data corresponding to parts of the physical agent 240. For example, each of the pixels in the subset corresponds to a respective one of the plurality of position values 248. In some embodiments, the object tracker 242 semantically identifies (e.g., by semantic segmentation) one or more parts of the physical agent 240 to determine one or more corresponding semantic values for "head", "left hand", "right hand", "face", etc. In some embodiments, the electronic device 210 includes a depth sensor that outputs depth data associated with parts of the physical agent 240, wherein the plurality of position values 248 includes one or more depth values indicated within the depth data. In some embodiments, the electronic device 210 uses data from a plurality of different environmental sensors (e.g., image sensor, depth sensor, ambient light sensor) to determine the plurality of position values 248.
[0044] The electronic device 210 includes a content recognizer 250 that identifies, based on the plurality of position values 248, a portion of computer-generated content that satisfies an occlusion criterion with respect to a corresponding part of the physical object. For example, referring to Figure 2D , the content recognizer 250 determines that the part of the physical agent 240 and the computer-generated content do not overlap with each other in Figure 2D and thus do not satisfy the occlusion criterion. In other words, the content recognizer 250 has not identified any portion of the computer-generated content that satisfies the occlusion criterion. As Figure 2D further shown, as indicated by the first movement line 260, the physical agent 240 starts moving left across the physical environment 200 in front of the physical cabinet 220.
[0045] Based on the movement of the physical agent 240, as Figure 2E shown, the computer-generated viewing screen 230 occludes the corresponding part (e.g., above the legs) of the physical agent 240. Accordingly, the corresponding part of the physical agent 240 is not displayed. Based on the updated plurality of position values 248 (as Figure 2Eas shown by the second fiducial line 246 in), the content recognizer 250 recognizes a portion of computer-generated content that meets an occlusion criterion relative to a corresponding portion of the physical agent 240. To this end, in some embodiments, the content recognizer 250 recognizes an overlapping region where a portion of the computer-generated content overlaps with the corresponding portion of the physical agent 240, as represented within the image data. The content recognizer 250 identifies the overlapping region as the portion of the computer-generated content that meets the occlusion criterion. In some embodiments, when more than a threshold number of pixels of the computer-generated content overlap with the corresponding portion of the physical agent 240, the content recognizer 250 determines that the portion of the computer-generated content meets the occlusion criterion. As another example, in some embodiments, in combination with identifying the overlapping region based on the image data, the content recognizer 250 uses one or more depth values to determine whether the occlusion criterion is met. Continuing this example, the content recognizer 250 determines that the overlapping portion of the computer-generated content meets the occlusion criterion because one or more depth values are greater than a first depth value associated with the overlapping portion of the computer-generated content. Although below the computer-generated content in Figure 2E a second portion (e.g., a leg) of the physical agent 240 may be visible, those of ordinary skill in the art will understand that in some embodiments, the physical agent 240 is fully occluded and the content recognizer 250 recognizes a portion of the computer-generated content that meets the occlusion criterion.
[0046] In addition, as Figure 2E shown, a distance line 251 indicates the distance between the fixation position (indicated by the first fiducial line 219) and the tracked portion of the physical agent 240 (indicated by the second fiducial line 246). In some embodiments, the electronic device 210 uses the distance to determine whether to display an object indicator, as described below.
[0047] The electronic device 210 displays an object indicator indicative of the corresponding portion of the physical agent 240. The object indicator overlaps with the portion of the computer-generated content that meets the occlusion criterion. By displaying the object indicator, the electronic device 210 provides the user 50 with an indicator of the location of the occluded portion of the physical agent 240. Accordingly, the user experience is enhanced and the electronic device 210 saves computational (e.g., processor and sensor) resources because the user 50 does not need to reposition the computer-generated content or the electronic device 210 to make the occluded portion of the physical agent 240 visible. According to various embodiments, the object indicator may be displayed in various ways.
[0048] For example, as Figure 2FAs shown, the electronic device 210 displays on the display 212 a contour 262 corresponding to the object indicator. The contour 262 is overlaid on the corresponding part of the physical agent 240. To this end, the electronic device 210 determines the contour 262 based on a plurality of position values 248. For example, the contour 262 indicates the outer boundary (e.g., silhouette) of the corresponding part of the physical agent 240. In some embodiments, the electronic device 210 synthesizes the contour 262 with a portion of the environmental data 244 associated with the corresponding part of the physical agent 240 (e.g., a subset of pixels of the image data).
[0049] According to various embodiments, the object indicator corresponds to a grid associated with the physical agent 240. The grid can have various textures. In some embodiments, the texture of the grid enables the user 50 to partially or fully view the physical agent 240 through the computer-generated viewing screen 230. For example, referring to Figure 2G , the electronic device 210 generates and displays a grid 263 that is slightly larger than the outer boundary of the physical agent 240. The electronic device 210 generates the grid 263 based on at least a portion of the plurality of position values 248. The grid 263 can include elongated bubbles, indicated by the dashed contour in Figure 2G . When synthesized with the corresponding part of the computer-generated viewing screen 230, the grid 263 enables the physical agent 240 to be viewable on the display 212, as shown in Figure 2G . In other words, the grid 263 enables penetration of the corresponding part of the computer-generated viewing screen 230, such that the user 50 can see the true physical representation of the physical agent 240 through the computer-generated viewing screen 230. In some embodiments, the electronic device 210 generates the grid 263 in response to determining that the physical agent 240 attempts to gain the attention of the user 50, e.g., when the physical agent 240 faces the user 50 and the arm of the physical agent 240 waves at the user 50.
[0050] In some embodiments, the grid indicates a shadow associated with the physical agent 240. For example, referring to Figure 2H , the electronic device 210 generates and displays a grid 264 that indicates a shadow associated with the physical agent 240. In some embodiments, the electronic device 210 generates the grid 264 that indicates the shadow (instead of penetration - e.g., via the grid 263) based on determining that the physical agent 240 is behind the computer-generated viewing screen 230 but does not attempt to attract the attention of the user 50. For example, the physical agent 240 has its back to the user 50, moves away from the user 50, and / or generally moves less than a threshold amount.
[0051] The grid 264 is associated with corresponding portions of the physical agent 240. The grid 264 has a cross-hatch pattern to distinguish the grid 264 from the background of the computer-generated viewing screen 230. In some embodiments, the grid 264 represents corresponding portions of the physical agent 240 as a collection of discrete geometric and topological elements. The electronic device 210 generates the grid 264 based on a plurality of position values 248. The grid 264 can be two-dimensional (2D), such as when the grid 264 includes a combination of triangular and quadrilateral elements. For example, the electronic device 210 determines a 2D grid based on image data representing corresponding portions of the physical agent 240.
[0052] The grid 263 and / or the grid 264 can be volumetric (e.g., three-dimensional (3D)), such as when the grid includes a combination of tetrahedra, quadrilateral pyramids, triangular prisms, and hexahedral elements. For example, the plurality of position values 248 includes a plurality of depth values, and the electronic device 210 determines a volumetric grid based on the plurality of depth values. In some embodiments, the electronic device 210 determines a volumetric grid based on point cloud data associated with corresponding portions of the physical agent 240.
[0053] In some embodiments, the electronic device 210 displays an object indicator in response to determining that a gaze position (indicated by the first fiducial 219) satisfies a proximity threshold relative to a portion of the computer-generated content. For example, referring to Figure 2F and Figure 2H , the electronic device 210 determines that the gaze position satisfies the proximity threshold because the distance between the gaze position and a portion of the physical agent 240 (indicated by the second fiducial 246) is less than a threshold distance. Accordingly, the electronic device 210 displays an object indicator, such as the contour 262 in Figure 2F or the grid 264 in Figure 2H . As a counterexample, referring to Figure 2M , the electronic device 210 stops displaying the object indicator (grid 264) because the gaze position stops satisfying the proximity threshold due to the physical agent 240 moving away from the gaze position - e.g., the distance between the gaze position and the physical agent 240 is greater than the threshold distance.
[0054] As Figure 2I shown, the physical agent 240 rotates so as to face away from the second wall 204, as indicated by the rotation indicator 266. Based on the rotation, the object tracker 242 determines updated plurality of position values 248 based on the environmental data 244, and the electronic device 210 updates the grid 264 accordingly. That is, as Figure 2J shown, the updated grid 264 has an orientation that reflects the rotation of the physical agent 240.
[0055] As Figure 2KAs shown, the physical agent 240 begins to move away from the second wall 204, as indicated by the second movement line 268. As the physical agent 240 moves, the object tracker 242 determines updated position values 248 based on the environmental data 244, and the electronic device 210 updates the grid 264 accordingly, as Figure 2L shown.
[0056] As Figure 2M shown, the physical agent 240 finishes moving, and the electronic device 210 determines updated position values 248 based on the environmental data 244. In some embodiments, the electronic device 210 determines that the gaze position does not meet the proximity threshold because the distance between the gaze position and a portion of the physical agent 240 is greater than a threshold amount. Thus, as Figure 2M shown, the electronic device 210 stops displaying the grid 264. Thus, in some cases, the electronic device 210 saves processing resources by not continuously displaying the object indicator (e.g., the grid 264). Another example of resource savings is shown in Figure 2D wherein, although the object tracker 242 identifies a portion of the physical agent 240, the electronic device 210 does not display the object indicator because the occlusion criterion is not met.
[0057] Figure 3 is an example of a block diagram of a system 300 for generating and displaying a grid corresponding to an occluded portion of a physical object according to some embodiments. Although relevant features are shown, those of ordinary skill in the art will recognize from this disclosure that, for the sake of brevity and to not obscure more relevant aspects of the exemplary embodiments disclosed herein, various other features are not shown. In some embodiments, the system 300 or portions thereof are integrated in an electronic device, such as the electronic device 210 described with reference to Figures 2A to 2M described.
[0058] To track a physical object, system 300 includes an object tracker 242. The object tracker 242 determines a plurality of position values 248 associated with the physical object based on a function of environmental data 244. To this end, system 300 includes one or more environmental sensors 310 that output the environmental data 244. For example, the environmental sensors 310 include a combination of an image sensor (e.g., a camera) 312, a depth sensor 314, and an ambient light sensor 316. For example, the image sensor 312 outputs one or more images of the physical environment. Continuing this example, the object tracker 242 identifies a first set of one or more pixels associated with a first part of the physical object, a second set of one or more pixels associated with a second part of the physical object, etc. within the image. The object tracker 242 can utilize computer vision techniques (optionally with the aid of a neural network) to identify the pixels of the image, such as instance segmentation or semantic segmentation. As another example, the depth sensor 314 outputs depth data associated with the physical environment. Continuing this example, based on the depth data, the object tracker 242 identifies a first depth value associated with a first part of the physical object, a second depth value associated with a second part of the physical object, etc. As yet another example, the ambient light sensor 316 senses the ambient light reflected from the physical object and outputs corresponding ambient light data. Continuing this example, based on the ambient light data, the object tracker 242 identifies a first ambient light value associated with a first part of the physical object, a second ambient light value associated with a second part of the physical object, etc.
[0059] In some embodiments, the plurality of position values 248 include position values based on data from different environmental sensors. For example, the object tracker 242 determines a first position value among the plurality of position values 248 by identifying a first set of one or more pixels of the image data from the image sensor 312, and determines a second position value among the plurality of position values 248 by identifying a first depth value based on the depth data from the depth sensor 314. Continuing this example, the first position value among the plurality of position values 248 and the second position value among the plurality of position values 248 may be associated with the same or different parts of the physical object.
[0060] System 300 includes a content recognizer 250. The content recognizer 250 identifies a portion of computer-generated content 320 that satisfies an occlusion criterion 340 with respect to a corresponding part of the physical object based on the plurality of position values 248. For example, referring to Figure 2B , the computer-generated content 320 includes a computer-generated viewing scene 230 that includes a computer-generated dog 232. As another example, the computer-generated content may be characterized by a static image or a video stream (e.g., a series of images) and may be 2D or 3D content.
[0061] In some specific implementations, the occlusion criterion 340 is based on the opacity of the computer-generated content 320. For example, the content recognizer 250 identifies a portion of the computer-generated content 320 that partially meets the occlusion criterion 340 because the portion of the computer-generated content 320 is associated with an opacity characteristic that exceeds a threshold, such as when a physical object cannot be viewed through the portion of the computer-generated content 320.
[0062] In some specific implementations, the occlusion criterion 340 is based on the position of the computer-generated content 320 relative to a physical object. For example, in some specific implementations, the content recognizer 250 identifies a portion of the computer-generated content 320 that at least partially overlaps a corresponding portion of the physical object. As an example, as Figure 2E shown, the content recognizer 250 identifies a portion of the computer-generated viewing frame 230 that overlaps a corresponding portion (e.g., above the legs) of the physical agent 240 based on a plurality of position values 248. Continuing with this example, the plurality of position values 248 can include a set of pixels from an image of the physical agent 240 from the image sensor that represent the corresponding portion.
[0063] In some specific implementations, the occlusion criterion 340 is based on the depth of the computer-generated content 320 relative to the depth of the physical object. As an example, as Figure 2E shown, the plurality of position values 248 includes one or more depth values associated with a portion of the physical agent 240. Accordingly, the content recognizer 250 identifies a portion of the computer-generated viewing frame 230 that meets the occlusion criterion 340 because the portion of the computer-generated viewing frame 230 is associated with a first depth value that is less than one or more depth values associated with the portion of the physical agent 240. In other words, the portion of the computer-generated viewing frame 230 is closer to the electronic device 210 than the portion of the physical agent 240.
[0064] On the other hand, at a given time, the depth value associated with a portion of the physical agent can be less than the depth value associated with the computer-generated viewing frame. For example, from the perspective of a user viewing the display, the portion of the physical agent is in front of the computer-generated viewing frame. In some specific implementations, based on the physical agent in front of the computer-generated viewing frame, the electronic device displays the portion of the physical agent that occludes the corresponding portion of the computer-generated viewing frame. Additionally, the electronic device can generate a grid without elongated bubbles, where the grid follows the silhouette (e.g., profile) of the physical agent. Thus, the grid can be used to make the corresponding portion of the computer-generated viewing frame transparent, such that the physical agent appears to truly occlude the corresponding portion of the computer-generated viewing frame.
[0065] In some specific implementations, the occlusion criterion 340 is based on a combination of relative position and relative depth. For example, referring to Figure 2D , although the first depth value associated with the computer-generated content is less than one or more depth values associated with a portion of the physical agent 240, the occlusion criterion 340 is not satisfied because there is no positional overlap between the computer-generated content and the portion of the physical agent 240. As a counterexample, referring to Figure 2E , the occlusion criterion 340 is satisfied because the computer-generated content overlaps with the portion of the physical agent 240.
[0066] In some specific implementations, in response to identifying a portion of the computer-generated content 320 that satisfies the occlusion criterion 340, the content recognizer 250 instructs the grid generator 330 to generate a grid. The grid indicates the corresponding portion of the physical object, and the grid overlaps with the portion of the computer-generated content. For example, the grid corresponds to the grid 263 described in reference Figure 2G or the grid 264 described in reference Figures 2H to 2M .
[0067] In some specific implementations, the grid generator 330 selectively generates a grid. For example, when the content recognizer 250 does not identify a portion of the computer-generated content 320 that satisfies the occlusion criterion 340 (e.g., as shown in Figure 2D ), the content recognizer 250 instructs the grid generator 330 not to generate a grid. As another example, in a specific implementation that includes an eye tracker 216 (e.g., described in reference Figures 2A to 2M ), the generation of the grid can be a function of the eye tracking data from the eye tracker 216. For example, based on the fixation position indicated in the eye tracking data, the grid generator 330 determines whether the fixation position satisfies a proximity threshold relative to the identified portion of the computer-generated content 320. In some specific implementations, the grid generator 330 generates a grid in response to determining that the fixation position satisfies the proximity threshold (e.g., as shown in Figure 2L ), and abandons generating a grid in response to determining that the fixation position does not satisfy the proximity threshold (e.g., as shown in Figure 2M ). Thus, the system 300 reduces resource utilization by avoiding continuously generating grids.
[0068] In some specific implementations, the system 300 includes a buffer 332 for storing the grid. To this end, the system 300 stores the grid from the grid generator 330 in the buffer 332. The buffer 332 can correspond to one of a cache, random access memory (RAM), etc. For example, the electronic device 210 includes a non-transitory memory corresponding to the buffer 332, and the electronic device 210 stores the grid 264 in the non-transitory memory.
[0069] In some embodiments, system 300 includes a compositing subsystem 350. According to various embodiments, system 300 provides a mesh from buffer 332 to the compositing subsystem 350, which in turn composites the mesh with environmental data 244 and computer-generated 320. For example, in response to receiving a mesh request from the compositing subsystem 350, system 300 retrieves the mesh from buffer 332 and provides the mesh to the compositing subsystem 350. The request can originate from content recognizer 250 when the content recognizer 250 identifies a portion of computer-generated content 320 that meets occlusion criterion 340. System 300 provides the output of the compositing subsystem 350 to display 212 for display. Thus, system 300 uses a common memory (e.g., buffer 332) for mesh storage and mesh retrieval during compositing. Compared to storing the mesh in a first memory during mesh generation, the compositing subsystem 350 thus composites the mesh with less latency and, at the same time, uses fewer computing resources by copying the mesh from the first memory to a second memory and retrieving the mesh from the second memory during compositing.
[0070] In some embodiments, the compositing subsystem 350 performs alpha compositing (e.g., alpha blending) based on the mesh and computer-generated content 320 to create an appearance of partial or full transparency. For example, referring to 2G, the compositing subsystem 350 performs alpha blending on a portion of physical agent 240 (e.g., captured by a camera) and a corresponding portion of computer-generated 320 such that the corresponding portion of computer-generated 320 appears partially or fully transparent depending on the portion of physical agent 240.
[0071] Figure 4 is an example of a flowchart of a method 400 for displaying an object indicator indicating an occluded portion of a physical object according to some embodiments. In various embodiments, method 400 or portions thereof are performed by an electronic device (e.g., Figure 1 electronic device 100 in Figures 2A to 2M electronic device 210 in
[0072] As represented by block 402, in some embodiments, method 400 includes obtaining environmental data from one or more environmental sensors. For example, referring to Figure 3, the environmental sensor includes a combination of an image sensor 312, a depth sensor 314, and an ambient light sensor 316. Other examples of environmental sensors include an Inertial Measurement Unit (IMU), a Simultaneous Localization and Mapping (SLAM) sensor, and a Visual Inertial Odometry (VIO) sensor. The environmental data can be a function of the current visible region associated with the respective sensor. For example, referring to Figure 2A , the image sensor integrated in the electronic device 210 has a pose that generally corresponds to the visible region 214 associated with the display 212. Thus, the image data from the image sensor includes corresponding representations of the first wall 202, the second wall 204, and the physical cabinet 220.
[0073] As shown in block 404, method 400 includes determining a plurality of position values associated with a physical object based on a function of environmental data. The plurality of position values are respectively associated with a plurality of positions of the physical object. The physical object can correspond to any type of physical (e.g., real-world) object, such as a physical agent (e.g., a person, an animal, or a robot) or an inanimate physical article (e.g., a stapler placed on a table). As an example, referring to Figure 2E , the object tracker 242 determines a first position value (indicated by the second marking 246) associated with the right hand of the physical agent 240, a second position value (not shown) associated with the left hand of the physical agent 240, a third position value (not shown) associated with the face of the physical agent 240, etc. The plurality of position values can include a combination of a set of pixels (within the image data from the image sensor), one or more depth values (within the depth data from the depth sensor), and one or more ambient light values (within the ambient light data from the ambient light sensor). In various embodiments, method 400 includes determining the plurality of position values by applying computer vision techniques to the environmental data, such as instance segmentation or semantic segmentation.
[0074] As represented in block 406, in some embodiments, method 400 includes displaying computer-generated content. For example, referring to Figure 2B , the electronic device 210 displays a computer-generated viewing screen 230 and a computer-generated dog 232 on the display 212.
[0075] In some embodiments, the computer-generated content is not associated with a physical object, such as when the display of the computer-generated content is independent of the physical object. In other words, the presence of a physical object within the visible region provided by the display does not affect the computer-generated content.
[0076] As shown in block 408, method 400 includes identifying a portion of computer-generated content that satisfies an occlusion criterion with respect to a corresponding portion of a physical object based on at least a portion of a plurality of position values. For example, a portion of the computer-generated content occludes a corresponding portion of the physical object, and the corresponding portion of the physical object would be visible but for the portion of the computer-generated content being displayed.
[0077] As shown in block 410, in some embodiments, the occlusion criterion is based on the position of the computer-generated content relative to the position of the physical object. To that end, in some embodiments, method 400 includes identifying a portion of computer-generated content 320 that at least partially overlaps a corresponding portion of the physical object. As an example, as Figure 2E shown, content recognizer 250 identifies a portion of computer-generated view 230 that overlaps a corresponding portion (e.g., above the legs) of physical agent 240 based on a plurality of position values 248. In the previous example, the plurality of position values 248 may include a set of pixels from an image of the corresponding portion of physical agent 240 from an image sensor.
[0078] As shown in block 412, in some embodiments, the occlusion criterion is based on the depth of the computer-generated content relative to the depth of the physical object. As an example, as Figure 2E shown, the plurality of position values 248 includes one or more depth values associated with a portion of physical agent 240. Accordingly, content recognizer 250 identifies a portion of computer-generated view 230 that satisfies occlusion criterion 340 because the portion of computer-generated view 230 is associated with a first depth value that is less than one or more depth values associated with the portion of physical agent 240.
[0079] As shown in block 414, in some embodiments, the occlusion criterion is partially based on the opacity associated with the computer-generated content. For example, referring to Figure 2E , content recognizer 250 identifies a portion of computer-generated view 230 that satisfies occlusion criterion 340 because that portion of computer-generated view 230 is associated with an opacity characteristic that exceeds a threshold. Accordingly, the corresponding portion of physical agent 240 cannot be viewed on the display through the portion of computer-generated view 230.
[0080] In some embodiments, the occlusion criterion is based on a combination of relative position, relative depth, and opacity associated with the computer-generated content.
[0081] As shown in block 416, in some embodiments, method 400 includes determining that a semantic value associated with a physical object satisfies an object criterion. To that end, method 400 includes obtaining the semantic value, such as by relative to an image from an image sensor (e.g.,Figure 3 performs semantic segmentation on the image data of the image sensor 312) therein. In some specific implementations, when the semantic value indicates a physical agent such as a person, an animal, or a robot, the semantic value meets the object criterion. In some specific implementations, when the semantic value indicates a predefined object type such as a trend or a common object type, the semantic value meets the object criterion. For example, the predefined object type is specified via user input (e.g., a text string). In some specific implementations, in response to determining that the semantic value meets the object criterion, method 400 proceeds to a portion of method 400 represented by block 422, which will be discussed below.
[0082] As shown in block 418, in some specific implementations, when the physical object corresponds to a physical agent, method 400 includes determining whether the physical agent meets a pose criterion or a movement criterion. At the end, method 400 includes determining pose characteristics associated with the physical agent or detecting movement characteristics associated with the physical agent based on the environmental data. For example, method 400 includes determining that the physical agent meets the pose criterion when the pose characteristics indicate that the physical agent is facing the device - e.g., the semantic segmentation output value is "face" or "eye". As another example, method 400 includes determining that the physical agent meets the movement criterion when the movement characteristics indicate that the physical agent's arm is waving. In some specific implementations, in response to determining that the physical agent meets the pose criterion or the movement criterion, method 400 proceeds to a portion of method 400 represented by block 422, which will be discussed below.
[0083] As shown in block 420, in some specific implementations, method 400 includes determining that the gaze position meets a proximity threshold relative to a portion of the computer-generated content. For example, referring to Figure 2E , the eye tracker 216 outputs eye tracking data 218 indicating the gaze position of the user 50. Additionally, the electronic device 210 determines that the gaze position meets the proximity threshold because the distance between the gaze position represented by the distance line 251 and the identified portion of the computer-generated viewing screen 230 is less than the threshold. Accordingly, the electronic device 210 displays an object indicator (e.g., Figure 2F the contour 262 in Figure 2H and Figure 2M the grid 264 in
[0084] As shown in block 422, in some embodiments, method 400 includes generating a grid associated with a physical object based on multiple position values. Thus, method 400 may include using the outputs of blocks 416, 418, and / or 420 to determine whether to generate a grid, thereby achieving optimization and resource savings. In some embodiments, a grid is generated in response to identifying a portion of computer-generated content that meets an occlusion criterion. Thus, an electronic device or system performing method 400 may selectively generate a grid, thereby avoiding resource utilization when generating a grid is not appropriate. Refer to Figure 3 , the grid generator 330 may generate a grid based on multiple position values 248. As shown in block 424, in some embodiments, the grid corresponds to a volumetric (e.g., 3D) grid. To this end, in some embodiments, method 400 includes obtaining multiple depth values within depth data from a depth sensor and generating a volumetric grid based at least in part on the multiple depth values. In some embodiments, method 400 includes applying a contour polygon function (e.g., marching squares function) with respect to image data representing a physical object in order to generate a grid.
[0085] As shown in block 426, in some embodiments, the grid enables penetration with respect to computer-generated content. For example, refer to Figure 2G , when the grid 263 is composited at the corresponding portion of the computer-generated viewing frame 230, the physical agent 240 is viewable on the display 212. For example, in some embodiments, method 400 includes using alpha blending in order to produce an apparent transparency with respect to computer-generated content.
[0086] As shown in block 428, in some embodiments, the grid indicates a shadow associated with a physical agent. For example, refer to Figure 2H , the grid 264 indicates the shadow of the physical agent 240 without indicating the features (e.g., eyes, nose) of the physical agent 240. In some embodiments, the darkness of the shadow is based on the depth of the computer-generated content relative to the depth of the physical object. For example, refer to Figure 2H , when the physical agent 240 moves closer to the computer-generated viewing frame 230 (e.g., the difference between the corresponding depth values decreases), the shadow associated with the grid 264 darkens, and vice versa as the physical agent 240 moves away from the computer-generated viewing frame 230. In some embodiments, when the difference between the corresponding depth values is greater than a threshold amount, method 400 includes stopping the display of the shadow grid. For example, when the physical agent is at a relatively large distance behind the computer-generated content (from the perspective of the user viewing the display), then the physical agent is less likely to attempt to draw the user's attention. Thus, displaying a shadow in this case is not likely to be helpful to the user.
[0087] As shown in block 430, in some embodiments, method 400 includes storing a mesh in a non-transitory memory. For example, an electronic device that executes method 400 includes a non-transitory memory and stores the mesh in the non-transitory memory. As another example, referring to Figure 3 , system 300 stores the mesh from mesh generator 330 in buffer 332.
[0088] As shown in block 432, method 400 includes displaying an object indicator indicative of a corresponding part of a physical object. The object indicator overlaps a part of the computer-generated content. In some embodiments, method 400 includes displaying the object indicator in response to identifying a part of the computer-generated content that meets an occlusion criterion, represented by block 408. The object indicator indicates the position of the physical object within the physical environment.
[0089] In some embodiments, displaying the object indicator corresponds to displaying a contour overlaid on a part of the computer-generated content, such as Figure 2F the contour 262 shown in. To this end, method 400 includes determining a contour associated with a corresponding part of the physical object based on a plurality of position values. In some embodiments, the contour satisfies a color contrast threshold and / or a brightness contrast threshold relative to the part of the computer-generated content such that the contour can be easily viewed overlaid on the part of the computer-generated content.
[0090] In some embodiments, displaying the object indicator corresponds to increasing a transparency characteristic associated with a part of the computer-generated content so that the corresponding part of the physical agent is viewable on the display. For example, method 400 includes applying a mask to a part of the computer-generated content. A compositing subsystem (e.g., compositing subsystem 350) can apply the mask.
[0091] As shown in block 434, in some embodiments, method 400 includes displaying the mesh as the object indicator, such as Figure 2G the mesh 263 shown in or Figure 2H the mesh 264 shown in. To this end, in some embodiments, as shown in block 436, method 400 includes compositing (e.g., via compositing subsystem 350) the mesh with a part of the environmental data associated with the physical object. Compositing includes retrieving the mesh from the non-transitory memory. Thus, method 400 utilizes a common non-transitory memory for mesh storage during mesh generation and for mesh retrieval during compositing, thereby reducing latency and computational cost.
[0092] The present disclosure describes various features, none of which alone can achieve the benefits described herein. It should be understood that the various features described herein may be combined, modified, or omitted, which will be apparent to those of ordinary skill in the art. Other combinations and sub-combinations other than those specifically described herein will be apparent to ordinary skill in the art and are intended to form part of the present disclosure. Various methods are described herein in connection with various flowchart steps and / or stages. It should be understood that in many cases, certain steps and / or stages may be combined together such that multiple steps and / or stages shown in the flowchart may be performed as a single step and / or stage. Additionally, certain steps and / or stages may be broken into additional sub-components to be performed independently. In some cases, the order of steps and / or stages may be rearranged and certain steps and / or stages may be entirely omitted. Additionally, the methods described herein should be understood to be broadly interpreted such that additional steps and / or stages other than those shown and described herein may also be performed.
[0093] Some or all of the methods and tasks described herein may be performed and fully automated by a computer system. In some cases, the computer system may include a plurality of different computers or computing devices (e.g., physical servers, workstations, storage arrays, etc.) that communicate and interoperate via a network to perform the functions. Each such computing device generally includes a processor (or processors) that executes program instructions or modules stored in a memory or other non-transitory computer-readable storage medium or device. The various functions disclosed herein may be implemented in such program instructions, but alternatively some or all of the disclosed functions may be implemented in dedicated circuitry (e.g., ASIC or FPGA or GP-GPU) of the computer system. In cases where the computer system includes multiple computing devices, these devices may or may not be located in the same location. The results of the disclosed methods and tasks may be persistently stored by transforming physical storage devices such as solid-state memory chips and / or disks into different states.
[0094] The various processes defined herein contemplate options for obtaining and utilizing a user's personal information. For example, such personal information may be utilized to provide an improved privacy screen on an electronic device. However, to the extent such personal information is collected, such information should be obtained with the user's informed consent. As described herein, the user should be aware of and in control of the use of their personal information.
[0095] Personal information will be used by appropriate parties only for legal and legitimate purposes. Parties utilizing such information will comply with privacy policies and practices that are at least in compliance with appropriate laws and regulations. Additionally, such policies should be comprehensive, user-accessible, and considered to be in compliance with or above government / industry standards. Additionally, except for any reasonable and legal purpose, parties shall not distribute, sell, or otherwise share such information.
[0096] However, a user may limit the extent to which parties can access or otherwise obtain personal information. For example, settings or other preferences may be adjusted so that the user can decide whether their personal information can be accessed by various entities. Additionally, although some of the features defined herein are described in the context of using personal information, aspects of these features may be implemented without the use of such information. For example, if user preferences, account names, and / or location history are collected, the information may be anonymized or otherwise generalized so that the information does not identify the corresponding user.
[0097] This disclosure is not intended to be limited to the specific implementations shown herein. Various modifications to the specific implementations described in this disclosure may be apparent to those skilled in the art, and the general principles defined herein may be applied to other specific implementations without departing from the spirit or scope of this disclosure. The teachings of the invention provided herein can be applied to other methods and systems and are not limited to the methods and systems described above, and elements and actions of the various specific implementations described above can be combined to provide more specific implementations. Thus, the novel methods and systems described herein can be implemented in many other forms; furthermore, various omissions, substitutions, and changes can be made to the forms of the methods and systems described herein without departing from the spirit of this disclosure. The appended claims and their equivalents are intended to cover such forms or modifications that fall within the scope and spirit of this disclosure.
Claims
1. A method, comprising: at an electronic device including one or more processors, non-transitory memory, one or more environmental sensors, and a display: displaying computer-generated content on the display; determining, based on environmental data from the one or more environmental sensors, a plurality of position values associated with a physical object; identifying, based on the plurality of position values indicating that the computer-generated content is located between the physical object and the user, a portion of the computer-generated content that satisfies a part of an occlusion criterion relative to the physical object; and in response to identifying that the occlusion criterion is satisfied: generating, based on the plurality of position values, a grid associated with the physical object; and displaying the grid on the display, wherein the grid overlaps with the portion of the computer-generated content.
2. The method according to claim 1, wherein displaying the grid corresponds to increasing a transparency characteristic associated with the portion of the computer-generated content.
3. The method according to claim 1, wherein the grid indicates a shadow associated with the physical object.
4. The method according to claim 1, the method further comprising storing the grid in the non-transitory memory, wherein displaying the grid includes compositing the grid with a portion of the environmental data associated with the physical object, and wherein the compositing includes retrieving the grid from the non-transitory memory.
5. The method according to claim 1, wherein the plurality of position values includes a first position value indicating a depth value.
6. The method according to claim 5, wherein identifying that the occlusion criterion is satisfied includes: determining, based on at least a portion of the plurality of position values including the first position value indicating the depth value, that the portion of the computer-generated content at least partially overlaps with the corresponding portion of the physical object on the display; and determining, based on at least a portion of the plurality of position values including the first position value indicating the depth value, that the portion of the computer-generated content is associated with a corresponding depth value that is less than a corresponding depth value associated with the physical object.
7. The method according to claim 5, wherein the one or more environmental sensors include a depth sensor that outputs depth data, wherein the plurality of position values includes a plurality of depth values based on the depth data, the depth values characterizing a physical agent, and wherein the grid corresponds to a volume grid based on the plurality of depth values.
8. The method according to claim 1, wherein identifying that the occlusion criterion is satisfied includes determining that the portion of the computer-generated content is associated with an opacity characteristic that exceeds a threshold.
9. The method according to claim 1, further comprising: obtaining, based on the environmental data, a semantic value associated with the physical object; and determining that the semantic value satisfies an object criterion; wherein generating the grid further responds to determining that the semantic value satisfies the object criterion.
10. The method according to claim 9, wherein When the semantic value indicates a physical agent, the semantic value satisfies the object criterion.
11. The method according to claim 9, wherein, when the semantic value indicates a predefined object type, the semantic value satisfies the object criterion.
12. The method according to claim 1, wherein, the physical object corresponds to a physical agent, the method further includes determining pose characteristics associated with the physical agent based on the environmental data, wherein generating the grid further responds to determining that the pose characteristics satisfy a pose criterion.
13. The method according to claim 1, wherein, the physical object corresponds to a physical agent, the method further includes detecting movement characteristics associated with the physical agent based on the environmental data, wherein generating the grid further responds to determining that the movement characteristics satisfy a movement criterion.
14. The method according to claim 1, wherein, the electronic device includes an eye tracker that outputs eye tracking data, wherein the eye tracking data indicates the gaze position of the user, and wherein generating the grid further responds to determining that the gaze position satisfies a proximity threshold relative to the portion of the computer-generated content.
15. The method according to claim 1, further includes determining a contour associated with the corresponding portion of the physical object based on the plurality of position values, wherein generating the grid is based on the contour.
16. The method according to claim 1, wherein, the one or more environmental sensors include an image sensor that outputs image data, and wherein the image data represents the physical object.
17. The method according to claim 1, wherein, the one or more environmental sensors include an ambient light sensor that senses ambient light from the physical environment and outputs corresponding ambient light data, and wherein the corresponding ambient light data is associated with the physical object.
18. The method according to claim 1, wherein, the one or more environmental sensors include a depth sensor that outputs depth sensor data, and wherein the depth sensor data indicates one or more depth values associated with the physical object.
19. A system, comprising: one or more environmental sensors; an object tracker for determining a plurality of position values associated with a physical object based on environmental data from the one or more environmental sensors; a content recognizer for recognizing a portion of computer-generated content that satisfies an occlusion criterion relative to a corresponding portion of the physical object based on the plurality of position values indicating that the computer-generated content is located between the physical object and the user; and a grid generator for generating a grid associated with the physical object based on the plurality of position values, wherein the grid generator generates the grid based on the content recognizer recognizing the portion of the computer-generated content that satisfies the occlusion criterion; and a display for displaying the grid overlapping the portion of the computer-generated content.
20. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs including instructions that, when executed by an electronic device having one or more processors, one or more environmental sensors, and a display, cause the electronic device to: display computer-generated content on the display; determine a plurality of position values associated with a physical object based on environmental data from the one or more environmental sensors; identify, based on the plurality of position values indicating that the computer-generated content is located between the physical object and the user, a portion of the computer-generated content that satisfies an occlusion criterion relative to the physical object; and in response to identifying that the occlusion criterion is satisfied: generate a grid associated with the physical object based on the plurality of position values; and display the grid on the display, wherein the grid overlaps the portion of the computer-generated content.
Citation Information
Patent Citations
Method and apparatus for selectively integrating sensory content
US10068369B2
Virtual reality display apparatus and display method thereof
US20170061696A1
Automatic placement of a virtual object in a three-dimensional space
US20180045963A1
Occlusion using pre-generated 3D models for augmented reality
US20190325661A1