Media processing system and method

JP2025503536A5Pending Publication Date: 2025-12-15QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024539367
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-13
Filing Date
2023-01-05
Publication Date
2025-12-15

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle unexpectedly captured scene parts in live streaming media, resulting in undesirable objects in user environment views, such as unauthorized facial images, and failure to effectively manage power and data compression efficiency.

Method used

By processing the image data, identifying and segmenting it into multiple areas, modifying the image data based on the object location and attributes, making the first area blurred without affecting the second area, and using object detection and semantic segmentation techniques, the hidden processing of objects that should not appear, including facial blur, resolution reduction or compression and other means.

Benefits of technology

Improve user privacy protection, reduce power consumption and data transmission requirements, while maintaining image quality, especially in extending the ability to process unauthorized objects in real-time in real-time in real-life devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Media processing systems and techniques are described. The media processing system receives image data representing an environment captured by an image sensor. The media processing system receives an indication of an object in the environment represented in the image data. The media processing system divides the image data into regions including a first region and a second region. The object is represented in one of the plurality of regions. The media processing system modifies the image data to obscure the first region without obscuring the second region based on the object represented in the one of the plurality of regions. The media processing system outputs the image data after modifying the image data. In some examples, the object is depicted in the first region and not in the second region. In some examples, the object is depicted in the second region and not in the first region.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001]

[0001] This application relates to media processing. More particularly, this application relates to systems and methods for obscuring and / or attenuating aspects of media corresponding to specific regions of an environment, while leaving other aspects of the media unobscured and / or unattenuated, e.g., based on object detection and / or semantic segmentation. [Background technology]

[0002]

[0002] Streaming media refers to media (e.g., video and / or audio) that is captured by a capture device and continuously provided from the capture device to one or more viewer devices over a network (e.g., the Internet) with little or no intermediate storage within the network elements. Streaming media may be provided from the capture device to one or more viewer devices while the capture device is still capturing media that is provided as a later portion of the streaming media, which is sometimes referred to as live streaming. Live streaming has little or no delay between capture and streaming, so there is little recourse if an unintended portion of a scene is captured.

[0003]

[0003] An extended reality (XR) device is a device that displays an environment to a user, for example through a head-mounted display (HMD) or a mobile handset. The environment is at least partially different from the real-world environment in which the user is located. A user can generally interactively change the view of their environment, for example by tilting or moving an HMD or other device. Virtual reality (VR), augmented reality (AR), and mixed reality (MR) are examples of XR. An XR device can include sensors that capture information from the environment. Because an XR device often provides a user with a primary view of the user's environment during use, the sensors of an XR device may sometimes capture unintended parts of a scene. Summary of the Invention

[0004] In some examples, systems and techniques for media processing are described. A media processing system receives image data captured by an image sensor. The image data represents (e.g., depicts) an environment. The media processing system receives an indication of an object in the environment to be depicted in the image data, for example by detecting the object using object detection. The media processing system divides the image data into regions based on the indication of the object, for example based on a location of the object in the image data. The regions include a first region and a second region. The object is represented in one of the plurality of regions. The media processing system modifies the image data to obscure the first region without obscuring the second region based on the object being represented in the one of the plurality of regions. The media processing system outputs the image data after modifying the image data. In some examples, the object is represented in the first region and not in the second region, and the media processing system obscures the first region because the object is therein. In some examples, an object is represented in the second region and not in the first region, and the media processing system obscures the first region because the object is not in it. In some examples, the object is a person. The media processing system may obscure the region to improve privacy, e.g., to obscure faces of people who were not intended to appear in the media. The media processing system may obscure the region in a manner that improves bandwidth usage and / or power consumption, e.g., by obscuring using increased compression in the modified region, reduced resolution in the modified region, etc.

[0005]

[0005] In one example, an apparatus for media processing is provided. The apparatus includes a memory and one or more processors (e.g., implemented in a circuit) coupled to the memory. The one or more processors are configured to receive image data representing an environment captured by an image sensor, receive an indication of an object in the environment represented in the image data, divide the image data into a plurality of regions including a first region and a second region, where the object is represented in one of the plurality of regions, modify the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions, and output the image data after modifying the image data.

[0006] In another example, a method of image processing is provided that includes receiving image data representative of an environment captured by an image sensor, receiving an indication of an object in the environment represented in the image data, dividing the image data into a plurality of regions, the plurality of regions including a first region and a second region, where the object is represented in one of the plurality of regions, modifying the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions, and outputting the image data after modifying the image data.

[0007]

[0007] In another example, a non-transitory computer readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to receive image data representing an environment captured by an image sensor, receive an indication of an object in the environment represented in the image data, divide the image data into a plurality of regions including a first region and a second region, where an object is represented in one of the plurality of regions, modify the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions, and output the image data after modifying the image data.

[0008]

[0008] In another example, an apparatus for image processing is provided that includes: means for receiving image data captured by an image sensor, the image data representing an environment, means for receiving an indication of an object in the environment represented in the image data, means for dividing the image data into a plurality of regions, the plurality of regions including a first region and a second region, an object being represented in one of the plurality of regions, means for modifying the image data based on the object being represented in one of the plurality of regions to obscure the first region without obscuring the second region, and means for outputting the image data after modifying the image data.

[0009] In some aspects, dividing the image data into a plurality of regions includes dividing the image data into a plurality of regions based on a determined location of the object, the object being located in at least one region and not located in at least one other region. In some aspects, the location of the object is determined from the image data. In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include detecting audio, the location of the object is determined based on attributes of the audio, the attributes including at least one of a location of the audio, a direction of the audio, an amplitude of the audio, or a frequency of the audio.

[0010] In some aspects, receiving an indication of the object in the environment includes detecting the object in the image data. In some aspects, receiving an indication of the object in the environment includes input via a user interface, the input being indicative of the object.

[0011] In some aspects, the object is represented in a first region and not represented in a second region, and modifying the image data to obscure the first region is based on the object being represented in the first region. In some aspects, the object is represented in a second region and not represented in the first region, and modifying the image data to obscure the first region without obscuring the second region is based on the object being represented in the second region and not represented in the first region.

[0012]

[0012] In some aspects, modifying the image data to obscure the first region includes modifying the image data using foveated compression of a peripheral area around the fixation point, where the second region includes the fixation point and the first region includes the peripheral area.

[0013] In some aspects, modifying the image data to obscure the first region includes modifying the image data to blur at least a portion of the first region. In some aspects, modifying the image data to obscure the first region includes modifying the image data to remove at least a portion of the first region. In some aspects, modifying the image data to obscure the first region includes modifying the image data to restore at least a portion of the first region. In some aspects, modifying the image data to obscure the first region includes modifying the image data to pixelate at least a portion of the first region.

[0014] In some aspects, modifying the image data to obscure the first region includes modifying the image data to reduce resolution of a first subset of image data representing the first region compared to a second subset of image data representing the second region. In some aspects, modifying the image data to obscure the first region includes modifying the image data to compress the first subset of image data representing the first region more than the second subset of image data representing the second region.

[0015] In some aspects, the object comprises at least a portion of a body of a person. In some aspects, the object comprises at least a portion of a face of a person. In some aspects, the object comprises at least a portion of a string of characters. In some aspects, the object comprises at least a portion of content to be displayed using the display.

[0016] In some aspects, outputting the image data includes displaying the image data using a display. In some aspects, outputting the image data includes transmitting the image data to a recipient device using a communications transceiver.

[0017]

[0017] In some aspects, one or more of the methods, devices, and computer-readable media described above further include receiving audio data captured by a microphone from an environment at a time corresponding to capture of the image data, detecting audio samples within the audio data corresponding to the object, modifying the audio data to attenuate the audio samples corresponding to the object, and outputting the audio data after modifying the audio data.

[0018] In some embodiments, at least one region is a region having a predetermined shape.

[0019]

[0019] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include receiving secondary image data from a second image sensor, the second image sensor having a different field of view than the first image sensor, the secondary image data captured by the second image sensor including a secondary image of the user, and segmenting the image data further based on the secondary image. In some aspects, the second image sensor captures a gesture or position of at least a portion of the user, and segmenting the image data includes defining a region corresponding to a gesture direction and / or position of at least a portion of the user. In some aspects, the gesture or position of the user includes a gaze direction of the user.

[0020] In some aspects, modifying the image data to obscure the first region reduces an amount of data used to code the first region. In some aspects, modifying the image data to obscure the first region includes at least one of increasing compression in the first region, increasing quantization in the first region, reducing resolution in the first region, cropping the first region, and / or pixelating the first region.

[0021] In some aspects, one or more of the methods, apparatus, and computer-readable media described above further include identifying an object, determining whether the object is to be displayed or obscured based on identifying the object, and in response to determining that the object is to be obscured, defining a first region to include the object. In some aspects, determining that the object is to be obscured includes determining that the object is included in a blacklist of objects to be obscured and / or determining that the object is not included in a whitelist of objects to be displayed.

[0022] In some aspects, the device is part of and / or includes a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a head mounted display (HMD) device, a wireless communication device, a mobile device (e.g., a mobile phone and / or mobile handset and / or a so-called "smartphone" or other mobile device), a camera, a personal computer, a laptop computer, a server computer, a vehicle or a computing device or component of a vehicle, another device, or a combination thereof. In some aspects, the device includes a camera or cameras that capture one or more images. In some aspects, the device further includes a display that displays one or more images, notifications, and / or other displayable data. In some aspects, the devices described above may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors.

[0023]

[0023] This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used independently to determine the scope of the claimed subject matter, which should be understood by reference to the entire specification of this patent, any or all drawings, and appropriate portions of each claim.

[0024]

[0024] The above, together with other features and embodiments, will become more apparent with reference to the following specification, claims, and accompanying drawings. [Brief description of the drawings]

[0025]

[0025] Exemplary embodiments of the present application are described in detail below with reference to the following drawings. [Figure 1]

[0026] 1 is a block diagram illustrating an example architecture of an image capture and processing system, in accordance with some examples. [Diagram 2]

[0027] 1 is a block diagram illustrating an example architecture of a media processing system that receives media captured by a sensor and executes a process to modify the media, according to some examples. [Figure 3A]

[0028] FIG. 1 is a perspective view illustrating a head mounted display (HMD) used as an extended reality (XR) system, according to some examples. [Figure 3B]

[0029] FIG. 3B is a perspective view illustrating the head mounted display (HMD) of FIG. 3A being worn by a user, according to some examples. [Figure 4A]

[0030] FIG. 1 is a perspective view illustrating the front of a mobile handset that includes a forward-facing camera and can be used as an extended reality (XR) system, according to some examples. [Figure 4B]

[0031] FIG. 1 is a perspective view illustrating the rear of a mobile handset that includes a rear-facing camera and can be used as an extended reality (XR) system, according to some examples. [Diagram 5]

[0032] FIG. 5 is a block diagram illustrating a process 500 for event-based image processing, according to some examples. [Figure 6]

[0033] FIG. 6 is a block diagram illustrating a process 600 for image processing based on detection of people in image data, according to some examples. [Figure 7A]

[0034] 1 is a conceptual diagram illustrating an example of an image of an environment and various modifications to the image to obscure portions of the environment, which are shown using dashed lines, according to some examples. [Figure 7B]

[0035] 1 is a conceptual diagram illustrating an example of an image of an environment and various modifications to the image to obscure portions of the depicted environment using shading, according to some examples. [Figure 8]

[0036] FIG. 1 is a conceptual diagram illustrating an example environmental soundscape and various modifications for attenuating aspects of the soundscape corresponding to different elements within the environment, according to some examples. [Figure 9]

[0037] FIG. 2 is a block diagram illustrating an example of a neural network that may be used for media processing operations, according to some examples. [Figure 10]

[0038] 1 is a flow diagram illustrating a process for media processing, according to some examples. [Figure 11]

[0039] FIG. 1 illustrates an example of a computing system for implementing certain aspects described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0026]

[0040] Specific aspects and embodiments of the present disclosure are provided below. As will be apparent to those skilled in the art, some of these aspects and embodiments may be applied independently, and some of them may be applied in combination. In the following description, for the purposes of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it will be apparent that various embodiments may be practiced without these specific details. The figures and descriptions are not intended to be limiting.

[0027]

[0041] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the following description of exemplary embodiments provides those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes may be made in the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0028]

[0042] A camera is a device that uses an image sensor to receive light and capture image frames, such as still images or video frames. The terms "image," "image frame," and "frame" are used interchangeably herein. A camera may be configured with various image capture and image processing settings. Different settings result in images with different appearances. Some camera settings, such as ISO, exposure time, aperture size, f / stop, shutter speed, focus, and gain, are determined and applied before or during the capture of one or more image frames. For example, settings or parameters may be applied to an image sensor to capture one or more image frames. Other camera settings, such as contrast, brightness, saturation, sharpness, levels, curves, or color changes, may constitute post-processing of one or more image frames. For example, settings or parameters may be applied to a processor (e.g., an image signal processor or ISP) to process one or more image frames captured by the image sensor.

[0029]

[0043] An extended reality (XR) system or device can provide virtual content to a user and / or combine a real-world view of a physical environment (scene) with a virtual environment (including the virtual content). The XR system facilitates user interaction with such a combined XR environment. The real-world view can include real-world objects (also called physical objects), such as people, vehicles, buildings, tables, chairs, and / or other real-world or physical objects. The XR system or device can facilitate interaction with different types of XR environments (e.g., a user can use the XR system or device to interact with an XR environment). The XR system can include a virtual reality (VR) system that facilitates interaction with an augmented reality (AR) environment, an MR system that facilitates interaction with a mixed reality (MR) environment, and / or other XR systems. Examples of XR systems or devices include head-mounted displays (HMDs), smart glasses, among others. In some cases, the XR device can track parts of the user (e.g., the user's hands and / or fingertips) to allow the user to interact with items of virtual content.

[0030]

[0044] Systems and techniques for media processing are described herein. A media processing system receives image data captured by an image sensor. The image data represents (e.g., depicts) an environment. The media processing system receives an indication of an object in the environment to be represented (e.g., depicted) in the image data, e.g., by detecting the object. The media processing system divides the image data into regions, e.g., based on the indication of the object, the detection of the object, the location of the object in the environment, and / or the location of the object in the image data. The regions include a first region and a second region. The object is represented (e.g., depicted) in one of the plurality of regions. The media processing system modifies the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions. The media processing system outputs the image data after modifying the image data. In some examples, the object is represented (e.g., depicted) in the first region and not in the second region, and the media processing system obscures the first region because the object is therein. In some examples, the object is represented (e.g., depicted) in the second region and not represented in the first region, and the media processing system obscures the first region because the object is not in it. In some examples, the object is a person, a face, a vehicle, a plant, an animal, a structure, a device, content displayed on a device, content written or drawn on the media, or a combination thereof.

[0031]

[0045] The media processing systems and techniques described herein provide many technical improvements over conventional media systems. For example, the media processing systems and techniques described herein can obscure regions in a manner that improves bandwidth usage and / or power consumption, for example, by obscuring using increased compression in the modified regions, reduced resolution in the modified regions, etc. In some examples, the media processing systems and techniques described herein can improve privacy and security, for example, by obscuring faces of people or objects that were not intended to appear in the media.

[0032]

[0046] Various aspects of the application are described with respect to the figures. FIG. 1 is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components used to capture and process images of one or more scenes (e.g., images of a scene 110). The image capture and processing system 100 can capture a standalone image (or photo) and / or capture a video including multiple images (or video frames) in a particular sequence. A lens 115 of the system 100 faces the scene 110 and accepts light from the scene 110. The lens 115 bends the light toward the image sensor 130. The light accepted by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is accepted by the image sensor 130. In some examples, the scene 110 is a scene in an environment, such as the environment facing the environment-facing sensor 210 of FIG. 2. In some examples, the scene 110 is a scene of at least a portion of a user, such as the user facing the user-facing sensor 205 of FIG. 2. For example, scene 110 may be a scene of one or both of a user's eyes and / or at least a portion of a user's face.

[0033]

[0047] The one or more controls 120 can control exposure, focus, and / or zoom based on information from the image sensor 130 and / or based on information from the image processor 150. The one or more controls 120 may include multiple mechanisms and components. For example, the control 120 may include one or more exposure controls 125A, one or more focus controls 125B, and / or one or more zoom controls 125C. The one or more controls 120 may also include additional controls beyond those shown, such as controls to control analog gain, flash, HDR, depth of field, and / or other image capture characteristics.

[0034]

[0048] The focus control mechanism 125B of the control mechanism 120 can obtain the focus setting. In some examples, the focus control mechanism 125B stores the focus setting in a memory register. Based on the focus setting, the focus control mechanism 125B can adjust the position of the lens 115 relative to the position of the image sensor 130. For example, based on the focus setting, the focus control mechanism 125B can move the lens 115 closer to or farther from the image sensor 130 by actuating a motor or servo, thereby adjusting the focus. In some cases, additional lenses, such as one or more microlenses above each photodiode of the image sensor 130, may be included in the system 100, each of which bends light received from the lens 115 toward a corresponding photodiode before the light reaches the photodiode. The focus setting may be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus settings may be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus settings may be referred to as image capture settings and / or image processing settings.

[0035]

[0049] The exposure control 125A of the control mechanism 120 can obtain an exposure setting. In some cases, the exposure control 125A stores the exposure setting in a memory register. Based on the exposure setting, the exposure control 125A can control the size of the aperture (e.g., aperture size or f / stop), the duration the aperture is open (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure setting may be referred to as an image capture setting and / or an image processing setting.

[0036]

[0050] The zoom control 125C of the control mechanism 120 can obtain the zoom setting. In some examples, the zoom control 125C stores the zoom setting in a memory register. Based on the zoom setting, the zoom control 125C can control the focal length of an assembly of lens elements (lens assembly) including the lens 115 and one or more additional lenses. For example, the zoom control 125C can control the focal length of the lens assembly by actuating one or more motors or servos to move one or more of the lenses relative to each other. The zoom setting may be referred to as an image capture setting and / or an image processing setting. In some examples, the lens assembly may include a parfocal zoom lens or a variable focus zoom lens. In some examples, the lens assembly may include a focusing lens (which may be the lens 115 in some cases) that first accepts light from the scene 110, and the light then passes through an afocal zoom system between the focusing lens (e.g., the lens 115) and the image sensor 130 before the light reaches the image sensor 130. In some cases, an afocal zoom system may include two positive (e.g., converging, convex) lenses of equal or similar focal lengths (e.g., within a threshold difference) with a negative (e.g., diverging, concave) lens between them. In some cases, the zoom control 125C moves one or more of the lenses in the afocal zoom system, such as one or both of the negative and positive lenses.

[0037]

[0051] Image sensor 130 includes one or more arrays of photodiodes or other light-sensitive elements. Each photodiode measures an amount of light that ultimately corresponds to a particular pixel in the image generated by image sensor 130. In some cases, different photodiodes may be covered by different color filters, and thus may measure light that matches the color of the filter covering the photodiode. For example, a Bayer color filter includes red, blue, and green filters, and each pixel of the image is generated based on red light data from at least one photodiode covered by a red filter, blue light data from at least one photodiode covered by a blue filter, and green light data from at least one photodiode covered by a green filter. Other types of color filters may use yellow, magenta, and / or cyan (also called "emerald") color filters instead of or in addition to red, blue, and / or green filters. Some image sensors may be completely devoid of color filters, and instead use different photodiodes (possibly stacked vertically) throughout the pixel array. Different photodiodes across the pixel array can have different spectral sensitivity curves and therefore respond to different wavelengths of light. Monochrome image sensors may also lack color filters and therefore no color depth.

[0038]

[0052] In some cases, image sensor 130 may alternatively or additionally include opaque and / or reflective masks that block light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which may be used for phase detection autofocus (PDAF). Image sensor 130 may also include analog gain amplifiers for amplifying analog signals output by the photodiodes and / or analog to digital converters (ADCs) for converting analog signals output from the photodiodes (and / or amplified by the analog gain amplifiers) to digital signals. In some cases, certain components or functions discussed with respect to one or more of control mechanisms 120 may instead or additionally be included within image sensor 130. The image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active-pixel sensor (APS), a complimentary metal-oxide semiconductor (CMOS), an N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0039]

[0053] Image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more of any other types of processors 1110 discussed with respect to computing system 1100. Host processor 152 may be a digital signal processor (DSP) and / or other types of processors. In some implementations, image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-chip or SoC) that includes host processor 152 and ISP 154. In some cases, the chip may include one or more input / output ports (e.g., input / output (I / O) ports 156), central processing units (CPUs), graphics processing units (GPUs), broadband modems (e.g., 3G, 4G or LTE, 5G, etc.), memory, connectivity components (e.g., Bluetooth™, Global Positioning System (GPS), etc.), any combination thereof, and / or other components.The I / O ports 156 may include any suitable input / output ports or interfaces according to one or more protocols or specifications, such as an Inter-Integrated Circuit 2 (I2C) interface, an Inter-Integrated Circuit 3 (I3C) interface, a Serial Peripheral Interface (SPI) interface, a serial General Purpose Input / Output (GPIO) interface, a Mobile Industry Processor Interface (MIPI) (e.g., a MIPI CSI-2 physical (PHY) layer port or interface, etc.), an Advanced High-performance Bus (AHB) bus, any combination thereof, and / or other input / output ports. In one illustrative example, the host processor 152 may communicate with the image sensor 130 using an I2C port and the ISP 154 may communicate with the image sensor 130 using a MIPI port.

[0040]

[0054] Image processor 150 may perform several tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form HDR images, image recognition, object recognition, feature recognition, accepting input, managing output, managing memory, or some combination thereof. Image processor 150 may store image frames and / or processed images in random access memory (RAM) 140 and / or 1120, read-only memory (ROM) 145 and / or 1125, a cache, a memory unit, another storage device, or some combination thereof.

[0041]

[0055] Various input / output (I / O) devices 160 may be connected to image processor 150. I / O devices 160 may include a display screen, a keyboard, a keypad, a touch screen, a track pad, a touch-sensitive surface, a printer, any other output device 1135, any other input device 1145, or some combination thereof. In some cases, captions may be entered into image processing device 105B through a physical keyboard or keypad of I / O device 160 or through a virtual keyboard or keypad of a touch screen of I / O device 160. I / O 160 may include one or more ports, jacks, or other connectors that enable a wired connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or send data to one or more peripheral devices. I / O 160 may include one or more wireless transceivers that enable a wireless connection between system 100 and one or more peripheral devices, through which system 100 may receive data from and / or transmit data to one or more peripheral devices. The peripheral devices may include any of the types of I / O devices 160 previously described, and may themselves be considered I / O devices 160 when coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector.

[0042]

[0056] In some cases, image capture and processing system 100 may be a single device. In some cases, image capture and processing system 100 may be two or more separate devices including image capture device 105A (e.g., a camera) and image processing device 105B (e.g., a computing device coupled to a camera). In some implementations, image capture device 105A and image processing device 105B may be coupled, for example, via one or more wires, cables, or other electrical connectors and / or wirelessly via one or more wireless transceivers. In some implementations, image capture device 105A and image processing device 105B may be separate from each other.

[0043]

[0057] As shown in Figure 1, a vertical dashed line divides the image capture and processing system 100 of Figure 1 into two portions, which respectively represent image capture device 105A and image processing device 105B. Image capture device 105A includes lens 115, control mechanism 120, and image sensor 130. Image processing device 105B includes image processor 150 (including ISP 154 and host processor 152), RAM 140, ROM 145, and I / O 160. In some cases, some components shown in image capture device 105A, such as ISP 154 and / or host processor 152, may be included within image capture device 105A.

[0044]

[0058] The image capture and processing system 100 may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smartphone, a mobile phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video gaming console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the image capture and processing system 100 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 wi-fi communication, wireless local area network (WLAN) communication, or any combination thereof. In some implementations, the image capture device 105A and the image processing device 105B may be different devices. For example, the image capture device 105A may include a camera device, and the image processing device 105B may include a computing device, such as a mobile handset, a desktop computer, or other computing device.

[0045]

[0059] Although image capture and processing system 100 is shown as including certain components, one skilled in the art will appreciate that image capture and processing system 100 may include many more components than those shown in FIG. 1. The components of image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of image capture and processing system 100 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing image capture and processing system 100.

[0046]

[0060] 2 is a block diagram illustrating an example architecture of a media processing system 200 that receives media captured by a sensor and performs a process to modify the media. In some examples, the media processing system 200 includes at least one image capture and processing system 100, an image capture device 105A, an image processing device 105B, or a combination or combinations thereof. In some examples, the media processing system 200 includes at least one computing system 1100. In some examples, the media processing system 200 includes at least one neural network 900.

[0047]

[0061] In some examples, the media processing system 200 includes one or more user-facing sensors 205. The user-facing sensors 205 capture sensor data that measures and / or tracks information about the user's body aspects and / or behavior by the user. In some examples, the user-facing sensors 205 include one or more cameras facing at least a portion of the user. The one or more cameras can include one or more image sensors that capture images of at least a portion of the user. For example, the user-facing sensors 205 can include one or more cameras focused on one or both eyes (and / or one or both eyelids) of the user, and the image sensors of the cameras capture images of the user's eye or eyes. The one or more cameras may also be referred to as eye capturing sensor(s). In some implementations, the one or more cameras can capture a series of images over time, which in some examples can be sequenced together in a temporal order, e.g., into a video. These series of images may depict or otherwise indicate, for example, the user's eye(s), pupil dilation, blinking (using the eyelids), squinting (using the eyelids), saccades, fixations, eye moisture levels, optokinetic reflexes or responses, vestibulo-ocular reflexes or responses, accommodation reflexes or responses, other attributes related to the eyes and / or eyelids described herein, or combinations thereof. In Figure 2, the one or more user-facing sensors 205 are shown as cameras facing the user's eyes and capturing images of the user's eyes.

[0048]

[0062] The user-facing sensors 205 may include one or more sensors that track information about the user's body and / or behavior, such as one or more cameras, image sensors, microphones, heart rate monitors, oximeters, biometric sensors, positioning receivers, Global Navigation Satellite System (GNSS) receivers, inertial measurement units (IMUs), accelerometers, gyroscopes, gyrometers, barometers, thermometers, altimeters, depth sensors, light detection and ranging (LIDAR) sensors, radio detection and ranging (RADAR) sensors, sound detection and ranging (SODAR) sensors, sound navigation and ranging (SONAR) sensors, time of flight (ToF) sensors, structured light sensors, other sensors discussed herein, or combinations thereof. In some examples, the one or more user-facing sensors 205 include at least one image capture and processing system 100, image capture device 105A, image processing device 105B, or a combination or combinations thereof. In some examples, the one or more user-facing sensors 205 include at least one input device 1145 of the computing system 1100. In some implementations, one or more of the user-facing sensor(s) 205 may supplement or refine sensor readings from other user-facing sensor(s) 205 and / or environmental facing sensor(s) 210. For example, inertial measurement units (IMUs), accelerometers, gyroscopes, or other sensors may be used by the gaze tracking engine 270 to refine the determination of the user's gaze.

[0049]

[0063] The one or more environmental facing sensors 210 of the media processing system 200 are one or more sensors pointed, oriented, and / or focused on the environment. In some examples, the one or more environmental facing sensors 210 face away from the user. The user facing sensor(s) 205 face a first direction and the environmental facing sensor(s) 210 face a second direction. In some examples, the second direction is parallel to the first direction. In some examples, the first direction and the second direction are opposite, opposite, and / or opposite to each other. In some examples, the one or more environmental facing sensors 210 can be pointed, oriented, and / or facing in a direction that the user's face faces. In some examples, the one or more environmental facing sensors 210 can be pointed, oriented, and / or facing in a direction that the media processing system 200 (or a side thereof) faces.

[0050]

[0064] The environmental facing sensors 210 capture sensor data that measures and / or tracks information about the real-world environment in front of and / or around the media processing system 200 and / or the user. In some examples, the environmental facing sensors 210 include one or more cameras facing at least a portion of the real-world environment. The one or more cameras can include one or more image sensors that capture images of at least a portion of the real-world environment. For example, the environmental facing sensors 210 can include one or more cameras focused on the real-world environment (e.g., the surroundings of the media processing system 200), where the image sensors of the cameras capture images of the real-world environment (e.g., the surroundings). Such cameras can capture a series of images over time, which in some examples can be sequenced together in a temporal order, e.g., into a video. These series of images can depict or otherwise show, for example, floors, ground, walls, ceilings, sky, water, plants, other people other than the user, parts of the user's body (e.g., arms or legs), structures, vehicles, animals, devices, other objects, or combinations thereof. 2, the one or more environmental facing sensors 210 are shown as cameras facing the house (e.g., structure) and person. In some examples, the one or more environmental facing sensors 210 include at least one image capture and processing system 100, image capture device 105A, image processing device 105B, or a combination or combinations thereof. In some examples, the one or more environmental facing sensors 210 include at least one input device 1145 of the computing system 1100. The environmental facing sensors 210 may include a camera, an image sensor, a positioning receiver, a GNSS receiver, an IMU, an accelerometer, a gyroscope, a gyrometer, a barometer, a thermometer, an altimeter, a depth sensor, a LIDAR sensor, a RADAR sensor, a SODAR sensor, a SONAR sensor, a ToF sensor, a structured light sensor, other sensors described herein, or a combination thereof.

[0051]

[0065] In some implementations, one or more of the environmental facing sensor(s) 210 may complement or refine sensor readings from other user facing sensor(s) 205 and / or environmental facing sensor(s) 210. For example, sensor data from cameras, image sensors, depth sensors, LIDAR sensors, RADAR sensors, SODAR sensors, SONAR sensors, ToF sensors, and / or structured light sensors may be combined or otherwise used together by object detection engine 225 to detect objects in the environment and / or by semantic segmentation engine 230 to segment a representation of the environment.

[0052]

[0066] In some examples, user input can be further used to detect objects in the environment. In an illustrative example, a touch screen user interface can receive a user touch input at a location on the touch screen where a preview image of the environment is displayed, and the location of the touch input can be used by the object detection engine 225 to identify a corresponding location in the preview image and / or other image data and / or other sensor data as having an object or being likely to have an object (e.g., having a reduced confidence threshold for object recognition). In another illustrative example, a mouse user interface can receive a click input at a location on the screen where a preview image of the environment is displayed, and the location of the click input can be used by the object detection engine 225 to identify a corresponding location in the preview image and / or other image data and / or other sensor data as having an object or being likely to have an object (e.g., having a reduced confidence threshold for object recognition).

[0053]

[0067] In some examples, one or more of the environment facing sensor(s) 210 may include one or more microphones that may record audio from the environment. In some examples, one or more of the environment facing sensor(s) 210 may include multiple microphones such that the direction and / or location of the audio in the environment may be determined from differences in the audio recorded at different microphones. In some examples, the audio may be further used to detect objects in the environment. For example, if a microphone(s) of the environment facing sensor(s) 210 detects a voice, the object detection engine 225 may increase the likelihood of detecting a person in the environment and / or in the image data. In some examples, attributes of the audio, such as the direction the audio is coming from, the direction the audio signal is traveling, the location of the audio (e.g., determined via triangulation), the amplitude of the audio, and / or the frequency of the audio, may suggest a position of the object to the object detection engine 225, which may increase the likelihood of detecting a person in that part of the environment and / or in the image data depicting that part of the environment (e.g., with a reduced confidence threshold for object recognition). For example, the direction of the audio and / or the location of the audio may identify the direction an object is relative to the microphone(s) (and / or environmental facing sensor(s) 210 or other parts of media processing system 200). The location of the audio, the amplitude of the audio, and / or the frequency of the audio may indicate how far away an object is relative to the microphone(s) (and / or environmental facing sensor(s) 210 or other parts of media processing system 200). The frequency of the audio may indicate whether an object is moving relative to the microphone(s) (and / or environmental facing sensor(s) 210 or other parts of media processing system 200), for example, based on the Doppler effect.Such indications may be used by the object detection engine 225 based on the audio to identify corresponding locations in the image data and / or other sensor data as having an object or likely to have an object (e.g., having a reduced confidence threshold for object recognition) based on, for example, direction (e.g., where to look for the object in the image), distance (e.g., whether to look for the object in the foreground or background), velocity (e.g., whether the object may include motion blur), or a combination thereof.

[0054]

[0068] In some examples, the media processing system 200 includes a virtual content generator 215 that generates virtual content. The virtual content may include a two-dimensional (2D) shape, a three-dimensional (3D) shape, a 2D object, a 3D object, a 2D model, a 3D model, a 2D animation, a 3D animation, a 2D image, a 3D image, a texture, a portion of another image, a character, a character string, or a combination thereof. In FIG. 2, the virtual content generated by the virtual content generator 215 is shown as a tetrahedron. In some examples, the virtual content generator 215 includes one or more software elements, such as one or more instruction sets corresponding to one or more programs executed on one or more processors of the media processing system 200, such as the processor 1110, the image processor 150, the host processor 152, the ISP 154, or a combination thereof, of the computing system 1100. In some examples, the virtual content generator 215 includes one or more hardware elements. For example, the virtual content generator 215 may include a processor, such as the processor 1110 of the computing system 1100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the virtual content generator 215 includes a combination of one or more software elements and one or more hardware elements.

[0055]

[0069] In some examples, media processing system 200 includes one or more output devices 240 configured and capable of outputting media. In some examples, output device(s) 240 include display(s) configured and capable of displaying visual media, such as images and / or videos. In some examples, output device(s) 240 include audio output device(s), such as loudspeakers or headphones, or connectors configured to couple media processing system 200 to loudspeakers or headphones. Audio output device(s) are configured and capable of playing audio media, such as music, sound effects, audio tracks corresponding to videos, recordings recorded by microphone(s) (e.g., of user-facing sensor(s) 205, environment-facing sensor(s) 210, and / or additional sensor(s) of media processing system 200), or combinations thereof. Output device(s) 240 may output media including a representation of the environment (e.g., as captured by environment facing sensor(s) 210), virtual content (e.g., as generated by virtual content generator 215), a combination of the representation of the environment and the virtual content (e.g., as generated by compositor 220), modification(s) to the representation(s) of the environment and / or the virtual content and / or combination (e.g., as generated by media modification engine 235), or a combination thereof. In some examples, output device(s) 240 may face a user of media processing system 200. For example, display(s) of output device(s) 240 may face a user of media processing system 200 and / or display visual media to (e.g., towards) a user of media processing system 200.Similarly, audio output device(s) of output device(s) 240 may face and / or play audio media to (e.g., towards) a user of media processing system 200. In some examples, output device(s) 240 includes output device 1135. In some examples, output device 1135 may include output device(s) 240.

[0056]

[0070] Media processing system 200 includes a compositor 220. The compositor 220 composes, composites, and / or combines virtual content (e.g., generated by virtual content generator 215) with the representation(s) of the environment. In some examples, the representation(s) of the environment are captured by environmental facing sensor(s) 210. In some examples, the representation(s) of the environment are visible to a user based on light from the environment reaching the user through a portion of media processing system 200 (e.g., through at least a portion of the display(s) of output device(s) 240).

[0057]

[0071] In some examples, the display(s) of output device(s) 240 of media processing system 200 function as optical “see-through” display(s) that allow light from the real-world environment (scene) surrounding media processing system 200 to pass across (e.g., through) the display(s) of output device(s) 240 to reach one or both eyes of a user. For example, the display(s) of output device(s) 240 may be at least partially transparent, semi-transparent, light permissive, light transmissive, or a combination thereof. In an illustrative example, the display(s) of output device(s) 240 include a transparent, semi-transparent, and / or light transmissive lens and a projector. The display(s) of output device(s) 240 may include a projector that projects virtual content onto the lens. The lenses may be, for example, eyeglass lenses, goggle lenses, contact lenses, lenses of a head mounted display (HMD) device, or a combination thereof. Light from the real world environment passes through the lenses and reaches one or both of the user's eyes. A projector projects virtual content onto the lenses so that the virtual content appears to be overlaid on top of the user's view of the environment from the perspective of one or both of the user's eyes. The compositor 220 may determine and / or modify display settings that control the positioning of the virtual content projected onto the lenses by the display(s) of the output device(s) 240.

[0058]

[0072] In some examples, the display(s) of the output device(s) 240 of the media processing system 200 include a lensless projector(s) as described above with respect to an optical see-through display. In such examples, the output device(s) 240 can use the projector(s) to project virtual content to one or both eyes of a user. In some examples, the projector of the display(s) of the output device(s) 240 can project virtual content to one or both retinas of one or both eyes of a user. In such examples, the display(s) of the output device(s) 240 may be referred to as an optical see-through display, a virtual retinal display (VRD), a retinal scan display (RSD), or a retinal projector (RP) display. In such examples, light from the real-world environment (scene) still reaches one or both eyes of the user. As the projector projects virtual content into the user's eye or eyes, the virtual content appears to be overlaid on the user's view of the environment from the perspective of one or both of the user's eyes. The compositor 220 can determine and / or modify display settings that control the positioning of the virtual content projected by the display(s) of the output device(s) 240 onto the user's eye(s).

[0059]

[0073] In some examples, the display(s) of output device(s) 240 of media processing system 200 are digital “pass-through” displays that allow a user of media processing system 200 to see a view of the environment by displaying a view of the environment on the display(s) of output device(s) 240. The view of the environment displayed on the digital pass-through display may be a view of the real-world environment surrounding media processing system 200 based on, for example, sensor data (e.g., images, videos, depth images, point clouds, other depth data, or combinations thereof) captured by one or more environmental facing sensors 210 of media processing system 200. In some examples, the view of the environment displayed on the digital pass-through display may include virtual content (e.g., generated by virtual content generator 215) and / or modifications (e.g., by media modification engine 235) that are incorporated into the view of the environment.

[0060]

[0074] The view of the environment displayed on the pass-through display may be a view of a virtual or mixed environment that is separate from but based on the real-world environment. For example, the virtual or mixed environment may include virtual objects and / or backgrounds, but may be mapped to an area and / or volume of space having dimensions based on dimensions of the area and / or volume of space in the real-world environment in which the user and media processing system 200 reside. Media processing system 200 may determine dimensions of the area and / or volume of space in the real-world environment in which the user and media processing system 200 reside. In some implementations, the environmental facing sensor(s) 210 of media processing system 200 may include a camera and / or image sensor that captures images of the environment (e.g., surroundings of media processing system 200) and / or a depth sensor (e.g., LIDAR, RADAR, SONAR, SODAR, ToF, structured light) that captures depth data (e.g., point cloud, depth image) of the environment. This can ensure that the user does not accidentally go down a set of stairs, bump into a wall or obstacle, or otherwise have a negative and / or potentially dangerous interaction with the real-world environment while exploring the virtual or mixed environment displayed on the display(s) of the output device(s) 240.

[0061]

[0075] In examples where the display(s) of the output device(s) 240 are digital pass-through displays, the media processing system 200 can use the compositor 220 to overlay the virtual content generated by the virtual content generator 215 over at least a portion of the environment captured using the environmental facing sensor(s) 210. In some examples, the compositor 220 can overlay the virtual content completely over the environment displayed on the display(s) of the output device(s) 240 such that, from the perspective of one or both eyes of a user looking at the display(s) of the output device(s) 240, the virtual content appears to be completely in front of the remainder of the environment displayed on the display(s) of the output device(s) 240. In some examples, the compositor 220 can overlay at least a portion of the virtual content over portions of the environment displayed on the display(s) of the output device(s) 240 such that, from the perspective of one or both eyes of a user looking at the display(s) of the output device(s) 240, the virtual content appears to be in front of some portions of the environment displayed on the display(s) of the output device(s) 240 but behind other portions of the environment displayed on the display(s) of the output device(s) 240. Thus, the compositor 220 can provide simulated depth to the virtual content and overlay portions of the environment displayed on the display(s) of the output device(s) 240 over portions of the virtual content.

[0062]

[0076] In examples where the display(s) of the output device(s) 240 are optical see-through displays, the media processing system 200 can use the compositor 220 to ensure that portions of the real-world environment are not overlaid by the virtual content generated by the virtual content generator 215. In some examples, the compositor 220 can only partially overlay the virtual content onto the real-world environment on the display such that the virtual content appears to be behind at least a portion of the real-world environment from the perspective of one or both eyes of a user viewing the display(s) of the output device(s) 240. In some examples, the compositor 220 can only partially overlay the virtual content onto the real-world environment on the display such that the virtual content appears to be behind at least a portion of the real-world environment and in front of other portions of the real-world environment from the perspective of one or both eyes of a user viewing the display(s) of the output device(s) 240. Thus, the compositor 220 can provide simulated depth to the virtual content and prevent portions of the real-world environment from being overlaid by the virtual content. The positioning of the virtual content relative to the environment can be identified and / or indicated by a display setting (e.g., a first display setting, a second display setting). The compositor 220 can determine and / or modify the display setting.

[0063]

[0077] Whether the display(s) of output device(s) 240 are optical see-through displays or digital pass-through displays, the display(s) of output device(s) 240 may, in some cases, provide a 3D view of the environment, virtual content, and / or modifications to the user. For example, media processing system 200 may output two slightly different perspectives to each of the user's two eyes on the display(s) of output device(s) 240 so that the display(s) of output device(s) 240 provide a 3D view to the user, in some cases providing a stereoscopic view of the environment incorporating the virtual content and / or modifications.

[0064]

[0078] The compositor 220 of the media processing system 200 may determine a display setting (e.g., a first display setting) for the display(s) of the output device(s) 240. In a media processing system 200 in which the display(s) of the output device(s) 240 are digital “pass-through” displays, the compositor 220 may generate an image that composes, composites, and / or combines a view of the environment (e.g., based on sensor data from the environmental facing sensors 210) with the virtual content generated by the virtual content generator 215. The display settings generated by the compositor 220 may indicate the position, orientation, depth, size, color, font size, font color, text language, layout, and / or other properties of the virtual content and / or of particular elements or portions of the virtual content. In media processing system 200 where the display(s) of output device(s) 240 are optical "see-through" displays, compositor 220 can generate display settings that indicate the position, orientation, depth, size, color, font size, font color, text language, and / or other characteristics of the virtual content and / or particular elements or portions of the virtual content as displayed by the display(s) of output device(s) 240 (e.g., as projected by the projector(s) of the display(s) of output device(s) 240 onto the lens(es) and / or eye(s)). In Figure 2, compositor 220 is shown as adding the virtual content (represented by the tetrahedrons) to a view of the environment (represented by the house and the people). In FIG. 2, output device(s) 240 are shown as a display that shows and / or provides a view of both the virtual content (represented by the tetrahedrons) and a view of the environment (represented by the house and people), as well as speakers that output audio corresponding to one or both of these.

[0065]

[0079] In some examples, the synthesizer 220 includes an ML system(s) and / or trained ML model(s) that receive as inputs the sensor data from the environmental facing sensor(s) 210, the virtual content generated by the virtual content generator 215, and / or the gaze data from the gaze tracking engine 270. The ML system(s) and / or trained ML model(s) output a combined media that includes at least a portion of the sensor data from the environmental facing sensor(s) 210 and at least a portion of the virtual content. In some cases, the ML system(s) and / or trained ML model(s) can position the virtual content based on the gaze data. In some examples, the ML system(s) and / or trained ML model(s) of the synthesizer 220 may include one or more neural networks (NNs) (e.g., neural network 900), one or more convolutional neural networks (CNNs), one or more trained time delay neural networks (TDNNs), one or more deep networks, one or more autoencoders, one or more deep belief nets (DBNs), one or more recurrent neural networks (RNNs), one or more generative adversarial networks (GANs), one or more other types of neural networks, one or more trained support vector machines (SVMs), one or more trained random forests (RANs), one or more trained neural networks (NN ... The systems may include one or more neural networks (e.g., neural forests, RFs), one or more computer vision systems, one or more deep learning systems, or a combination thereof.

[0066]

[0080] In some examples, the compositor 220 includes a software element, such as a set of instructions corresponding to a program executing on a processor, such as the processor 1110, the image processor 150, the host processor 152, the ISP 154 of the computing system 1100, or a combination thereof. In some examples, the compositor 220 includes one or more hardware elements. For example, the compositor 220 may include a processor, such as the processor 1110, the image processor 150, the host processor 152, the ISP 154 of the computing system 1100, or a combination thereof. In some examples, the compositor 220 includes a combination of one or more software elements and one or more hardware elements.

[0067]

[0081] The media processing system 200 includes an object detection engine 225. In some examples, the object detection engine 225 receives visual media data (e.g., images, videos) from the environment facing sensor(s) 210, the virtual content generator 215, and / or the compositor 220. The object detection engine 225 detects, recognizes, classifies, and / or tracks one or more feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s). The object detection engine 225 can include one or more machine learning (ML) systems having one or more trained ML models. The object detection engine 225 may perform feature detection, feature extraction, feature recognition, feature tracking, object detection, object recognition, object tracking, face detection, face recognition, face tracking, person detection, person recognition, person tracking, animal detection, animal recognition, animal tracking, device detection, device recognition, device tracking, vehicle detection, vehicle recognition, vehicle tracking, classification, or combinations thereof. The object detection engine 225 may perform these operations by inputting the visual media data into trained ML model(s) and receiving detections of feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) as output of the trained ML model(s). This detection may identify locations and / or areas in which feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) are located in the visual media data.

[0068]

[0082] The ML system(s) and / or trained ML model(s) of the object detection engine 225 may include one or more NNs, one or more CNNs, one or more TDNNs, one or more deep networks, one or more autoencoders, one or more DBNs, one or more RNNs, one or more GANs, one or more trained SVMs, one or more trained RFs, one or more computer vision systems, one or more deep learning systems, or a combination thereof. In some examples, the object detection engine 225 generates a confidence level associated with each detection of feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) within the visual media data and reports the detection (e.g., to the semantic segmentation engine 230 and / or the media modification engine 235) if the confidence level meets or exceeds a predetermined confidence level threshold.

[0069]

[0083] In some examples, the object detection engine 225 receives audio media data (e.g., sound clips, music clips, audio samples, and / or recordings) from the microphone(s) of the environment facing sensor(s) 210, the virtual content generator 215, and / or the synthesizer 220. The object detection engine 225 detects, recognizes, classifies, and / or tracks one or more sound clips, music clips, audio samples, and / or audio recordings in the audio media data. In some examples, the object detection engine 225 detects, recognizes, classifies, and / or tracks audio corresponding to a particular object (e.g., a person, a vehicle, a device, or other object). The object detection engine 225 can include one or more machine learning (ML) systems having one or more trained ML models. The object detection engine 225 may perform audio feature detection, audio feature extraction, audio feature recognition, audio feature tracking, voice detection, voice recognition, voice tracking, device sound detection, device sound recognition, device sound tracking, animal sound detection, animal sound recognition, animal sound tracking, animal detection, vehicle sound recognition, vehicle sound tracking, vehicle sound detection, vehicle sound recognition, vehicle sound tracking, object sound detection, object sound recognition, object sound tracking, classification, or combinations thereof. The object detection engine 225 may perform these operations by inputting the audio media data into trained ML model(s) and receiving as output of the trained ML model(s) detection of sounds corresponding to audio feature(s), object(s), voice(s), animal(s), device(s), and / or vehicle(s).This detection can identify audio characteristics (e.g., frequency and / or amplitude and / or sound direction) and / or time(s) of sounds in the audio media data where sounds occur that correspond to audio feature(s), object(s), voice(s), animal(s), device(s), and / or vehicle(s). In Figure 2, the object detection engine 225 is shown as a bounding box around people in the media, but not around a house or tetrahedron.

[0070]

[0084] The ML system(s) and / or trained ML model(s) may include one or more NNs, one or more CNNs, one or more TDNNs, one or more deep networks, one or more autoencoders, one or more DBNs, one or more RNNs, one or more GANs, one or more trained SVMs, one or more trained RFs, one or more deep learning systems, or combinations thereof. In some examples, the object detection engine 225 generates a confidence level associated with each detection of a sound corresponding to an audio feature(s), object(s), voice(s), animal(s), device(s), and / or vehicle(s) in the audio media data, and reports the detection (e.g., to the semantic segmentation engine 230 and / or the media modification engine 235) if the confidence level meets or exceeds a predetermined confidence level threshold.

[0071]

[0085] In some examples, the object detection engine 225 receives gaze data from the gaze tracking engine 270 and uses the gaze data as input to the ML system and / or trained ML model(s) of the object detection engine 225. If the gaze data indicates that the user is looking at a particular region of the environment, the object detection engine 225 may reduce its confidence threshold for that region of the environment, such that the object detection engine 225 indicates detection of an object(s) in the region if the confidence meets or exceeds a predefined reduced confidence threshold, even if it does not meet or exceed a standard confidence threshold.

[0072]

[0086] In some examples, the object detection engine 225 detects, recognizes, and / or tracks a part or parts of the user's body, such as one or more of the user's hands and / or feet. In some examples, the user's hands or feet may be one of the object or objects detected in the media by the object detection engine 225. In some examples, the object detection engine 225 detects, recognizes, and / or tracks one or more objects held and / or touched by the user's hand(s). In some examples, the object detection engine 225 detects, recognizes, and / or tracks one or more objects standing on and / or touched by the user's foot or both feet. In some examples, the object detection engine 225 detects, recognizes, and / or tracks one or more objects pointed at by the user and / or gestures towards them using the user's hand or feet. In some examples, the object detection engine 225 may reduce its confidence threshold for an area of ​​the environment where one or more hands or feet of the user are holding, touching, pointing, gesturing towards, or a combination thereof. Thus, the object detection engine 225 may indicate detection of an object(s) in an area if the confidence meets or exceeds a predetermined reduced confidence threshold, even if the confidence does not meet or exceed a standard confidence threshold.

[0073]

[0087] In some examples, the object detection engine 225 includes a software element, such as a set of instructions corresponding to a program executing on a processor, such as the processor 1110, the image processor 150, the host processor 152, the ISP 154 of the computing system 1100, or a combination thereof. In some examples, the object detection engine 225 includes one or more hardware elements. For example, the object detection engine 225 may include a processor, such as the processor 1110, the image processor 150, the host processor 152, the ISP 154 of the computing system 1100, or a combination thereof. In some examples, the object detection engine 225 includes a combination of one or more software elements and one or more hardware elements.

[0074]

[0088] The media processing system 200 includes a semantic segmentation engine 230. The semantic segmentation engine 230 divides the media (e.g., the media captured by the environmental facing sensor(s) 210, the virtual content generated by the virtual content generator 215, and / or the combined media generated by the compositor 220) into segments. In some examples, the semantic segmentation engine 230 may divide the media into segments or regions based on the location(s) of one or more feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) detected in the visual media data by the object detection engine 225. For example, the semantic segmentation engine 230 may divide one or more images into a first region and a second region. The first region includes one or more feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) detected in the visual media data by the object detection engine 225. The second region lacks (does not include and / or is missing) one or more feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) detected in the visual media data by the object detection engine 225.

[0075]

[0089] In some examples, the semantic segmentation engine 230 may divide the media into segments or regions based on the location(s) and / or direction(s) of sound(s) corresponding to audio feature(s), object(s), sound(s), animal(s), device(s), and / or vehicle(s) detected in the audio media data by the object detection engine 225. For example, the semantic segmentation engine 230 may divide one or more images into a first region and a second region. The first region may be based on the location(s) and / or direction(s) of sound(s) corresponding to audio feature(s), object(s), sound(s), animal(s), device(s), and / or vehicle(s) detected in the audio media data by the object detection engine 225. The first region includes the location(s) and / or direction(s) of the sound(s) corresponding to the audio feature(s), object(s), voice(s), animal(s), device(s), and / or vehicle(s) detected in the audio media data by the object detection engine 225. The second region is devoid of (does not include and / or is missing) the location(s) and / or direction(s) of the sound(s) corresponding to the audio feature(s), object(s), voice(s), animal(s), device(s), and / or vehicle(s) detected in the audio media data by the object detection engine 225. In FIG. 2, the semantic segmentation engine 230 is shown as including separate boxes defining separate regions around the person, the house, and the tetrahedron (virtual content).

[0076]

[0090] In some examples, the region is two-dimensional (2D), e.g., the media includes a two-dimensional image, video, or other media. In some examples, the region has or includes a polygonal shape, e.g., a square, a rectangle, a quadrilateral, a triangle, a pentagon, a hexagon, or another polygonal shape. In some examples, the region has or includes a rounded shape, e.g., a circle, an oval, or another rounded shape. In some examples, the region is three-dimensional (3D), e.g., the media includes a three-dimensional depth image, a point cloud, video depth data, or other media. In some examples, the region has or includes a polyhedral shape, e.g., a cube, a rectangular prism, a quadrilateral prism, a triangular prism, a pentagonal prism, a hexagonal prism, a tetrahedron, a pyramid, or another polyhedral shape. In some examples, the region has or includes a rounded 3D shape, including, for example, a sphere, an ellipsoid, a cylinder, a cone, or another rounded shape. In some examples, the boundary of the region in the media is or is based on the boundary in the media of one or more feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) detected in the media by the object detection engine 225. In some examples, the boundary of the region in the media is based on a fractional or decimal semantic segmentation of the media, such as the left and right halves, or the top and bottom halves, or a diagonal half, or a quadrant, or a horizontal or vertical third, or another similar semantic segmentation. In some examples, the boundary of the region in the media is based on a central region of the media and / or a peripheral region around the central region. In some examples, the boundaries of the region in the media are based on gaze data from the gaze tracking engine 270, based on the gaze region the user is looking at according to the gaze tracking engine 270, and a peripheral region around the gaze region. The peripheral region may be within the user's peripheral vision in some examples. The peripheral region may be outside the user's vision in some examples.

[0077]

[0091] In some examples, the semantic segmentation engine 230 may include an ML system and / or a trained ML model. In some examples, the semantic segmentation engine 230 may perform these semantic segmentation operations by inputting the media data, the object detection data from the object detection engine 225, and / or the gaze data from the gaze tracking engine 270 into a trained ML model(s) and receiving as an output of the trained ML model(s) an indication of distinct regions resulting from the semantic segmentation, or the location and / or boundaries of the regions. The ML system(s) and / or the trained ML model(s) may include one or more NNs, one or more CNNs, one or more TDNNs, one or more deep networks, one or more autoencoders, one or more DBNs, one or more RNNs, one or more GANs, one or more trained SVMs, one or more trained RFs, one or more deep learning systems, or a combination thereof. In some examples, the semantic segmentation engine 230 generates a confidence level associated with the semantic segmentation into regions. In some examples, if the confidence level meets or exceeds a predetermined confidence level threshold, the semantic segmentation engine 230 outputs the region or an indication thereof.

[0078]

[0092] In some examples, the semantic segmentation engine 230 receives gaze data from the gaze tracking engine 270 and uses the gaze data as input to the ML system and / or trained ML model(s) of the semantic segmentation engine 230. If the gaze data indicates that the user is looking at a particular region of the environment, the semantic segmentation engine 230 can segment the media such that the region is one of, contains, or is included by the regions in which the semantic segmentation engine 230 segments the media. In some examples, if the gaze data indicates that the user is looking at a particular region of the environment, the semantic segmentation engine 230 can reduce the confidence threshold of the semantic segmentation engine 230 for semantic segmentation based on that region of the environment.

[0079]

[0093] In some examples, the semantic segmentation engine 230 can receive data from the object detection engine 225 and can divide the media into segments based on detection of a region or regions of the environment where one or more hands or feet of the user are holding, touching, pointing, gesturing, or a combination thereof. For example, the semantic segmentation engine 230 can divide the media into a first region and a second region. The first region includes the region or regions of the environment where one or more hands or feet of the user are holding, touching, pointing, gesturing, or a combination thereof. The second region lacks (does not include and / or is missing) the region or regions of the environment where one or more hands or feet of the user are holding, touching, pointing, gesturing, or a combination thereof.

[0080]

[0094] In some examples, the semantic segmentation engine 230 includes a software element, such as a set of instructions corresponding to a program executing on a processor, such as the processor 1110 of the computing system 1100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the semantic segmentation engine 230 includes one or more hardware elements. For example, the semantic segmentation engine 230 may include a processor, such as the processor 1110 of the computing system 1100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the semantic segmentation engine 230 includes a combination of one or more software elements and one or more hardware elements.

[0081]

[0095] The media processing system 200 includes a media modification engine 235. The media modification engine 235 modifies media before the media is output using the output device(s) 240 and / or transceiver(s) 245 of the media processing system 200. The media modified by the media modification engine 235 may include media captured by the environment facing sensor(s) 210, virtual content generated by the virtual content generator 215, and / or combined media generated by the compositor 220. The media modification engine 235 may modify a portion(s) of media to obscure and / or attenuate the portion(s) of media, possibly without obscuring and / or attenuating other portion(s) of the media. The portion(s) of media may be referred to as a region, subset, area, and / or aspect of the media.

[0082]

[0096] The media modification engine 235 may, for example, modify one or more first region(s) of visual media data of the media (e.g., image(s), video(s)) without modifying one or more second region(s) of the media based on information from the semantic segmentation engine 230, the object detection engine 225, or both. For example, the media modification engine 235 may modify the first region(s) of the visual media data to obscure the first region(s) of the visual media data without obscuring the second region(s) of the visual media data. The first region(s) and second region(s) may be identified using the object detection engine 225 and / or the semantic segmentation engine 230. In one illustrative example, the first region(s) include one or more feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) detected in the media by the object detection engine 225, while the second region(s) lack (do not include) such detection(s). In another illustrative example, the second region(s) include one or more feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) detected in the media by the object detection engine 225, while the first region(s) lack (do not include) such detection(s). Thus, in some examples, the media modification engine 235 obscures areas that contain detected feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s), while leaving other areas unobscured.In some examples, the media modification engine 235 leaves areas containing the detected feature(s), object(s), face(s), person(s), animal(s), device(s), and / or vehicle(s) unobscured and obscures other areas.

[0083]

[0097] The media modification engine 235 can obscure the region(s) of the media visual media data in a variety of ways. For example, the media modification engine 235 can blur the region(s), scramble the region(s), pixelate the region(s), pixelate the region(s), pixelate the region(s), mosaic the region(s), crop the region(s), compress the region(s) more than other region(s) of the visual media data using image and / or video compression techniques, reduce the resolution of the region(s) compared to the resolution of other region(s) of the visual media data, etc. The region(s) of the visual media data may be obscured by reducing the image resolution of the region(s) of the visual media data, quantizing the region(s) more strongly than other region(s) of the visual media data during image and / or video compression, removing the region(s), replacing the region(s) with other data (e.g., a color, a pattern, another image), inpainting the region(s) (e.g., using interpolation based on one or more surrounding pixels and / or regions), or combinations thereof. In some examples, the media modification engine 235 may obscure the region(s) of the visual media data with a clear, sharp boundary between the obscured and unobscured regions. In some examples, the media modification engine 235 may obscure the region(s) of the visual media data with a gradual gradient boundary between the obscured and unobscured regions, for example, as shown in FIG. 7B. In some examples, the media modification engine 235 may obscure an area or areas of the visual media data using foveated compression, foveated blur, foveated resolution reduction, foveated pixelation, foveated shading, other foveated image processing, or a combination thereof.For example, in embodiments in which the media modification engine 235 may modify one or more first regions(s) of visual media data of a media (e.g., image(s), video(s)) without modifying one or more second regions(s) of the visual media data, the one or more second regions(s) may include a fixation point, e.g., a fixation(s) by an eye(s) of the user 205 as determined by the gaze tracking engine 270, described in more detail below, and the one or more first regions(s) may include a peripheral area around the fixation point, in which case the media modification engine 235 may modify the visual media data to obscure the one or more first regions(s) by modifying the visual media data using foveated compression of the peripheral area around the fixation point. In FIG. 2, the semantic segmentation engine 230 is shown as having pixelated and / or made pixellated the house and tetrahedrons (virtual content), but not the people.

[0084]

[0098] In some examples, using increased compression, increased quantization, resolution reduction, cropping, and / or pixelation of a region or regions of media to obscure a region or regions of visual media data can result in bandwidth savings, storage space savings, and / or power savings. For example, the modified media may require less data (e.g., a fewer number of bits) to store and / or transmit, and thus bandwidth may be saved within media processing system 200 and / or in transferring media from media processing system 200 to a recipient device. The modified media may also require less energy to encode and / or decode, and thus power may be saved on both the encoding side (e.g., to store the modified media) and the decoding side (e.g., to display and / or play the modified media).

[0085]

[0099] The media modification engine 235 may, for example, modify a first portion of the audio media data of the media based on information from the semantic segmentation engine 230, the object detection engine 225, or both to attenuate, mute, and / or remove a first portion of the audio media data without attenuating, mute, and / or removing a second portion of the audio media data. In some examples, in response to detection of particular sound(s) in the audio media data by the object detection engine 225 and / or semantic segmentation of those sound(s) from other audio in the audio media data by the semantic segmentation engine 230, the media modification engine 235 may attenuate, mute, and / or remove those sound(s) from the audio media data that correspond to audio feature(s), object(s), voice(s), animal(s), device(s), and / or vehicle(s). For example, the media modification engine 235 may attenuate, mute, and / or remove a particular person's voice from the audio media data in response to detection and / or recognition of the person's voice in the audio media data by the object detection engine 225 and / or semantic segmentation of the person's voice from other audio in the audio media data by the semantic segmentation engine 230.

[0086]

[0100] In some examples, the media modification engine 235 may include an ML system and / or a trained ML model. In some examples, the media modification engine 235 may perform these modification operations by inputting the media data, the object detection data from the object detection engine 225, the semantic segmentation data from the semantic segmentation engine 230, and / or the gaze data from the gaze tracking engine 270 into the trained ML model(s) and receiving as output of the trained ML model(s) the modified media with modification(s) to the portion(s) of the media and / or modification(s) in the portion(s) to be modified. The ML system(s) and / or trained ML model(s) may include one or more NNs, one or more CNNs, one or more TDNNs, one or more deep networks, one or more autoencoders, one or more DBNs, one or more RNNs, one or more GANs, one or more trained SVMs, one or more trained RFs, one or more deep learning systems, or combinations thereof. In some examples, the media modification engine 235 generates a confidence level associated with the modification(s). In some examples, the media modification engine 235 outputs and / or performs the modification(s) on the media if the confidence level meets or exceeds a predefined confidence level threshold.

[0087]

[0101] In some examples, the media modification engine 235 receives gaze data from the gaze tracking engine 270. In some examples, the media modification engine 235 can use the gaze data to determine whether to modify a particular region or leave the region unmodified. In some examples, the media modification engine 235 uses the gaze data as input to its ML system and / or trained ML model(s).

[0088]

[0102] In some examples, the media modification engine 235 includes a software element, such as a set of instructions corresponding to a program executing on a processor, such as the processor 1110, the image processor 150, the host processor 152, the ISP 154 of the computing system 1100, or a combination thereof. In some examples, the media modification engine 235 includes one or more hardware elements. For example, the media modification engine 235 may include a processor, such as the processor 1110, the image processor 150, the host processor 152, the ISP 154 of the computing system 1100, or a combination thereof. In some examples, the media modification engine 235 includes a combination of one or more software elements and one or more hardware elements.

[0089]

[0103] As mentioned above, media processing system 200 includes output device(s) 240 that media processing system 200 can use to output media after modifying the media using media modification engine 235, e.g., by displaying visual media data of the media using display(s) of output device(s) 240 and / or by playing audio media data of the media using audio output device(s) of output device(s) 240. Media processing system 200 also includes one or more transceivers 245 that media processing system 200 can use to output media after modifying the media using media modification engine 235, e.g., by transmitting the media to a recipient device. The recipient device can output media using its own output device(s), e.g., by displaying visual media data of the media using display(s) of output device(s) and / or playing audio media data of the media using audio output device(s) of output device(s). The transceiver(s) 245 may include a wired or wireless transceiver(s), a communication interface(s), an antenna(s), a connection, a coupling, a coupling system, or a combination thereof. In some examples, the transceiver(s) 245 may include a communication interface 1140 of the computing system 1100. In some examples, the communication interface 1140 of the computing system 1100 may include the transceiver(s) 245. In FIG. 2, the transceiver(s) 245 is shown as a wireless transceiver(s) 245 transmitting media data, which is shown to include representations of a person, a house, and a tetrahedron (virtual content).

[0090]

[0104] In some examples, the media processing system 200 includes a gaze tracking engine 270. The gaze tracking engine 270 can receive sensor data from the user-facing sensor(s) 205 and detects, recognizes, and / or tracks a user's gaze (e.g., where the user is looking, what the user is looking at in the environment and / or medium), the user's facial expression(s), and / or the user's gestures based on the sensor data. In some examples, the sensor data that the gaze tracking engine 270 receives from the user-facing sensor(s) 205 includes image(s) and / or video(s) of the user's eye(s). In some examples, the sensor data that the gaze tracking engine 270 receives from the user-facing sensor(s) 205 includes depth data (e.g., point cloud, depth image) of the user's eye(s). The gaze tracking engine 270 can detect, recognize, and / or track a user's gaze, a user's facial expression(s), and / or a user's gestures based on one or more attributes of the user's eye(s) and / or face detected in sensor data from the user-facing sensor(s) 205. The attributes may include, for example, the position(s) of the eye(s) of the user 205, the movement(s) of the eye(s) of the user 205, the position(s) of the eyelid(s) of the user 205, the movement(s) of the eyelid(s) of the user 205, the position(s) of the eyebrow(s) of the user 205, the movement(s) of the eyebrow(s) of the user 205, the pupil dilation(s) of the eye(s) of the user 205, the fixation(s) by the eye(s) of the user 205. The information may include, for example, eye movement, eyelid movement, eyelid saccade ...In FIG. 2, the gaze tracking engine 270 is shown as identifying both the direction the user's eyes are looking (indicated by the solid black arrow) and the angle that that direction has changed over time (indicated by the dashed black curved arrow).

[0091]

[0105] In some examples, the gaze tracking engine 270 can include an ML system and / or a trained ML model. In some examples, the gaze tracking engine 270 can perform gaze tracking operations by inputting sensor data from a user-facing sensor(s) including, for example, an image of the eye(s) of a user(s) and / or media data (e.g., to determine what the user's gaze is directed at in the media data) into a trained ML model(s). The gaze tracking engine 270 can receive gaze data as an output of the trained ML model(s) indicating where the user is looking, what the user is looking at in the media, various eye movements and / or other eye attributes, or a combination thereof. The ML system(s) and / or trained ML model(s) may include one or more NNs, one or more CNNs, one or more TDNNs, one or more deep networks, one or more autoencoders, one or more DBNs, one or more RNNs, one or more GANs, one or more trained SVMs, one or more trained RFs, one or more deep learning systems, or combinations thereof. In some examples, the gaze tracking engine 270 generates a confidence level associated with the gaze tracking. In some examples, the gaze tracking engine 270 outputs the gaze data if the confidence level meets or exceeds a predetermined confidence level threshold.

[0092]

[0106] In some examples, the gaze tracking engine 270 includes a software element, such as a set of instructions corresponding to a program executed on a processor, such as the processor 1110 of the computing system 1100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the gaze tracking engine 270 includes one or more hardware elements. For example, the gaze tracking engine 270 can include a processor, such as the processor 1110 of the computing system 1100, the image processor 150, the host processor 152, the ISP 154, or a combination thereof. In some examples, the gaze tracking engine 270 includes a combination of one or more software elements and one or more hardware elements.

[0093]

[0107] In some examples, the media processing system 200 performs modifications to the media using the media modification engine 235 based on the object(s) detected by the object detection engine 225, the semantic segmentation of the media by the semantic segmentation engine 230, and / or the gaze data detected by the gaze tracking engine 270 to improve privacy. For example, the environmental facing sensor(s) 210 may sometimes capture parts of an environment that include people or objects that may not be the intended focus of the media. Such people may not have consented to appear in the media, in some cases. This may be of particular concern when the media processing system 200 is an XR device and / or a live streaming device, which may increase the probability that the environmental facing sensor(s) 210 may capture parts of an environment that include such people or objects. This is because the XR device may often be the primary lens through which a user views their environment, and thus a user may point their XR device at someone or something without realizing what the user is pointing the XR device at. Live streaming devices often have little or no delay between the capture and transmission of their respective media, which leaves little or no recourse if unintended or unwanted people or objects appear in the environment captured by the environmental facing sensor(s) 210. The use of modifications by the media modification engine 235 can limit the people or objects that appear explicitly in the media to only those on an approved list (e.g., a white list) and / or those that do not appear on a blocked list (e.g., a black list). The use of modifications by the media modification engine 235 may obscure and / or attenuate the visual and / or audio portions of the media that correspond to people or objects that appear on a blocked list (e.g., a black list) or that do not appear on an approved list (e.g., a white list).Thus, redaction of media by the media redaction engine 235 provides a powerful, near-instantaneous privacy enhancement that is useful in situations where a user does not have enough time to perform media edits.

[0094]

[0108] In some examples, media processing system 200 performs modification using media modification engine 235 based on identities of particular people or objects detected, for example, based on whether those identities appear on an approved list (e.g., a whitelist) or a blocked list (e.g., a blacklist). In some examples, the approved list and / or blocked list may be based on information from messages, emails, event invitations, schedules, calendars, contact lists, or combinations thereof. In some examples, the approved list may be automatically generated by media processing system 200 to include invitees to an event that appear on a calendar, schedule. In some examples, the blocked list may be automatically generated by media processing system 200 to include anyone not invited to an event. In some examples, the reverse is true, since event invitees are on the blocked list and / or non-invites are on the approved list. In some examples, the approved list may be automatically generated by media processing system 200 to include people to whom a message was sent, such as people who appear in the "to" field, "cc" field, and / or "bcc" field of an email. In some examples, the block list may be automatically generated by media processing system 200 to include anyone to whom a message was not sent. In some examples, the reverse is true, because a message recipient is on the block list and / or a non-recipient is on the approved list. In some examples, the approved list may be automatically generated by media processing system 200 to include a single person, such as the sender of a message or email or the host of an event, and / or someone else is placed on the block list. In some examples, the reverse is true, because a single person is on the block list and / or other people are on the approved list.

[0095]

[0109] In some examples, the media processing system 200 may use its object detection engine 225 in two passes. For example, the object detection engine 225 may perform a preliminary coarse pass to determine whether a type of object or a sound associated with a type of object is present in the media at all. For example, the object detection engine 225 may determine whether any faces are present in the visual media data and / or whether any sounds are present in the audio media data. If the object detection engine 225 detects the presence of a type of object or a sound associated with a type of object in the preliminary coarse pass, the object detection engine 225 may perform a more detailed pass. For example, if a first pass of the object detection engine 225 determines that one or more faces are present in the visual media data and / or determines that one or more faces are present in the audio media data, the object detection engine 225 may perform a more detailed pass to determine whether the object detection engine 225 recognizes any of the detected faces and / or whether the object detection engine 225 recognizes any of the detected sounds.

[0096]

[0110] In some examples, the media processing system 200 includes a feedback engine 260. The feedback engine 260 can detect feedback received from a user interface. The feedback engine 260 can detect feedback regarding one engine of the media processing system 200 received from another engine of the media processing system 200, for example, whether one engine decides to use data from the other engine. The feedback can be feedback regarding composition by the compositor 220, object detection by the object detection engine 225, semantic segmentation by the semantic segmentation engine 230, media modification by the media modification engine 235, gaze tracking by the gaze tracking engine 270, or a combination thereof. The feedback received by the feedback engine 260 can be positive feedback or negative feedback. For example, if one engine of the media processing system 200 uses data from another engine of the media processing system 200, the feedback engine 260 can interpret this as positive feedback. If one engine of the media processing system 200 rejects data from another engine of the media processing system 200, the feedback engine 260 may interpret this as negative feedback. Positive feedback may also be based on attributes of the sensor data from the user-facing sensor(s) 205, such as the user smiling, laughing, nodding, saying a positive statement (e.g., "yes," "confirmed," "OK," "next"), or otherwise reacting positively to the media. Negative feedback may also be based on attributes of the sensor data from the user-facing sensor(s) 205, such as the user grimacing, crying, shaking their head (e.g., in a "no" motion), saying a negative statement (e.g., "no," "different," "not good," "not this"), or otherwise reacting negatively to the virtual content.

[0097]

[0111] In some examples, feedback engine 260 provides feedback to one or more ML systems of media processing system 200 as training data for updating the one or more ML systems of media processing system 200. For example, feedback engine 260 may provide feedback as training data to the ML system(s) and / or trained ML model(s) of synthesizer 220, object detection engine 225, semantic segmentation engine 230, media modification engine 235, and / or gaze tracking engine 270. Positive feedback may be used to strengthen and / or strengthen weights associated with the output of the ML system(s) and / or trained ML model(s). Negative feedback may be used to weaken and / or remove weights associated with the output of the ML system(s) and / or trained ML model(s).

[0098]

[0112] In some examples, feedback engine 260 includes a software element, such as a set of instructions corresponding to a program executing on a processor, such as processor 1110, image processor 150, host processor 152, ISP 154, or a combination thereof, of computing system 1100. In some examples, feedback engine 260 includes one or more hardware elements. For example, feedback engine 260 may include a processor, such as processor 1110, image processor 150, host processor 152, ISP 154, or a combination thereof, of computing system 1100. In some examples, feedback engine 260 includes a combination of one or more software elements and one or more hardware elements.

[0099]

[0113] In some examples, the media processing system 200 may alter the segmentation of the environment by the semantic segmentation engine 230 and / or the portion(s) of the media that are obscured and / or receive attenuation correction from the media correction engine 235 based on the user's gaze (e.g., as detected by the gaze tracking engine 270), based on a gesture by the user (e.g., as detected by the gaze tracking engine 270 and / or the object detection engine 225), based on a command(s) spoken by the user (e.g., "blur this," "obscure this," "don't blur that," "don't obscure that"), or based on a combination thereof.

[0100]

[0114] 3A is a perspective view 300 showing a head mounted display (HMD) 310 used as an extended reality (XR) system 200. The HMD 310 may be, for example, an augmented reality (AR) headset, a virtual reality (VR) headset, a mixed reality (MR) headset, an extended reality (XR) headset, or some combination thereof. The HMD 310 may be an example of a media processing system 200. The HMD 310 includes a first camera 330A and a second camera 330B along the front of the HMD 310. The first camera 330A and the second camera 330B may be examples of the environment facing sensor 210 of the media processing system 200. The HMD 310 includes a third camera 330C and a fourth camera 330D that face the user's eye(s) when the user's eye(s) face the display(s) 340. The third camera 330C and the fourth camera 330D may be examples of user-facing sensors 205 of the media processing system 200. In some examples, the HMD 310 may have only a single camera with a single image sensor. In some examples, the HMD 310 may include one or more additional cameras in addition to the first camera 330A, the second camera 330B, the third camera 330C, and the fourth camera 330D. In some examples, the HMD 310 may include one or more additional sensors in addition to the first camera 330A, the second camera 330B, the third camera 330C, and the fourth camera 330D, which may also include other types of user-facing sensors 205 and / or environmental facing sensors 210 of the media processing system 200. In some examples, the first camera 330A, the second camera 330B, the third camera 330C, and / or the fourth camera 330D may be examples of the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.

[0101]

[0115] The HMD 310 may include one or more displays 340 visible to a user 320 wearing the HMD 310 on the user's 320 head. The one or more displays 340 of the HMD 310 may be examples of one or more displays of the output device(s) 240 of the media processing system 200. In some examples, the HMD 310 may include one display 340 and two viewfinders. The two viewfinders may include a left viewfinder for the left eye of the user 320 and a right viewfinder for the right eye of the user 320. The left viewfinder may be oriented so that the left eye of the user 320 sees the left side of the display. The right viewfinder may be oriented so that the right eye of the user 320 sees the right side of the display. In some examples, the HMD 310 may include two displays 340 including a left display that displays content to the left eye of the user 320 and a right display that displays content to the right eye of the user 320. The display(s) 340 of the HMD 310 may be a digital "pass-through" display or an optical "see-through" display.

[0102]

[0116] The HMD 310 may include one or more earpieces 335 that may function as speakers and / or headphones to output audio to one or more ears of a user of the HMD 310. Although one earpiece 335 is shown in FIGS. 3A and 3B, it should be understood that the HMD 310 may include two earpieces, one for each ear (left and right) of the user. In some examples, the HMD 310 may also include one or more microphones (not shown). The one or more microphones may be examples of the user-facing sensors 205 and / or the environment-facing sensors 210 of the media processing system 200. In some examples, the audio output by the HMD 310 through the one or more earpieces 335 to the user may include or be based on audio recorded using the one or more microphones.

[0103]

[0117] FIG. 3B is a perspective view 350 showing the head mounted display (HMD) of FIG. 3A being worn by a user 320. The user 320 wears the HMD 310 on the user's 320 head over the user's 320 eyes. The HMD 310 can capture images using a first camera 330A and a second camera 330B. In some examples, the HMD 310 displays one or more output images to the user's 320 eyes using a display(s) 340. In some examples, the output images can include virtual content generated by the virtual content generator 215, composited using the compositor 220, and / or displayed by the display(s) of the output device(s) 240. The output images can be based on images captured by the first camera 330A and the second camera 330B, for example, with the virtual content overlaid. The output images may provide a stereoscopic view of the environment, possibly with virtual content overlaid and / or other modifications. For example, the HMD 310 may display a first display image based on an image captured by the first camera 330A to the right eye of the user 320. The HMD 310 may display a second display image based on an image captured by the second camera 330B to the left eye of the user 320. For example, the HMD 310 may provide overlaid virtual content in the display image overlaid on the images captured by the first camera 330A and the second camera 330B. The third camera 330C and the fourth camera 330D may capture images of the eyes before, during, and / or after the user views the display image displayed by the display(s) 340. In this manner, sensor data from the third camera 330C and / or the fourth camera 330D can capture the reaction of the user's eyes (and / or other parts of the user) to the virtual content. The earpieces 335 of the HMD 310 are shown within the ears of the user 320.The HMD 310 may output audio to the user 320 through earpiece 335 and / or through another earpiece (not shown) of the HMD 310 in the other ear (not shown) of the user 320.

[0104]

[0118] 4A is a perspective view 400 showing the front of a mobile handset 410 that includes a forward-facing camera and can be used as an extended reality (XR) system 200. The mobile handset 410 may be an example of a media processing system 200. The mobile handset 410 may be, for example, a mobile phone, a satellite phone, a portable game console, a music player, a health tracking device, a wearable device, a wireless communication device, a laptop, a mobile device, any other type of computing device or computing system described herein, or a combination thereof.

[0105]

[0119] The front surface 420 of the mobile handset 410 includes a display 440. The front surface 420 of the mobile handset 410 includes a first camera 430A and a second camera 430B. The first camera 430A and the second camera 430B may be examples of user-facing sensors 205 of the media processing system 200. The first camera 430A and the second camera 430B may face the user, including the user's eye(s), while content (e.g., modified media output by the media modification engine 235) is displayed on the display 440. The display 440 may be an example of the display(s) of the output device(s) 240 of the media processing system 200.

[0106]

[0120] The first camera 430A and the second camera 430B are shown within a bezel around the display 440 on the front face 420 of the mobile handset 410. In some examples, the first camera 430A and the second camera 430B can be located in a notch or cutout cut out of the display 440 on the front face 420 of the mobile handset 410. In some examples, the first camera 430A and the second camera 430B can be under-display cameras located between the display 440 and the remainder of the mobile handset 410, so that light passes through a portion of the display 440 before reaching the first camera 430A and the second camera 430B. The first camera 430A and the second camera 430B in the perspective view 400 are forward-facing cameras. The first camera 430A and the second camera 430B face in a direction perpendicular to the plane of the front face 420 of the mobile handset 410. The first camera 430A and the second camera 430B may be two of one or more cameras of the mobile handset 410. In some examples, the front face 420 of the mobile handset 410 may have only a single camera.

[0107]

[0121] In some examples, the front surface 420 of the mobile handset 410 may include one or more additional cameras in addition to the first camera 430A and the second camera 430B. The one or more additional cameras may also be examples of the user-facing sensors 205 of the media processing system 200. In some examples, the front surface 420 of the mobile handset 410 may include one or more additional sensors in addition to the first camera 430A and the second camera 430B. The one or more additional sensors may also be examples of the user-facing sensors 205 of the media processing system 200. In some cases, the front surface 420 of the mobile handset 410 includes two or more displays 440. The one or more displays 440 of the front surface 420 of the mobile handset 410 may be examples of the display(s) of the output device(s) 240 of the media processing system 200. For example, the one or more displays 440 may include one or more touch screen displays.

[0108]

[0122] The mobile handset 410 may include one or more speakers 435A and / or other audio output devices (e.g., earphones or headphones or connectors thereto) that can output audio to one or more ears of a user of the mobile handset 410. While one speaker 435A is shown in FIG. 4A, it should be understood that the mobile handset 410 can include more than one speaker and / or other audio devices. In some examples, the mobile handset 410 can also include one or more microphones (not shown). The one or more microphones may be examples of the user-facing sensors 205 and / or the environmental facing sensors 210 of the media processing system 200. In some examples, the mobile handset 410 can include one or more microphones along and / or adjacent to the front surface 420 of the mobile handset 410, which are examples of the user-facing sensors 205 of the media processing system 200. In some examples, audio output by the mobile handset 410 to the user through one or more speakers 435A and / or other audio output devices may include or be based on audio recorded using one or more microphones.

[0109]

[0123] 4B is a perspective view 450 showing a rear view 460 of a mobile handset that includes a rear-facing camera and can be used as an extended reality (XR) system 200. The mobile handset 410 includes a third camera 430C and a fourth camera 430D on the rear view 460 of the mobile handset 410. The third camera 430C and the fourth camera 430D in the perspective view 450 are rear-facing. The third camera 430C and the fourth camera 430D may be examples of the environmental facing sensor 210 of the media processing system 200 of FIG. 2. The third camera 430C and the fourth camera 430D face in a direction perpendicular to the plane of the rear view 460 of the mobile handset 410.

[0110]

[0124] The third camera 430C and the fourth camera 430D may be two of the one or more cameras of the mobile handset 410. In some examples, the back surface 460 of the mobile handset 410 may have only a single camera. In some examples, the back surface 460 of the mobile handset 410 may include one or more additional cameras in addition to the third camera 430C and the fourth camera 430D. The one or more additional cameras may also be examples of the environmental facing sensors 210 of the media processing system 200. In some examples, the back surface 460 of the mobile handset 410 may include one or more additional sensors in addition to the third camera 430C and the fourth camera 430D. The one or more additional sensors may also be examples of the environmental facing sensors 210 of the media processing system 200. In some examples, the first camera 430A, the second camera 430B, the third camera 430C, and / or the fourth camera 430D may be examples of the image capture and processing system 100, the image capture device 105A, the image processing device 105B, or a combination thereof.

[0111]

[0125] The mobile handset 410 may include one or more speakers 435B and / or other audio output devices (e.g., earphones or headphones or connectors thereto) that can output audio to one or more ears of a user of the mobile handset 410. While one speaker 435B is shown in FIG. 4B, it should be understood that the mobile handset 410 can include more than one speaker and / or other audio device. In some examples, the mobile handset 410 can also include one or more microphones (not shown). The one or more microphones may be examples of the user-facing sensors 205 and / or the environmental facing sensors 210 of the media processing system 200. In some examples, the mobile handset 410 can include one or more microphones along and / or adjacent to the back surface 460 of the mobile handset 410, which are examples of the environmental facing sensors 210 of the media processing system 200. In some examples, audio output by the mobile handset 410 to the user through one or more speakers 435B and / or other audio output devices may include or be based on audio recorded using one or more microphones.

[0112]

[0126] The mobile handset 410 may use the display 440 on the front face 420 as a pass-through display. For example, the display 440 may display an output image. The output image may be based on an image captured by the third camera 430C and / or the fourth camera 430D, e.g., with virtual content overlaid and / or with modification by the media modification engine 235. The first camera 430A and / or the second camera 430B may capture images of the user's eye (and / or other parts of the user) before, during, and / or after the output image including the virtual content is displayed on the display 440. In this manner, sensor data from the first camera 430A and / or the second camera 430B may capture a reaction of the user's eye (and / or other parts of the user) to the virtual content.

[0113]

[0127] FIG. 5 is a block diagram illustrating a process 500 for event-based image processing. Process 500 is performed by a media processing system, such as media processing system 200, the media processing system of FIG. 6, and / or the media processing system of FIG. 10. Process 500 begins with image data 502 of an environment 505 being captured (e.g., by environmental facing sensor(s) 210) and / or received by the media processing system. The media processing system activates multiple media processing engines 510, such as a foveated compression engine 515 (which may be part of a media modification engine 235 that uses foveated compression to obscure), a gaze tracking engine 520 (which may be an example of a gaze tracking engine 270), an object detection engine 525 (which may be an example of an object detection engine 225 or an aspect thereof), and an audio recognition engine 530 (which may be an example of an audio aspect of the object detection engine 225 or a audio aspect thereof).

[0114]

[0128] The media processing system may perform event detection 535 in a first region 540 of the environment 505. The event detection 535 may include gaze detection 545 of a user's gaze looking at the first region 540 (using the gaze tracking engine 520), object detection 550 in the first region 540 (using the object detection engine 525), hand detection 555 of a user's hand in or pointing to the first region 540 (using the object detection engine 525), audio detection 560 of audio coming from the first region 540 and / or referencing the first region 540 or an object in the first region 540 (using the audio recognition engine 530), or combinations thereof. In response to the event detection 535, the media processing system may perform modification 565 of the image data 502 (e.g., using the media modification engine 235). The modification 565 may modify the first region 540 without modifying a second region 570 that is different from the first region 540, modify the second region 570 without modifying the first region 540, modify both the first region 540 and the second region 570, etc. The media processing system outputs the modified image data 575 of the environment 505, for example, by displaying the modified image data 575, playing audio corresponding to the modified image data 575, and / or transmitting the modified image data 575 to a recipient device using a communications transceiver.

[0115]

[0129] Figure 6 is a block diagram illustrating a process 600 for image processing based on detection of a person in image data. Process 600 is performed by a media processing system, such as media processing system 200, the media processing system of Figure 5, and / or the media processing system of Figure 10. Process 500 begins with image data 502 of an environment 505 being captured (e.g., by environment facing sensor(s) 210) and / or received by the media processing system. The media processing system performs object detection 550 to detect an object in the first region 540, for example using object detection engine 225 and / or object detection engine 525. In some examples, the object is a person 605.

[0116]

[0130] The media processing system performs image processing 610 based on detection of the person 605 (or other object). The image processing 610 may include face detection, recognition, and / or tracking 615 of the person 605. The image processing 610 may include semantic segmentation 620 of the face 625 and body 630 of the person 605. For example, the face 625 and body 630 of the person 605. The media processing system generates modified image data 575 based on the image modification 635 and outputs the modified image data 575. The media processing system outputs the modified image data 575 of the environment 505, for example, by displaying the modified image data 575, playing audio corresponding to the modified image data 575, and / or transmitting the modified image data 575 to a recipient device using a communication transceiver.

[0117]

[0131] The media processing system performs image modification 635 based on the image processing 610, for example, by applying blur 640 to the face 625, applying a reduced bit rate 645 to the face 625, applying increased compression 650 to the face 625, applying inpainting 655 to the face 625, applying pixelation 660 to the face 625, or a combination thereof.

[0118]

[0132] 7A is a conceptual diagram 700 illustrating an example of an image of an environment 705 and various modifications to the image to obscure portions of the environment, which are shown using dashed lines. The image of the environment 705 depicts a room with four people and a laptop 735. The four people include a person 730 and three other people. The image of the environment 705 is processed by a media processing system, such as the media processing system 200, the media processing system of FIG. 5, and / or the media processing system of FIG. 6. The image of the environment 705 is processed by the media processing system to generate a modified image of the environment 710, a modified image of the environment 715, and / or a modified image of the environment 720. The portions of the modified image of the environment 710, the modified image of the environment 715, and the modified image of the environment 720 that are obscured are shown with dashed black lines. The portions of the modified image of the environment 710, the modified image of the environment 715, and the modified image of the environment 720 that are not obscured are shown with solid black lines.

[0119]

[0133] In the modified image of the environment 710, everything in the room except the person 730 and laptop 735 has been obscured by the media modification engine 235. In some examples, the person 730 and laptop 735 appear on an approved list (e.g., a whitelist) and / or everything else in the room appears on a blocked list (e.g., a blacklist).

[0120]

[0134] In the modified image of environment 715, person 730 is obscured by media modification engine 235, but everything else in the room (including the three other people and the laptop) remains unobscured. In some examples, person 730 appears on a blocked list (e.g., a blacklist) and / or everything else in the room appears on an approved list (e.g., a whitelist).

[0121]

[0135] In the modified image of environment 720, the three people other than person 730 are obscured by media modification engine 235, but everything else in the room (including person 730 and laptop 735) remains unobscured. In some examples, person 730 appears on an approved list (e.g., a whitelist) and / or all others other than person 730 are on a blocked list (e.g., a blacklist).

[0122]

[0136] Figure 7B is a conceptual diagram 750 illustrating an example image of environment 705 and various modifications to the image to obscure portions of the environment shown using shading. In Figure 7B, the image of environment 705 is processed by a media processing system to generate a modified image of environment 755, a modified image of environment 760, and / or a modified image of environment 765. In the modified image of environment 755, the modified image of environment 760, and / or the modified image of environment 765, regions are obscured using gradual, gradient, and / or foveated obscuration techniques. Areas shaded using the darker shading pattern of FIG. 7B are more heavily obscured (e.g., more heavily blurred, compressed, pixelated, pixilated, tessellated, darkened, brightened, restored, scrambled, and / or reduced in resolution), while areas shaded using the lighter shading pattern of FIG. 7B are less obscured or remain unobscured.

[0123]

[0137] In the modified image of the environment 755, everything in the room has been obscured by the media modification engine 235 except for the person 730 and laptop 735. In some examples, the person 730 and laptop 735 appear on an approved list (e.g., a whitelist) and / or everything else in the room appears on a blocked list (e.g., a blacklist). The obscuration is progressive, and parts of the environment around the person 730 and laptop 735 remain unobscured or less obscured than other parts of the environment.

[0124]

[0138] In the modified image of environment 760, person 730 is obscured by media modification engine 235, but everything else in the room (including the three other people and the laptop) remains unobscured. In some examples, person 730 appears on a blocked list (e.g., blacklist) and / or everything else in the room appears on a blocked list (e.g., blacklist). The obscuration is progressive, and parts of the environment around person 730 are more obscured than other parts of the environment.

[0125]

[0139] In the modified image of environment 765, the faces of the three people other than person 730 are obscured by the media modification engine 235, while everything else in the room (including person 730 and laptop 735) remains unobscured. In some examples, person 730 appears on an approved list (e.g., a whitelist) and / or all other people other than person 730 are on a blocked list (e.g., a blacklist). The obscuration is progressive, and parts of the environment around the faces of the three people other than person 730 are more obscured than other parts of the environment.

[0126]

[0140] As described above, obscuring the region(s) in FIGS. 7A-7B may include blurring the region(s), pixelating the region(s), pixelating the region(s), mosaicizing the region(s), cropping the region(s), compressing the region(s) more strongly than other region(s) of the visual media data using image and / or video compression techniques, reducing the image resolution of the region(s) relative to the resolution of the other region(s) of the visual media data, quantizing the region(s) more strongly than other region(s) of the visual media data during image and / or video compression, removing the region(s), replacing the region(s) with other data (e.g., a color, a pattern, another image), inpainting the region(s) (e.g., using interpolation based on one or more surrounding pixels and / or regions), or combinations thereof.

[0127]

[0141] FIG. 8 is a conceptual diagram 800 illustrating an example soundscape of an environment 805 and various modifications to attenuate aspects of the soundscape corresponding to different elements within the environment. The soundscape of environment 805 is shown in FIG. 8 as a depiction of a room with four people and a laptop 735, who are also depicted in the images of environment 705 in FIGS. 7A-7B. The soundscape of environment 805 is processed by a media processing system, such as media processing system 200, the media processing system of FIG. 5, and / or the media processing system of FIG. 6. The soundscape of environment 805 is processed by the media processing system to generate a modified soundscape of environment 810, a modified soundscape of environment 815, and / or a modified soundscape of environment 820.

[0128]

[0142] The soundscape of environment 805, the modified soundscape of environment 810, the modified soundscape of environment 815, and the modified soundscape of environment 820 include speaker icons above each of the four people in the environments to indicate the sound(s) (e.g., voices) from each of the four people. The soundscape of environment 805, the modified soundscape of environment 810, the modified soundscape of environment 815, and the modified soundscape of environment 820 include a speaker icon above laptop 735 to indicate the sound(s) from laptop 735. The soundscape of environment 805, the modified soundscape of environment 810, the modified soundscape of environment 815, and the modified soundscape of environment 820 include a speaker icon in the upper left corner to indicate the sound(s) from the rest of the environment. The crossed out speaker icons represent sounds that have been attenuated, muted, and / or removed by the media modification engine 235. The speaker icons that are not crossed out represent sounds that remain unattenuated, muted, or removed by the media modification engine 235 .

[0129]

[0143] Portions of the modified soundscape of environment 810, the modified soundscape of environment 815, and the modified soundscape of environment 820 in which the corresponding sound(s) are attenuated, muted, and / or eliminated are shown with dashed black lines and illustrated thereon. Portions of the modified soundscape of environment 810, the modified soundscape of environment 815, and the modified soundscape of environment 820 in which the corresponding sound(s) are not attenuated, muted, and / or eliminated are shown with solid black lines.

[0130]

[0144] In the modified soundscape of the environment 810, sounds from everything in the environment other than the person 730 and laptop 735 in the room are attenuated, muted, and / or removed by the media modification engine 235. In some examples, the person 730 and laptop 735 appear on an approved list (e.g., a whitelist) and / or everything else in the room appears on a blocked list (e.g., a blacklist).

[0131]

[0145] In the modified soundscape of the environment 815, the sound from the person 730 is attenuated, muted, and / or removed by the media modification engine 235, while everything else in the room (including the three other people and the laptop) remains unattenuated, muted, or removed. In some examples, the person 730 appears on a blocked list (e.g., a blacklist) and / or everything else in the room appears on an approved list (e.g., a whitelist).

[0132]

[0146] In the modified soundscape of environment 820, sounds from the three people other than person 730 are attenuated, muted, and / or removed by media modification engine 235, while everything else in the room (including person 730 and laptop 735) remains unattenuated, muted, or removed. In some examples, person 730 appears on an approved list (e.g., a whitelist) and / or all others other than person 730 are on a blocked list (e.g., a blacklist).

[0133]

[0147] In some examples, the visual aspects of the media may be obscured, as in FIG. 7A or FIG. 7B, and the audio aspects of the media may be attenuated, muted, and / or removed, as in FIG. 8.

[0134]

[0148] 9 is a block diagram illustrating an example of a neural network (NN) 900 that may be used for media processing operations. The neural network 900 may include any type of deep network, such as a convolutional neural network (CNN), an autoencoder, a deep belief net (DBN), a recurrent neural network (RNN), a generative adversarial network (GAN), and / or other types of neural networks. The neural network 900 may be an example of one of the trained neural networks of the media processing system 200, the synthesizer 220, the object detection engine 225, the semantic segmentation engine 230, the media modification engine 235, and / or the gaze tracking engine 270, the foveated compression engine 515, the gaze tracking engine 520, the object detection engine 525, the audio recognition engine 530, the face tracking 615, the semantic segmentation 620, or combinations thereof.

[0135]

[0149] The input layer 910 of the neural network 900 includes input data. The input data of the input layer 910 can include data representing pixels of one or more input image frames. In some examples, the input data of the input layer 910 includes data representing pixels of image data (e.g., images captured by the user-facing sensor 205, media captured by the environment-facing sensor 210, virtual content generated by the virtual content generator 215, and / or composite images generated by the compositor 220), image(s) captured by the third camera 330C, image(s) captured by the fourth camera 330D, image(s) captured by the first camera 430A, image(s) captured by the second camera 430B, image data 502 of the environment, and / or metadata corresponding to the image data. In some examples, the input data of the input layer 910 includes gaze data from the gaze tracking engine 270, object detection data from the object detection engine 225, semantic segmentation data from the semantic segmentation engine 230, or a combination thereof.

[0136]

[0150] The image may include image data from an image sensor including raw pixel data (e.g., including a single color per pixel based on a Bayer filter) or processed pixel values ​​(e.g., RGB pixels for an RGB image). The neural network 900 includes multiple hidden layers 912A, 912B through 912N. The hidden layers 912A, 912B through 912N include "N" hidden layers, where "N" is an integer greater than or equal to 1. The number of hidden layers may be adapted to include as many layers as required for a given application. The neural network 900 further includes an output layer 914 that provides an output resulting from the processing performed by the hidden layers 912A, 912B through 912N.

[0137]

[0151] In some examples, the output layer 914 may provide an output image, such as a combined image generated by the compositor 220, modified media output by the media modification engine 235, modified image data 575 of the environment 505, or a combination thereof. In some examples, the output layer 914 may provide gaze data from the gaze tracking engine 270, object detection data from the object detection engine 225, semantic segmentation data from the semantic segmentation engine 230, or a combination thereof.

[0138]

[0152] Neural network 900 is a multi-layered neural network of interconnected filters. Each filter can be trained to learn features that represent the input data. Information related to the filters is shared between different layers, with each layer retaining the information as it is processed. In some cases, neural network 900 can include a feed-forward network, in which there are no feedback connections where the output of the network is fed back to itself. In some cases, network 900 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading the input.

[0139]

[0153] In some cases, information may be exchanged between layers through interconnections of nodes and nodes between various layers. In some cases, the network may include a convolutional neural network, which may not connect every node in one layer to every other node in the next layer. In a network in which information is exchanged between layers, the nodes of the input layer 910 may activate a set of nodes in the first hidden layer 912A. For example, as shown, each of the input nodes of the input layer 910 may be connected to each of the nodes of the first hidden layer 912A. The nodes of the hidden layer may transform the information of each input node by applying an activation function (e.g., a filter) to this information. The information derived from the transformation may then be passed to the nodes of the next hidden layer 912B, activating those nodes, which may perform their own designated function. Exemplary functions include convolution functions, downsampling, upscaling, data transformation, and / or any other suitable function. The output of the hidden layer 912B may then activate the nodes of the next hidden layer, and so on. The output of the final hidden layer 912N may activate one or more nodes in the output layer 914, which provides the processed output image. In some cases, a node in the neural network 900 (e.g., node 916) is shown as having multiple output lines, however, the node has a single output and all lines shown as outputting from the node represent the same output value.

[0140]

[0154] In some cases, each node or interconnection between nodes can have a weight, which is a set of parameters derived from training of the neural network 900. For example, the interconnections between nodes can represent information learned about the interconnected nodes. The interconnections can have adjustable numerical weights that can be adjusted (e.g., based on a training data set), allowing the neural network 900 to be adaptive to the input and to learn as more and more data is processed.

[0141]

[0155] The neural network 900 is pre-trained to process features from the data in the input layer 910 using different hidden layers 912A, 912B through 912N to provide output through the output layer 914.

[0142]

[0156] 10 is a flow diagram illustrating a process for a media processing operation. Process 1000 may be performed by a media processing system. In some examples, the media processing system may include, for example, image capture and processing system 100, image capture device 105A, image processing device 105B, image processor 150, ISP 154, host processor 152, media processing system 200, HMD 310, mobile handset 410, the media processing system of FIG. 5, the media processing system of FIG. 6, the media processing system of FIG. 7A, the media processing system of FIG. 7B, the media processing system of FIG. 8, neural network 900, computing system 1100, processor 1110, or a combination thereof.

[0143]

[0157] At operation 1005, the media processing system is configured to receive and may receive image data captured by the image sensor, the image data being representative of (e.g., depicting) an environment. In some examples, the media processing system includes an image sensor connector that couples and / or connects the image sensor to the remainder of the media processing system (e.g., including a processor and / or memory of the media processing system). In some examples, the media processing system receives image data from the image sensor by receiving image data from, through, and / or using the image sensor connector.

[0144]

[0158] Examples of image sensors include image sensor 130, user-facing sensor(s) 205, environment-facing sensor(s) 210, first camera 330A, second camera 330B, first camera 430A, second camera 430B, third camera 430C, fourth camera 430D, an image sensor that captures image data 502, an image sensor that captures images of the environment 705, an image sensor used to capture images used as input data for the input layer 910 of the NN 900, an input device 1145, another image sensor described herein, another sensor described herein, or a combination thereof.

[0145]

[0159] Examples of image data include image data captured using image capture and processing system 100, image data captured using an image sensor(s) of user facing sensor(s) 205, image data captured using an image sensor(s) of environment facing sensor(s) 210, image data captured using first camera 330A, image data captured using second camera 330B, image data captured using first camera 430A, image data captured using second camera 430B, image data captured using third camera 430C, image data captured using fourth camera 430D, image data 502, an image of the environment 705, an image used as input data for input layer 910 of NN 900, another image described herein, another set of image data described herein, or a combination thereof.

[0146]

[0160] Examples of environments include scene 110, the user facing user facing sensor(s) 205, the environment facing environment facing sensor(s) 210, the environment in which HMD 310 is present, the environment in which first camera 330A and / or second camera 330B capture image data, the environment in which mobile handset 410 is present, the environment in which first camera 430A and / or second camera 430B and / or third camera 430C and / or fourth camera 430D capture image data of environment 505, an environment depicted in an image of environment 705, an environment represented in a soundscape of environment 805, another environment or scene described herein, or a combination thereof.

[0147]

[0161] At operation 1010, the media processing system is configured to receive and may receive an indication of an object in the environment represented in the image data. In some examples, the object represented in the image data includes an object depicted in the image data. Examples of objects include an object detected using object detection engine 225, an object detected using object detection engine 525, an object corresponding to event detection 535, an object outputting audio detected by audio recognition engine 530, an object detected using object detection 550, a hand detected using hand detection 555, an object outputting audio detected by audio detection 560, a person 605, a face tracked in face tracking 615, a face 625, a body 630, a person 730, a laptop 735, other people in the image of environment 705, other objects described herein, or combinations thereof. In some examples, the object may be a person, an animal, a vehicle, a plant, a structure, a device, content displayed on a device, printed content printed on a medium, written content written on a medium, drawn content drawn on a medium, or combinations thereof.

[0148]

[0162] In some aspects, receiving an indication of an object in the environment includes detecting the object in the image data, for example using the object detection engine 225, event detection 535, audio recognition engine 530, object detection 550, hand detection 555, audio detection 560, face tracking 615, NN900, or a combination thereof.

[0149]

[0163] In some aspects, receiving an indication of an object in the environment includes input through a user interface, the input indicating the object. In some examples, the input through the user interface can be touch input through a touchscreen interface, touch input through a trackpad interface, click input through a mouse interface, button input through a button interface, keyboard input through a keyboard interface, keypad input through a keypad interface, gaze input through user-facing sensor(s) 205 that is interpreted using gaze tracking engine 270, voice command input through a microphone and / or speech recognition system, text command input using a keyboard.

[0150]

[0164] At operation 1015, the media processing system is configured and capable of dividing the image data into a plurality of regions. The plurality of regions includes a first region and a second region. The object is represented in one of the plurality of regions. In some examples, dividing the image data into the plurality of regions is performed by the semantic segmentation engine 230 and / or the semantic segmentation 620. In some examples, one of the plurality of regions in which the object is depicted is the first region. In some examples, one of the plurality of regions in which the object is depicted is the second region. Examples of the plurality of regions include the regions of the images of FIGS. 7A, 7B, and 8. For example, examples of the multiple regions include regions corresponding to person 730, laptop 735, other people in the image of environment 705, other objects in the image of environment 705, background areas in the image of environment 705, regions of the image of environment 705, regions of different contours in the modified image of environment 710, regions of different contours in the modified image of environment 715, regions of different contours in the modified image of environment 720, regions of different shading in the modified image of environment 755, regions of different shading in the modified image of environment 760, regions of different shading in the modified image of environment 765, regions of different contours in the modified soundscape of environment 810, regions of different contours in the modified soundscape of environment 815, and regions of different contours in the modified soundscape of environment 820.

[0151]

[0165] In some aspects, dividing the image data into a plurality of regions includes dividing the image data into a plurality of regions based on a determined location of an object. The object is located in at least one region and is not located in at least one other region. In some aspects, the location of the object is determined from the image data, for example, based on object detection. In some aspects, the media processing system is configured to detect audio, and can detect the audio, and the location of the object is determined based on attributes of the audio, the attributes including at least one of a location of the audio, a direction of the audio, an amplitude of the audio, or a frequency of the audio.

[0152]

[0166] In some embodiments, at least one region is a region having a predetermined shape, such as a square, a rectangle, a circle, a triangle, a polygon, a portion of any of these shapes, or a combination thereof.

[0153]

[0167] At operation 1020, the media processing system is configured to, and can, modify the image data to obscure a first region without obscuring a second region based on the object being represented in one of the multiple regions. Examples of modifying the image data include media modification engine 235, modification 565, image modification 635, blurring 640, reduced bit rate 645, increased compression 650, inpainting 655, pixelation 660, modified image data 575, modified image of environment 710, modified image of environment 715, modified image of environment 720, modified image of environment 755, modified image of environment 760, modified image of environment 765, modified soundscape of environment 810, modified soundscape of environment 815, and modified soundscape of environment 820.

[0154]

[0168] In some aspects, the object is represented in a first region and not represented in a second region, and modifying the image data to obscure the first region is based on the object being represented in the first region and / or not being represented (e.g., missing from) the second region.

[0155]

[0169] In some aspects, the object is represented in the second region and not represented in the first region, and modifying the image data to obscure the first region without obscuring the second region is based on the object being represented in the second region and / or not being represented in the first region (e.g., being missing from the first region).

[0156]

[0170] In some aspects, modifying the image data to obscure the first region includes modifying the image data using foveated compression of a peripheral area around the fixation point. In some examples, the second region includes the fixation point and the first region includes the peripheral area, such as in the modified image of environment 755. In some examples, the first region includes the fixation point while the second region includes the peripheral area, such as in the modified image of environment 760 and the modified image of environment 765.

[0157]

[0171] In some aspects, modifying the image data to obscure the first region includes modifying the image data to blur at least a portion of the first region, to remove at least a portion of the first region, to inpaint at least a portion of the first region, to pixelate or pixelate at least a portion of the first region, or a combination thereof.

[0158]

[0172] In some embodiments, modifying the image data to obscure the first region includes modifying the image data to reduce resolution of a first subset of image data depicting the first region compared to a second subset of image data depicting the second region. In some embodiments, modifying the image data to obscure the first region includes modifying the image data to compress the first subset of image data depicting the first region more than the second subset of image data depicting the second region.

[0159]

[0173] In some aspects, modifying the image data to obscure the first region reduces an amount of data used to code the first region. In some aspects, modifying the image data to obscure the first region includes at least one of increasing compression in the first region, increasing quantization in the first region, reducing resolution in the first region, cropping the first region, and / or pixelating the first region.

[0160]

[0174] At operation 1025, the media processing system is configured to and may output the image data after modifying the image data. In some aspects, outputting the image data includes displaying the image data using output device(s) 240, such as a display. In some aspects, outputting the image data includes transmitting the image data to a recipient device using a communication transceiver, such as transceiver(s) 245 and / or communication interface 1140.

[0161]

[0175] In some aspects, the object includes at least a portion of a person's body, such as in the modified image of environment 710, the modified image of environment 715, the modified image of environment 720, the modified image of environment 755, the modified image of environment 760, and the modified image of environment 765. In some aspects, the object includes at least a portion of a person's face, such as in the modified image of environment 765. In some aspects, the object includes at least a portion of a string of characters, such as, for example, string of characters displayed on laptop 735. In some aspects, the object includes at least a portion of content displayed using a display, such as, for example, content displayed using the display of laptop 735.

[0162]

[0176] In some aspects, the media processing system is configured to receive and can receive audio data captured by a microphone from the environment. The audio data is captured at a time corresponding to the capture of the image data. The media processing system detects audio samples corresponding to the object in the audio data. The media processing system modifies the audio data to attenuate the audio samples corresponding to the object and outputs the audio data after modifying the audio data. Examples of such modifications of the audio data include a modified soundscape of the environment 810, a modified soundscape of the environment 815, and a modified soundscape of the environment 820. In some aspects, outputting the audio data includes playing the image data using the output device(s) 240, such as speakers and / or headphones. In some aspects, outputting the audio data includes transmitting the audio data to a recipient device using a communication transceiver, such as the transceiver(s) 245 and / or the communication interface 1140.

[0163]

[0177] In some aspects, the media processing system is configured to receive and can receive secondary image data from a second image sensor. Examples of secondary image data include any of the examples listed above for image data. Examples of the second image sensor include any of the examples listed above for the image sensor and the third camera 330C and / or the fourth camera 330D. The second image sensor has a different field of view than the image sensor. The secondary image data captured by the second image sensor includes a secondary image of the user. The segmentation of the image data in operation 1015 is further based on the secondary image. In some aspects, the second image sensor captures a gesture or position of at least a portion of the user, and segmenting the image data includes defining a region corresponding to a direction and / or position of the gesture of at least a portion of the user. In some aspects, the user's gesture or position includes the user's gaze direction, for example, as determined by the gaze tracking engine 270 when the secondary image sensor is one of the user-facing sensor(s) 205 (e.g., the third camera 330C, the fourth camera 330D, the first camera 430A, the second camera 430B).

[0164]

[0178] In some aspects, the media processing system is configured and capable of identifying the object. The media processing system can determine whether the object is displayed or obscured based on identifying the object. In some examples, the media processing system can define a first region to include the object in response to determining whether the object is obscured or displayed. In some aspects, determining that the object is obscured includes determining that the object is included in a blacklist of objects to be obscured and / or determining that the object is not included in a whitelist of objects to be displayed. In some aspects, determining that the object is displayed includes determining that the object is included in a whitelist of objects to be displayed and / or determining that the object is not included in a blacklist of objects to be obscured.

[0165]

[0179] In some examples, a media processing system may include means for receiving image data captured by an image sensor, where the image data depicts an environment; means for receiving an indication of an object in the environment represented in the image data; means for dividing the image data into a plurality of regions, where the plurality of regions includes a first region and a second region, where the object is represented in one of the plurality of regions; means for modifying the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions; and means for outputting the image data after modifying the image data.

[0166]

[0180] In some examples, the means for receiving image data includes image capture and processing system 100, image capture device 105A, image processing device 105B, image processor 150, ISP 154, host processor 152, image sensor 130, user facing sensor(s) 205, environment facing sensor(s) 210, first camera 330A, second camera 330B, first camera 430A, second camera 430B, third camera 430C, fourth camera 430D, an image sensor that captures image data 502, an image sensor that captures an image of the environment 705, an image sensor used to capture images used as input data for input layer 910 of NN 900, input device 1145, another image sensor described herein, another sensor described herein, or a combination thereof.

[0167]

[0181] In some examples, the means for receiving an indication of an object in the environment includes image processor 150, ISP 154, host processor 152, object detection engine 225, object detection engine 525, event detection 535, audio recognition engine 530, object detection 550, hand detection 555, audio detection 560, face tracking 615, NN 900, computing system 1100, processor 1110, or combinations thereof.

[0168]

[0182] In some examples, the means for dividing the image data into multiple regions includes the image processor 150, the ISP 154, the host processor 152, the semantic segmentation engine 230, the semantic segmentation 620, the NN 900, the computing system 1100, the processor 1110, or a combination thereof.

[0169]

[0183] In some examples, the means for modifying image data includes image processor 150, ISP 154, host processor 152, media modification engine 235, modification 565, image modification 635, blurring 640, reduced bit rate 645, increased compression 650, restoration 655, pixelation 660, NN 900, computing system 1100, processor 1110, or combinations thereof.

[0170]

[0184] In some examples, the means for outputting image data includes the image processor 150, the ISP 154, the host processor 152, the output device(s) 240, the transceiver(s) 245, the computing system 1100, the output device 1135, the communication interface 1140, or a combination thereof.

[0171]

[0185] In some examples, the processes described herein (e.g., process 500, process 600, and process 1000, as well as the processes of Figures 1, 2, 7A, 7B, 8, 9, and / or 11, and / or other processes described herein) may be performed by a computing device or apparatus. In some examples, the processes described herein may be performed by processing system 100, image capture device 105A, image processing device 105B, image processor 150, ISP 154, host processor 152, media processing system 200, HMD 310, mobile handset 410, the media processing system of Figure 5, the media processing system of Figure 6, the media processing system of Figure 7A, the media processing system of Figure 7B, the media processing system of Figure 8, neural network 900, the media processing system of Figure 11, computing system 1100, processor 1110, or a combination thereof.

[0172]

[0186] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, a wearable device (e.g., a VR headset, an AR headset, AR glasses, a network-connected watch or smartwatch, or other wearable device), a server computer, an autonomous vehicle or autonomous vehicle computing device, a robotic device, a television, and / or any other computing device having resource capabilities to perform the processes described herein. In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) configured to perform steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other component(s). The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0173]

[0187] Components of a computing device may be implemented with circuitry. For example, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuitry (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuitry), and / or may include and / or be implemented using computer software, firmware, or any combination thereof, to perform various operations described herein.

[0174]

[0188] The processes described herein are illustrated as logical flow diagrams, block diagrams, or conceptual diagrams, whose operations represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the described operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement a process.

[0175]

[0189] Additionally, the processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that collectively execute on one or more processors, by hardware, or a combination thereof. As mentioned above, the code may be stored on a computer-readable or machine-readable storage medium, for example in the form of a computer program comprising a number of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0176]

[0190] Fig. 11 illustrates an example of a system for implementing certain aspects of the present technology. In particular, Fig. 11 illustrates an example of a computing system 1100, which can be, for example, any computing device constituting an internal computing system, a remote computing system, a camera, or any of its components whose components communicate with each other using a connection 1105. The connection 1105 can be a physical connection using a bus, or a direct connection to a processor 1110, such as in a chipset architecture. The connection 1105 can also be a virtual connection, a network connection, or a logical connection.

[0177]

[0191] In some embodiments, computing system 1100 is a distributed system in which the functionality described in this disclosure may be distributed within a data center, multiple data centers, a peer network, etc. In some embodiments, one or more of the system components described represent many such components, each performing some or all of the functionality described by that component. In some embodiments, the components may be physical or virtual devices.

[0178]

[0192] The exemplary system 1100 includes at least one processing unit 1110 (CPU or processor) and connections 1105 coupling various system components to the processor 1110, including system memory 1115, such as read only memory (ROM) 1120 and random access memory (RAM) 1125. The computing system 1100 may include a cache 1112 of high-speed memory directly connected to, adjacent to, or integrated as part of the processor 1110.

[0179]

[0193] The processor 1110 may include any general purpose processor, hardware or software services such as services 1132, 1134, and 1136 stored in a storage device 1130 configured to control the processor 1110, and special purpose processors where software instructions are embedded in the actual processor design. The processor 1110 may essentially be a completely self-contained computing system including multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0180]

[0194] To enable user interaction, computing system 1100 includes input devices 1145, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. Computing system 1100 can also include output devices 1135, which can be one or more of a number of output mechanisms. In some cases, a multi-modal system may enable a user to provide multiple types of input / output to communicate with computing system 1100. Computing system 1100 can include a communication interface 1140, which can generally govern and manage user input and system output.The communications interface may be any of the following: audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, Apple® Lightning® port / plug, Ethernet port / plug, fiber optic port / plug, proprietary wired port / plug, BLUETOOTH® wireless signal transmission, BLUETOOTH® low energy (BLE) wireless signal transmission, IBEACON® wireless signal transmission, radio-frequency identification (RFID) wireless signal transmission, near-field communications (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WLAN), and Bluetooth® wireless signal transmission. The wireless communication device may perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including those utilizing WiMAX (Wireless Access), infrared (IR) communications wireless signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad-hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or any combination thereof.The communication interface 1140 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine the location of the computing system 1100 based on reception of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite system (BDS), and the European Galileo GNSS. There is no constraint to operate with any particular hardware arrangement, and therefore the basic features herein may be easily substituted for improved hardware or firmware arrangements as they are developed.

[0181]

[0195] The storage device 1130 may be a non-volatile and / or non-transitory and / or computer readable memory device, such as a magnetic cassette, a flash memory card, a solid state memory device, a digital versatile disk, a cartridge, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, a flash memory, a memristor memory, any other solid state memory, a compact disc read only memory (CD-ROM) optical disk, a rewritable compact disc (CD) optical disk, a digital video disk (DVD) optical disk, a blu-ray disc (BDD) optical disk, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a memory stick card, a smart card chip, an EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (ICC), a circuit, IC chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random-access memory (RRAM),The memory may be a hard disk or other type of computer readable medium capable of storing data that is accessible by a computer, such as a memory, RRAM / ReRAM, phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.

[0182]

[0196] The storage devices 1130 may include software services, servers, services, etc. that cause the system to perform functions when code defining such software is executed by the processor 1110. In some embodiments, hardware services that perform certain functions may include software components stored in a computer-readable medium in conjunction with the necessary hardware components, such as the processor 1110, connections 1105, output devices 1135, etc., to perform the functions.

[0183]

[0197] The term "computer-readable medium" as used herein includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, storing, or transporting instruction(s) and / or data. Computer-readable media may also include non-transitory media on which data may be stored and which do not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact disks (CDs) or digital versatile disks (DVDs), flash memory, memory, or memory devices. Computer-readable media may have code and / or machine-executable instructions stored thereon, which may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted using any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0184]

[0198] In some embodiments, computer readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when mentioned, non-transitory computer readable storage media specifically excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0185]

[0199] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, it will be understood by those skilled in the art that the embodiments may be practiced without these specific details. For ease of explanation, in some cases, the present technology may be presented as including individual functional blocks, including devices, device components, steps or routines in a method embodied in software, or functional blocks comprising a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form so as not to obscure the embodiments in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail so as to avoid obscuring the embodiments.

[0186]

[0200] Particular embodiments may be described above as a process or method that is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although the flowcharts may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagrams. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or to the main function.

[0187]

[0201] The processes and methods according to the examples described above may be implemented using computer-executable instructions stored on or otherwise available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions, or in some cases configure a general-purpose computer, a special-purpose computer, or a processing device to perform a function or group of functions. Portions of the computer resources used may be accessible over a network. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, network-attached storage devices, etc.

[0188]

[0202] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) performing the necessary tasks may be stored in a computer-readable or machine-readable medium. A processor or processors may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices or other small-footprint personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, and the like. The functionality described herein may also be embodied in peripheral devices or add-in cards. Such functionality may also be implemented on a circuit board among different chips, or on different processes executing in a single device, as further examples.

[0189]

[0203] The instructions, media for propagating such instructions, computing resources for executing such instructions, and other structures supporting such computing resources are exemplary means for providing the functionality described in this disclosure.

[0190]

[0204] In the above description, aspects of the present application are described with reference to specific embodiments thereof, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while exemplary embodiments of the present application have been described in detail herein, it should be understood that the inventive concepts may be embodied and employed in various other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. The various features and aspects of the present application described above may be used individually or jointly. Moreover, the embodiments may be utilized in any number of environments and applications other than those described herein without departing from the broader spirit and scope of the present specification. Thus, the present specification and drawings should be regarded as illustrative and not restrictive. For purposes of illustration, the methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in an order different from that described.

[0191]

[0205] Those skilled in the art will appreciate that the less than ("<") and greater than (">") symbols or terminology used herein may be replaced with the less than or equal to ("≦") and greater than or equal to ("≧") symbols, respectively, without departing from the scope of the present specification.

[0192]

[0206] When a component is described as being "configured to" perform certain operations, such configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuitry (e.g., a microprocessor or other suitable electronic circuitry) to perform the operations, or any combination thereof.

[0193]

[0207] The phrase "coupled to" refers to any component that is physically connected, either directly or indirectly, to another component and / or that is in communication, either directly or indirectly, with another component (e.g., connected to the other component via a wired or wireless connection and / or other suitable communication interface).

[0194]

[0208] Claim language or other language reciting "at least one of" a set and / or "one or more" of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, a claim language reciting "at least one of A and B" means A, B, or A and B. In another example, a claim language reciting "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more" of a set does not limit the set to the items listed in the set. For example, a claim language reciting "at least one of A and B" can mean A, B, or A and B, and can further include items not listed in the set of A and B.

[0195]

[0209] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present application.

[0196]

[0210] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device, or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise a memory or data storage medium, such as a random access memory (RAM), such as a synchronous dynamic random access memory (SDRAM), a read-only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read-only memory (EEPROM), a FLASH memory, a magnetic or optical data storage medium, etc. The techniques may additionally or alternatively be realized at least in part by a computer-readable communications medium, such as a propagated signal or wave, that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer.

[0197]

[0211] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but the processor may alternatively be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, the term "processor" as used herein may refer to any of the above structures, any combination of the above structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects the functionality described herein may be provided in dedicated software or hardware modules configured for encoding and decoding, or may be incorporated within a combined video encoder-decoder (CODEC).

[0198]

[0212] Exemplary aspects of the present disclosure include the following.

[0199]

[0213] Aspect 1: An apparatus for media processing, the apparatus comprising: a memory; and one or more processors coupled to the memory, the one or more processors configured to receive image data depicting an environment captured by an image sensor; receive an indication of an object in the environment represented in the image data; divide the image data into a plurality of regions including a first region and a second region, where the object is represented in one of the plurality of regions; modify the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions; and output the image data after modifying the image data.

[0200]

[0214] Aspect 2. The apparatus of aspect 1, wherein to divide the image data into multiple regions, one or more processors are configured to divide the image data into multiple regions based on a determined location of the object, and the object is located in at least one region and not located in at least one other region.

[0201]

[0215] Example 3. The apparatus of example 1 or 2, wherein the location of the object is determined from image data.

[0202]

[0216] Aspect 4. An apparatus as described in any of aspects 1 to 3, wherein the one or more processors are configured to detect audio, and the location of the object is determined based on attributes of the audio, the attributes including at least one of the location of the audio, the direction of the audio, the amplitude of the audio, or the frequency of the audio.

[0203]

[0217] Aspect 5. An apparatus as described in any of Aspects 1-4, wherein the one or more processors are configured to detect objects within the image data to receive an indication of an object within the environment.

[0204]

[0218] Aspect 6. An apparatus as described in any of aspects 1-5, wherein to receive an indication of an object within the environment, one or more processors are configured to receive input via a user interface, the input indicating the object.

[0205]

[0219] Embodiment 7. An apparatus described in any of embodiments 1 to 6, wherein an object is represented in a first region and not in a second region, and modifying the image data to obscure the first region is based on the object being depicted in the first region.

[0206]

[0220] Example 8. An apparatus as described in any of Examples 1 to 7, wherein the object is represented in the second region and not represented in the first region, and modifying the image data to obscure the first region without obscuring the second region is based on the object being depicted in the second region and not represented in the first region.

[0207]

[0221] Aspect 9. An apparatus as described in any of aspects 1-8, wherein modifying the image data to obscure the first region includes modifying the image data using foveated compression of a peripheral area around the fixation point, wherein the second region includes the fixation point and the first region includes the peripheral area.

[0208]

[0222] Embodiment 10. The apparatus of any of embodiments 1-9, wherein modifying the image data to obscure the first region includes modifying the image data to blur at least a portion of the first region.

[0209]

[0223] Embodiment 11. An apparatus as described in any of embodiments 1-10, wherein modifying the image data to obscure the first region includes modifying the image data to remove at least a portion of the first region.

[0210]

[0224] Example 12. An apparatus as described in any of Examples 1 to 11, wherein modifying the image data to obscure the first region includes modifying the image data to repair at least a portion of the first region.

[0211]

[0225] Embodiment 13. The apparatus of any of embodiments 1-12, wherein modifying the image data to obscure the first region includes modifying the image data to pixelate at least a portion of the first region.

[0212]

[0226] Embodiment 14. An apparatus as described in any of embodiments 1-13, wherein modifying the image data to obscure the first region includes modifying the image data to reduce resolution of a first subset of image data depicting the first region compared to a second subset of image data depicting the second region.

[0213]

[0227] Embodiment 15. The apparatus of any of embodiments 1-14, wherein modifying the image data to obscure the first region includes modifying the image data to compress a first subset of image data depicting the first region more than a second subset of image data depicting the second region.

[0214]

[0228] Aspect 16. An apparatus described in any of aspects 1 to 15, wherein the object includes at least a portion of a person.

[0215]

[0229] Aspect 17. An apparatus described in any of aspects 1 to 16, wherein the object includes at least a portion of a person's face.

[0216]

[0230] Aspect 18. An apparatus described in any of aspects 1 to 17, wherein the object includes at least a portion of a character string.

[0217]

[0231] Aspect 19. An apparatus according to any of aspects 1 to 18, wherein the object comprises at least a portion of content displayed using the display.

[0218]

[0232] Aspect 20. The device of any of aspects 1 to 19, further comprising a display, and wherein the one or more processors are configured to display the image data using the display to output the image data.

[0219]

[0233] Aspect 21. The apparatus of any of aspects 1 to 20, further comprising a communications transceiver, wherein the one or more processors are configured to transmit the image data to a recipient device using the communications transceiver to output the image data.

[0220]

[0234] Aspect 22. An apparatus as described in any of aspects 1 to 21, wherein one or more processors are configured to receive audio data captured by a microphone from an environment at a time corresponding to the capture of image data, detect audio samples corresponding to an object within the audio data, modify the audio data to attenuate the audio samples corresponding to the object, and output the audio data after modifying the audio data.

[0221]

[0235] Embodiment 23. A device according to any one of embodiments 1 to 22, wherein at least one region is a region having a predetermined shape.

[0222]

[0236] Aspect 24. An apparatus as described in any of aspects 1 to 23, wherein the one or more processors are configured to receive secondary image data from a second image sensor, the second image sensor having a different field of view than the first image sensor, the secondary image data captured by the second image sensor includes a secondary image of the user, and the segmentation of the image data is further based on the secondary image.

[0223]

[0237] Aspect 25. The device of aspect 24, wherein the second image sensor captures a gesture or position of at least a portion of the user, and dividing the image data includes defining a region corresponding to a direction and / or position of the gesture of at least a portion of the user.

[0224]

[0238] Aspect 26. The device of aspect 25, wherein the user's gesture or position includes the user's gaze direction.

[0225]

[0239] Example 27. An apparatus described in any of examples 1 to 26, wherein modifying the image data to obscure the first region reduces the amount of data used to code the first region.

[0226]

[0240] Aspect 28. The apparatus of aspect 27, wherein modifying the image data to obscure the first region includes at least one of increasing compression in the first region, increasing quantization in the first region, reducing resolution in the first region, cropping the first region, and / or pixelating the first region.

[0227]

[0241] Example 29. The apparatus of any of examples 1 to 28, further comprising: identifying an object; determining whether the detected object is to be displayed or obscured; and, when the object is determined to be obscured, defining a first region to include the object.

[0228]

[0242] Aspect 30. The apparatus of aspect 29, wherein determining that the object is to be obscured includes determining that the object is included in a blacklist of objects to be obscured and / or determining that the object is not included in a whitelist of objects to be displayed.

[0229]

[0243] Aspect 31. A method for media processing, the method including: receiving image data depicting an environment captured by an image sensor; receiving an indication of an object in the environment represented in the image data; dividing the image data into a plurality of regions, the plurality of regions including a first region and a second region, the object being represented in one of the plurality of regions; modifying the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions; and outputting the image data after modifying the image data.

[0230]

[0244] Aspect 32. The method of aspect 31, wherein dividing the image data into a plurality of regions includes dividing the image data into a plurality of regions based on a determined location of the object, the object being located in at least one region and not located in at least one other region.

[0231]

[0245] Aspect 33. The method of aspect 31 or 32, wherein the location of the object is determined from image data.

[0232]

[0246] Aspect 34. A method as described in any of aspects 31 to 33, further comprising detecting audio, wherein the location of the object is determined based on attributes of the audio, the attributes including at least one of the location of the audio, the direction of the audio, the amplitude of the audio, or the frequency of the audio.

[0233]

[0247] Example 35. The method of any of examples 31 to 34, wherein receiving an indication of an object within the environment includes detecting the object within the image data.

[0234]

[0248] Aspect 36. The method of any of aspects 31-35, wherein receiving an indication of an object in the environment includes input via a user interface, the input indicating the object.

[0235]

[0249] Embodiment 37. A method according to any of embodiments 31 to 36, wherein an object is represented in a first region and not represented in a second region, and modifying the image data to obscure the first region is based on the object being depicted in the first region.

[0236]

[0250] Embodiment 38. A method according to any of embodiments 31 to 37, wherein the object is represented in the second region and not represented in the first region, and modifying the image data to obscure the first region without obscuring the second region is based on the object being depicted in the second region and not represented in the first region.

[0237]

[0251] Aspect 39. A method according to any of aspects 31 to 38, wherein modifying the image data to obscure the first region includes modifying the image data using foveated compression of a peripheral area around the fixation point, wherein the second region includes the fixation point and the first region includes the peripheral area.

[0238]

[0252] Embodiment 40. The method of any of embodiments 31-39, wherein modifying the image data to obscure the first region includes modifying the image data to blur at least a portion of the first region.

[0239]

[0253] Embodiment 41. The method of any of embodiments 31-40, wherein modifying the image data to obscure the first region includes modifying the image data to remove at least a portion of the first region.

[0240]

[0254] Embodiment 42. A method according to any of embodiments 31 to 41, wherein modifying the image data to obscure the first region includes modifying the image data to repair at least a portion of the first region.

[0241]

[0255] Embodiment 43. The method of any of embodiments 31-42, wherein modifying the image data to obscure the first region includes modifying the image data to pixelate at least a portion of the first region.

[0242]

[0256] Embodiment 44. A method according to any of embodiments 31 to 43, wherein modifying the image data to obscure the first region includes modifying the image data to reduce resolution of a first subset of image data depicting the first region compared to a second subset of image data depicting the second region.

[0243]

[0257] Embodiment 45. The method of any of embodiments 31-44, wherein modifying the image data to obscure the first region includes modifying the image data to compress a first subset of image data depicting the first region more than a second subset of image data depicting the second region.

[0244]

[0258] Example 46. A method according to any of examples 31 to 45, wherein the object includes at least a portion of a person.

[0245]

[0259] Example 47. A method according to any of examples 31 to 46, wherein the object includes at least a portion of a person's face.

[0246]

[0260] Aspect 48. A method according to any one of aspects 31 to 47, wherein the object includes at least a portion of a character string.

[0247]

[0261] Example 49. The method of any of examples 31 to 48, wherein the object comprises at least a portion of the content to be displayed using the display.

[0248]

[0262] Embodiment 50. The method of any of embodiments 31 to 49, wherein outputting the image data includes displaying the image data using a display.

[0249]

[0263] Embodiment 51. The method of any of embodiments 31 to 50, wherein outputting the image data includes transmitting the image data to a recipient device using a communications transceiver.

[0250]

[0264] Aspect 52. A method according to any of aspects 31 to 51, further comprising: receiving audio data captured by a microphone from an environment at a time corresponding to the capture of image data; detecting audio samples corresponding to the object within the audio data; modifying the audio data to attenuate the audio samples corresponding to the object; and outputting the audio data after modifying the audio data.

[0251]

[0265] Embodiment 53. The method of any one of embodiments 31 to 52, wherein at least one region is a region having a predetermined shape.

[0252]

[0266] Aspect 54. The method of any of aspects 31 to 53, further comprising receiving secondary image data from a second image sensor, the second image sensor having a different field of view than the first image sensor, the secondary image data captured by the second image sensor including a secondary image of the user, and the segmentation of the image data is further based on the secondary image.

[0253]

[0267] Aspect 55. The method of aspect 54, wherein a second image sensor captures a gesture or position of at least a portion of the user, and dividing the image data includes defining a region corresponding to the direction and / or position of the gesture of at least a portion of the user.

[0254]

[0268] Aspect 56. A method according to any of aspects 31 to 55, wherein the user's gesture or position includes the user's gaze direction.

[0255]

[0269] Embodiment 57. A method according to any of embodiments 31 to 56, wherein modifying the image data to obscure the first region reduces the amount of data used to code the first region.

[0256]

[0270] Aspect 58. The method of aspect 57, wherein modifying the image data to obscure the first region includes at least one of increasing compression in the first region, increasing quantization in the first region, reducing resolution in the first region, cropping the first region, and / or pixelating the first region.

[0257]

[0271] Aspect 59. A method according to any of aspects 31 to 58, further comprising: identifying an object; determining whether the object is to be displayed or obscured based on identifying the object; and in response to determining that the object is to be obscured, defining a first region to include the object.

[0258]

[0272] Aspect 60. The method of aspect 59, wherein determining that the object is to be obscured includes determining that the object is included in a blacklist of objects to be obscured and / or determining that the object is not included in a whitelist of objects to be displayed.

[0259]

[0273] Aspect 61: A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to receive image data depicting an environment captured by an image sensor, receive an indication of an object in the environment represented in the image data, divide the image data into a plurality of regions including a first region and a second region, where the object is represented in one of the plurality of regions, modify the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions, and output the image data after modifying the image data.

[0260]

[0274] Aspect 62: The non-transitory computer-readable medium of aspect 61, further comprising any of aspects 2 to 30 and / or any of aspects 32 to 60.

[0261]

[0275] Aspect 63: An apparatus for image processing, the apparatus comprising: means for receiving image data captured by an image sensor, the image data depicting an environment; means for receiving an indication of an object in the environment represented in the image data; means for dividing the image data into a plurality of regions, the plurality of regions including a first region and a second region, the object being represented in one of the plurality of regions; means for modifying the image data to obscure the first region without obscuring the second region based on the object being represented in one of the plurality of regions; and means for outputting the image data after modifying the image data.

[0262]

[0276] Example 64: The apparatus described in example 63, further comprising means for performing the operations according to any of examples 2 to 30 and / or any of examples 32 to 60.

Claims

1. 1. An apparatus for media processing, said apparatus comprising: a first display; and a second display; and a first camera configured to capture one or more images of a real-world environment; a second camera configured to capture one or more images of the real-world environment; and a third camera configured to capture one or more images of a first eye of a user of the device; a fourth camera configured to capture one or more images of the user's second eye; and a first speaker; a second speaker; at least one inertial measurement unit (IMU); at least one memory; one or more processors coupled to the at least one memory, wherein the one or more processors: generating a first virtual content for display; outputting the first virtual content for display on the first display; receiving image data captured by the first camera, the image data representing the real-world environment; Detecting people within the image data; generating display content in response to detecting the person represented in the image data, the display content including an image of the first virtual content within a second region of the display content and at least a portion of the person within a first region of the display content; outputting the display content for display on the first display; acquiring gaze data associated with the user's gaze; generating second virtual content based on the gaze data; The apparatus is configured to:

2. The device described in claim 1, wherein at least the portion of the person includes the person's face.

3. The device described in claim 1, wherein at least the portion of the person further includes a portion of the person's body.

4. The at least one processor: The device of claim 1 configured to perform at least one hand detection.

5. The at least one processor: generating additional display content in response to detecting the at least one hand within a third region of the display content, the additional display content including an image of the at least one hand; The apparatus of claim 4 configured to:

6. The at least one processor:

5. The device of claim 4, configured to, in response to detecting the at least one hand, determine a gesture associated with the at least one hand, and wherein the at least one processor is configured to track the gaze of the user.

7. The device described in claim 4, wherein the at least one hand is a hand of the person represented in the image data.

8. The device described in claim 1, wherein the first display comprises a digital pass-through display.

9. The at least one processor: generating a virtual keypad; outputting the virtual keypad for display on at least one of the first display or the second display; The apparatus of claim 1 configured to:

10. The at least one processor: The device of claim 1 , configured to detect the person in the image data based on user input, the user input comprising at least one of a touch input or a gesture input.

11. The device described in claim 11, wherein the at least one processor is configured to output additional virtual content for display on the second display, the virtual content and the additional virtual content forming a stereoscopic view.

12. The device described in claim 1, wherein the first region of the displayed content includes the virtual content.

13. The device described in claim 12, wherein the image of at least a portion of the person attenuates the virtual content in at least a portion of the first region.

14. The device described in claim 12, wherein the display content includes a gradient boundary between the first region and the second region.

15. A method for media processing, comprising: generating a first virtual content for display; outputting the first virtual content for display on the first display; receiving image data captured by the first camera, the image data representing the real-world environment; Detecting people within the image data; generating display content in response to detecting the person represented in the image data, the display content including an image of the first virtual content within a second region of the display content and at least a portion of the person within a first region of the display content; outputting the display content for display on the first display; acquiring gaze data associated with the user's gaze; generating second virtual content based on the gaze data; A method comprising: