Foveal magnification
Patent Information
- Application Number
- CN202610385974.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2026-03-13
- Filing Date
- 2026-03-27
- Publication Date
- 2026-09-29
Smart Images

Figure CN122845846A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates in general to systems, methods, and apparatuses that use gaze and head posture to magnify portions of a source image for viewing via electronic devices such as head-mounted displays (HMDs). Background Technology
[0002] Existing technologies that enable users to view content on a device’s high-resolution display can be improved relative to power and computational requirements to provide the desired viewing experience. Summary of the Invention
[0003] The various embodiments disclosed herein include devices, systems, and methods for implementing a process of magnifying only a small portion of a source image based on a region of interest determined based on user gaze, user head pose, camera calibration information, etc. For example, a source image with low resolution or quality can be transmitted and / or decoded for display via a high-resolution display, such as, for example, a 4K resolution display. In some embodiments, the source image may include content frames that have been distorted due to wide-angle lens distortion and / or other distortions imposed to optimize transmission and encoding. In some embodiments, the source image may, in particular, include immersive media, pass-through content, etc.
[0004] In some implementations, a first portion of the source image corresponding to the region of interest can be magnified and combined with the remaining, unmagnified portion of the source image for presentation via a high-resolution display. Combining the first portion of the source image with the remaining portion may include mapping the first portion (i.e., the magnified portion) into the user view space for combination with the remaining portion (i.e., the unmagnified portion) of the source image, which can be additionally mapped into the user view space. This process can help reduce power / computational requirements relative to scenarios associated with, for example, pixel throughput limitations of video decoders, network bandwidth issues, and low-resolution cameras.
[0005] In some implementations, information associated with a user's gaze can be excluded from access by systems outside the electronic device. For example, when a zoom-in process is performed on different devices, the system or software requesting the zoom-in of the aforementioned source image portion (outside the system or device performing the zoom-in) may be unable to access the user's gaze information, thereby protecting user data privacy.
[0006] In some embodiments, an electronic device has a processor (e.g., one or more processors) that executes instructions stored in a non-transitory computer-readable medium to perform a method. The method performs one or more steps or processes. In some embodiments, the electronic device acquires source content at a first resolution to present to a user via one or more displays. In some embodiments, the user's gaze direction and head posture can be identified based on sensor data acquired via one or more sensors, and a region of interest corresponding to a region of the source content can be identified based on the identified user gaze direction and head posture. In some embodiments, a region of the source content can be magnified to a second resolution greater than the first resolution, and combined content can be generated by combining the magnified region with the second resolution with the remaining region of the source content with the first resolution. In some embodiments, a view of the combined content can be presented to a user via one or more displays.
[0007] According to some embodiments, an apparatus includes one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the apparatus, cause the apparatus to perform or cause to perform any of the methods described herein. According to some embodiments, an apparatus includes: one or more processors, non-transitory memory, and components for performing or causing to perform any of the methods described herein. Attached Figure Description
[0008] To enable those skilled in the art to understand this disclosure, more detailed descriptions can be made with reference to aspects of some exemplary embodiments, some of which are shown in the accompanying drawings.
[0009] Figures 1A to 1B Exemplary electronic devices operating in a physical environment according to some specific implementations are illustrated.
[0010] Figure 2 Examples are given of views of distorted images in content warp space and views of rendered images in content view space, based on some specific implementations.
[0011] Figure 3 An example is illustrated of a pipeline configured to dynamically zoom in on portions of an image to optimize viewing via a device such as an HMD, according to some specific implementations.
[0012] Figure 4An example is illustrated of a protected system synthesizer pipeline configured, according to some specific implementations, to dynamically magnify portions of an image for optimized viewing via a device such as an HMD.
[0013] Figure 5 This is a block diagram illustrating an example system of a user device, based on some specific implementations, for magnifying only a small portion of a source image based on a region of interest.
[0014] Figure 6 This is a flowchart representation of an exemplary method for magnifying a small portion of a source image based on a region of interest (ROI) in some specific implementations, the ROI being determined based on user gaze, user head pose, and / or camera calibration information.
[0015] Figure 7 It is a block diagram based on some specific implementations of electronic devices.
[0016] As is customary practice, various features illustrated in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Furthermore, some drawings may not depict all components of a given system, method, or apparatus. Finally, throughout the specification and drawings, the same reference numerals may be used to denote the same features. Detailed Implementation
[0017] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects and / or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.
[0018] Figures 1A to 1B Exemplary electronic devices 105 and 110 are illustrated in physical environment 100. Figures 1A to 1B In the example, physical environment 100 is a room including table 120. Electronic devices 105 and 110 may include one or more cameras, microphones, depth sensors, or other sensors that can be used to capture and evaluate information about physical environment 100 and objects within it, as well as information about user 102 of electronic devices 105 and 110. Information about physical environment 100 and / or user 102 may be used to provide visual and audio content, and / or to identify the current location of physical environment 100 and / or the location of user 102 within physical environment 100.
[0019] In some implementations, a view of the extended reality (XR) environment may be provided to one or more participants (e.g., user 102 and / or other participants not shown) via electronic devices 105 (e.g., wearable devices such as HMDs) and / or 110 (e.g., handheld devices such as mobile devices, tablet computing devices, laptops, etc.). This XR environment may include a view of the 3D environment generated from camera images and / or depth camera images of the physical environment 100, and a representation of user 102 based on camera images and / or depth camera images of user 102. Such an XR environment may include virtual content positioned at a 3D location relative to a 3D coordinate system associated with the XR environment, which may correspond to the 3D coordinate system of the physical environment 100.
[0020] In some implementations, electronic devices (such as device 105, such as an HMD) can be configured to provide head posture and gaze-related video enhancement by intelligently identifying areas of the image that the user is focusing on (e.g., the area the user is viewing) and applying targeted magnification only to that area. This targeted magnification process can reduce power and computational requirements by eliminating the need to magnify the entire image or frame of the video.
[0021] In some implementations, the electronic device may be configured to acquire source content at a first (low) resolution for presentation to a user via the device's display. For example, the source content may include content frames or images that have been distorted due to wide-angle lens distortion and / or other types of distortion imposed to optimize the transmission and encoding process.
[0022] In some implementations, the user's gaze direction and head posture can be identified (e.g., via sensor data), and a region of interest (ROI) corresponding to a region of the source content (e.g., the region may include an object or part of an object of the source content) can be predicted based on the identified gaze direction and head posture. For example, determining where the user is looking relative to one or more previous frames of the source content can be used to determine the ROI and potential zoom strategies to be used in subsequent frames. In some implementations, one or more previous frames may include blank frames (e.g., used to determine the user's gaze direction and head posture) that appear before the display of initiating content (e.g., immersive or pass-through content).
[0023] In some specific implementations, the region of the source content associated with the predicted region of interest can be magnified to a second (high) resolution (e.g., 4K) greater than the low resolution of the source content.
[0024] In some implementations, combined content can be generated by combining a magnified area with a second resolution with the remaining area of the source content with a first resolution. A view of the combined content can then be presented to the user via the display of an electronic device.
[0025] Figure 2 Examples include a view 200a of a distorted image 205a in a content distortion space 215 and a view 200b of a rendered image 205b in a content view space 220, according to some specific implementations.
[0026] In some implementations, distorted image 205a represents an image that has been distorted due to wide-angle lens distortion and / or alternative distortion types applied to an image or frame to optimize transmission encoding (e.g., source content in content distortion space 215, such as immersive media, pass-through content, etc.). Content distortion space includes a coordinate space for creating the distortion or pre-distortion of the content, taking into account specified factors such as, for example, lens distortion (e.g., distortion introduced by an HMD lens). In some implementations, distorted image 205a is generated in response to receiving media content (e.g., an image or frame) and decoding that media content using, for example, a hardware-accelerated decoder. The resulting decoded frame includes distorted images captured using a fisheye lens or wide-angle camera to capture a large FOV. Subsequently, each distorted image (e.g., distorted image 205a) can be projected onto a grid using camera calibration data (e.g., camera orientation relative to the scene) to correct the distorted image.
[0027] In some implementations, the rendered image 205b represents an image that has been rendered (without distortion) into 3D space and projected onto a (user) view space 220 for each eye based on the user's gaze and head pose. The view space 220 includes spaces for defining the scene for positioning and orientation relative to the camera.
[0028] In some implementations, a magnification process is applied to the distorted image 205a to generate a rendered image 205b comprising a high-resolution portion 209 (e.g., 4K×4K) and a low-resolution portion 217 (e.g., the remainder of a scene at a lower resolution, such as 2K×2K). In some implementations, the magnification process may include magnifying only a portion 207 of the distorted image 205a based on a region of interest (i.e., within portion 207), which is determined based on user gaze, user head pose, camera calibration information, etc. Subsequently, the (magnified) portion 207 (e.g., corresponding to the region of interest) of the distorted image 205a may be mapped to the user view space 220 for combination with the unmagnified portion 212 (e.g., the remainder of the distorted image 205a) to create combined content mapped into view space 220, generating a rendered image 205b comprising the high-resolution portion 209 and the low-resolution portion 217 for user presentation via a device such as an HMD.
[0029] The aforementioned amplification process can help reduce power / computational requirements relative to scenarios associated with, for example, pixel throughput limitations of video decoders, network bandwidth issues, and low-resolution cameras.
[0030] Figure 3 An example is illustrated of a pipeline 300 configured, according to some specific implementation, to dynamically magnify portions of an image for optimized viewing via a device such as an HMD. For instance, pipeline 300 implements a process for dynamically magnifying only the region or area of the image predicted to be viewed by the user (e.g., the area or area the user is currently viewing) by sampling from multiple textures (e.g., low-resolution and high-resolution textures) and intelligently blending these textures based on head pose, camera calibration information, and user gaze localization (e.g., based on determining that eye localization satisfies specified localization criteria to determine that a person is looking at a specified portion of the image).
[0031] In some implementations, pipeline 300 optimizes rendering for a foveal display system by integrating gaze, head pose, and camera calibration information to dynamically magnify regions of interest within low-resolution (e.g., 2K×2K) source content (e.g., one or more images). Therefore, pipeline 300 is configured to selectively apply magnification to improve visual quality relative to the user's focal region (of the image) while reducing computational overhead in peripheral regions of the image. Similarly, pipeline 300 is configured to use hybrid techniques to achieve a seamless transition between the magnified resolution region and the lower (basic) resolution region.
[0032] In some implementations, pipeline 300 acquires source content 308 (e.g., images or frames of immersive media, pass-through content, etc.) at a basic resolution (e.g., low resolution, such as 2K×2K) for transmission via transmission module 320 and decoding via video decoder 322.
[0033] In some specific implementations, calibration data 306 is obtained to adjust pipeline 300 relative to factors such as gaze accuracy, texture blending, and environmental adjustment.
[0034] In some implementations, the user's head pose 302 can be tracked in real time to ensure that the mesh moves dynamically with the user's head, thereby maintaining content alignment with the user. For example, the user's head pose 302 can be configured to ensure that the undistorted image remains aligned with the user's gaze direction vector 304 (e.g., a 3D vector originating from the center of the eye or camera positioning) and head movement. The user's gaze direction vector 304 can be used to determine the focal region (e.g., region of interest) associated with the user's attention. For example, the gaze direction vector 304 and the head pose 302 (e.g., a head pose matrix) can be combined to compute gaze localization in 3D space. Subsequently, the gaze localization can be mapped to screen coordinates to apply magnification to the user's region of interest. In some implementations, the pipeline 300 can be configured to maintain privacy by ensuring that external (third-party) applications do not have direct access to gaze information during the magnification process.
[0035] In some specific implementations, the view-to-content warp space module 310 can be configured to warp a portion of the source content 308 (e.g., as...). Figure 2 The portion 207 shown in the illustration is distorted for magnification. The portion of the source content 308 identified for magnification can be associated with the region of interest determined via region of interest calculation 312 based on the gaze direction vector 304 and head pose 304 (and calibration data 306).
[0036] In some specific implementations, the portion of the source content 308 that is identified for magnification (i.e., the region of interest) is input into the smart amplifier 314 to increase the pixel count by using, for example, rule-based algorithms or artificial intelligence (AI) to recover or generate details that are not present in the original source content 308 (e.g., low-resolution content), thereby enhancing the resolution and quality of that portion of the source image (e.g., magnification).
[0037] In some implementations, the smart amplifier 314 may include a hardware amplifier, which may include, for example, multi-scale retinex (MSR) hardware to enhance the contrast and dynamic range of the portion of the source image associated with the region of interest.
[0038] In some implementations, the smart amplifier 314 may include a scaling and enhancement engine (ASE) configured to enhance image quality in the system-on-chip (SoC). In some implementations, the ASE may be configured to analyze the image portion by calculating the first derivative of the image portion to detect edges (of the image portion) and perform magnification relative to the detected edges and associated directions, thereby enabling new pixels to be aligned with the detected edges to preserve the appearance of sharpness and detail in the image.
[0039] In some implementations, the smart amplifier 314 may include an application programming interface (API) with direct access to the GPU for graphics optimization. Similarly, the API can be configured to handle tasks related to image resizing, resampling, and rendering.
[0040] In some specific implementations, the smart amplifier 314 may use machine learning, deterministic rule-based algorithms, and / or computer vision algorithms to enhance the resolution (magnification) of a portion of the region of interest of an image relative to a scene associated with bandwidth limitations, low-resolution content, or a device with low rendering capabilities.
[0041] In some implementations, the rendering module 324 can be configured as a texture sampling mechanism that obtains magnified content 318 (associated with the region of interest of the image) from the smart amplifier 314 and the remaining area of the source content 308 from the video decoder 322. The rendering module 324 is also configured to generate combined content for rendering via, for example, an HMD, by combining the magnified content 318 with the remaining area of the source content 308.
[0042] In some implementations, the rendering module 324 may include a blending mechanism for selecting and blending low-resolution textures (outside the region of interest) and high-resolution textures (associated with the region of interest) based on where each shaded pixel is located relative to the user's gaze area.
[0043] Figure 4 An example is illustrated of a synthesizer pipeline 400 configured, according to some specific implementation, to dynamically magnify portions of an image for optimized viewing via a device such as an HMD. Pipeline 400 is implemented to maintain gaze privacy and enhance external (third-party) applications without exposing sensitive user data, such as, for example, user gaze patterns.
[0044] In some implementations, pipeline 400 executes an external (third-party) application process 405 configured to track user head pose 402 to adjust the camera viewpoint for each eye buffer. Similarly, interpupillary distance (IPD) 504 can be determined to ensure proper stereo separation between non-foveal eye buffers 406, thereby adjusting camera settings to ensure accurate visual perception of depth.
[0045] In some implementations, pipeline 400 performs a synthesizer process 409 configured to reproject source content from non-foveal eye buffers 406 (e.g., left and right) at a basic (low) resolution. Subsequently, information associated with an updated head pose 410 (e.g., latest head pose) is applied to non-foveal eye buffers 406 to adjust head movement and realign the content with the user's current viewpoint. In response, a region of interest (ROI) calculation 412 is determined based on a gaze direction vector 414 (not detectable via external application process 405) and the updated head pose 410. The ROI calculation 412 defines a region of interest (e.g., a region centered on the user's fovea or gaze point) that identifies a portion of the source content for magnification to achieve better sharpness and visual fidelity.
[0046] In some implementations, the portion of the source content that is identified for magnification (e.g., a designated object) (i.e., a region of interest) is input into a smart amplifier 418 to increase the pixel count by using, for example, deterministic rule-based algorithms or artificial intelligence (AI) to recover or generate details that are not present in the original source content (e.g., low-resolution content), thereby enhancing the resolution and quality of that portion of the source image (e.g., magnification) (to create magnified content 420).
[0047] In some specific implementations, the smart amplifier 418 may include, in particular, hardware amplifiers (e.g., MSR), ASE, API, machine learning modules, deterministic rule-based algorithms and / or computer vision algorithms, as described above relative to... Figure 3 The intelligent amplifier 314 is described.
[0048] In some implementations, the magnified content 420 (associated with the region of interest of the image) from the smart amplifier 418 is blended with the remaining portion (e.g., the lower-resolution portion) of the reprojected source content via the compositing module 424. This blending process achieves a smooth transition between the high-resolution and lower-resolution portions of the source content (e.g., the image) for rendering via, for example, an external third-party application. This ensures that the external third-party application receives an accurate, high-quality eye buffer for rendering, while maintaining head pose and IPD considerations and excluding user gaze information from access by the external third-party application.
[0049] Figure 5 This is a block diagram illustrating an example system of a user device 500 for magnifying only a small portion of a source image based on a region of interest (ROI), determined based on user gaze, user head pose, and / or camera calibration information. The user device 500 (e.g., an HMD) may include a non-AI-based interface 502, a hardware-based process 504, a non-AI-based interface 506, and a lightweight engine 508. The hardware-based process 504 may include, or alternatively may be, a deterministic process, a deterministic rule-based process, or a computer vision-based process. Similarly, the hardware-based process 504 may be executed via hardware amplifiers such as multi-scale Retinex hardware, scaling and enhancement engines, APIs, machine learning modules, etc.
[0050] In some implementations, a non-AI-based (e.g., hardware-based) interface 502 is configured to accept input data (e.g., from sensors, databases, etc.) including source content (e.g., images, video frames, etc.) to perform a process (e.g., hardware-based, deterministic rule-based, computer vision-based, etc.) for dynamically magnifying only the region or area of the image or frame predicted to be viewed by the user (e.g., the area or area the user is currently viewing). This process may include sampling from multiple textures (e.g., low-resolution and high-resolution textures) and intelligently blending textures based on user gaze localization, head pose, and camera calibration information.
[0051] Input 510 from a non-AI-based interface 502 can be fed to a hardware-based process 504. The hardware-based process 504 may include one or more learning-based and / or non-learning-based models for sensing, synthesizing, and inferring information. Those skilled in the art will understand that the hardware-based process 504 may include any suitable number of hardware-based processes to generate output 512 based on input 510.
[0052] In some implementations, the hardware-based process 504 can be used in conjunction with, for example, computer vision and / or deterministic rule-based techniques that can be used to magnify a portion of the content frame / image (e.g., corresponding to a region of interest) and combine it with the unmagnified portion of the content frame / image to help reduce power / computational requirements relative to scenarios associated with, for example, pixel throughput limitations of video decoders, network bandwidth issues, low-resolution cameras, etc.
[0053] In some implementations, user device 500 may optionally choose to use lightweight engine 508 instead of hardware-based process 504 to amplify a portion of the content frame / image (e.g., corresponding to the region of interest) and combine it with the unamplified portion of the content frame / image. Lightweight engine 408 may be a non-learning network.
[0054] In some implementations, the non-AI-based interface 406 is configured to accept results from the hardware-based process 504 and / or the lightweight engine 508 to provide a combined content consisting of a magnified portion of the content (e.g., a high-resolution portion of an image) and the remainder of the source content (e.g., a high-resolution portion of an image), thereby providing an enhanced view of the content.
[0055] For example, the output from hardware-based process 504 and / or lightweight engine 508 is provided to non-AI-based interface 506, where interface 506 provides an enhanced view of the content.
[0056] Those skilled in the art will understand that the hardware-based process 504 may include any suitable machine learning model that is well-known or widely available, such as regression techniques, classification techniques, neural networks, and deep learning networks. For example, the hardware-based process 504 may include neural networks such as artificial neural networks (ANNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), adversarial networks (GANs), reinforcement learning models (RLMs), encoder / decoder networks, and / or Transformer-based models (e.g., bidirectional encoder representations from Transformers (BERT), generative pre-trained Transformers (GPT), and / or multimodal large language models (LLMs)). Additionally or alternatively, those skilled in the art will understand that the hardware-based process may be any suitable non-learning process, such as rule-based systems, computer vision-based systems, heuristic methods, decision trees, knowledge-based systems, statistical or stochastic systems, and expert systems.
[0057] In cases where the hardware-based process 504 includes a machine learning-based model, the hardware-based process can be trained to analyze the original image relative to detected regions of interest determined based on user gaze, user head pose, camera calibration information, etc., to magnify a portion of the original image (e.g., an object corresponding to the region of interest), and combine the magnified portion of the original image with the unmagnified portion to generate an augmented view of the content using one or more well-known or widely available training techniques (such as supervised learning, semi-supervised learning, unsupervised learning, and / or reinforcement learning techniques). Training data may include images, gaze data, head pose data, camera calibration information, etc.
[0058] In some implementations, where the hardware-based process 504 includes a machine learning-based model, the machine learning-based model can be trained on a large dataset that includes gaze data, head pose data, and camera calibration information.
[0059] In some implementations, both synthetic and non-synthetic data can be used to train the hardware-based process 504.
[0060] In some implementations, the hardware-based process 504 may be deployed as one or more generative models, wherein content is automatically generated by one or more computers in response to a request to generate content. The automatically generated content may optionally be generated on the device (e.g., at least in part by the computer system that received the request to generate content) and / or generated outside the device (e.g., at least in part by one or more nearby computers available via a local network or one or more computers available via the Internet). This automatically generated content may optionally include visual content (e.g., images, graphics, and / or video), audio content, and / or text content.
[0061] In some implementations, novel, automatically generated content is referred to as generative content (e.g., generative images, generative graphics, generative videos, generative audio, and / or generative text). Generative content is typically generated based on prompt input 510 for hardware-based process 504. Non-AI-based interface 502 may optionally include one or more preprocessing steps to adjust the input (e.g., adjusting user-provided prompts, creating system-generated prompts, and / or selecting AI models) before the AI model uses the input to generate an output image. Non-AI-based interface 506 may optionally include one or more postprocessing steps to adjust the output 512 of hardware-based process 504 (e.g., passing AI model output to different AI models, scaling up, scaling down, cropping, formatting, and / or adding or removing metadata) before the output 512 of hardware-based process 504 is used for other purposes (e.g., being provided to different software processes for further processing) or presented to the user (e.g., visually or auditorily).
[0062] The cue input 510 for generating generative content may include one or more of the following: one or more words (e.g., written or spoken natural language cues), one or more images and / or magnified portions of images, one or more pictures and / or one or more videos with magnified portions of frames. A generative pre-trained transformer model is a type of LLM that can efficiently generate novel generative content based on the cue input 510. In some embodiments, the hardware-based process 504 uses the cue input 510, which includes text, to generate various generative texts, generative audio content, and / or generative visual content (such as magnified visual content). In other embodiments, the hardware-based process 504 uses the cue input 510, which includes visual content and / or audio content, to generate generative text (e.g., transcriptions of audio and / or descriptions of visual content). In still other embodiments, the hardware-based process 504 uses the cue input 510, which includes multiple types of content (e.g., text, images, audio, video, and / or other sensor data), to generate generative content. The cue input 510 may sometimes also include values for one or more parameters indicating the importance of various parts of the cue. Some input prompts 510 include a structured set of instructions that can be provided to the hardware-based process 504, which includes wording, specified style, relevant context (e.g., starting content and / or one or more examples) and / or roles for the hardware-based process 504.
[0063] Generative content is typically cue-based, but not deterministically selected from pre-generated content; rather, it is generated using the cue as a starting point. In some implementations, pre-existing content (e.g., visual content) is used as part of the cue for creating the generative content (e.g., pre-existing content is used as a starting point for creating the generative content). For example, cue input 510 may request modifications to the visual content to include, exclude, or amplify content specified by the cue (e.g., removing identifying features from the visual content, adding features to the visual content to change its visual style, and / or amplifying a portion of the visual content described in the cue, etc.). In some implementations, a random or pseudo-random seed is used as part of the cue input 510 for creating the generative content (e.g., random or pseudo-random seed content is used as a starting point for creating the generative content). For example, when a portion of an image is amplified via a diffusion model, random noise patterns are iteratively denoised based on the cue input 510 to amplify the image portion. While a particular type of hardware-based process 504 has been described herein, it should be understood that a variety of different hardware-based processes can be used to generate generative content based on cue.
[0064] In cases where the hardware-based process 504 is a non-learning-based system, the hardware-based process 504 can use a predefined set of rules or a predefined structure to make decisions based on the input seen by the process. For example, the hardware-based process 504 can be used to optimize rendering for a foveal display system by integrating gaze, head pose, and camera calibration information to dynamically determine and selectively magnify the user's focal region (region of interest) within low-resolution source content (e.g., one or more images), while reducing computational overhead in the peripheral regions of the source content. Similarly, the hardware-based process 504 can be used to achieve a seamless transition between magnified resolution regions and lower (basic) resolution regions using hybrid techniques.
[0065] Similarly, hardware-based process 504 can be used to enable extrapolation techniques to selectively magnify portions of an image.
[0066] Some implementations described herein may include the use of learning-based and / or non-learning-based processes. This use may include collecting, preprocessing, encoding, labeling, organizing, analyzing, recommending, and / or generating data. Entities that collect, share, and / or otherwise utilize user data should provide transparency and / or obtain user consent when collecting such data. This disclosure recognizes that the use of data in hardware-based processes can be used to benefit users. For example, the data can be used to train models that can be deployed to improve the performance, accuracy, and / or functionality of applications and / or services. Therefore, the use of data enables hardware-based processes to adapt and / or optimize operations to provide a more personalized, efficient, and / or enhanced user experience. Such adaptation and / or optimization may include customizing content, recommendations, and / or interactions for individual users, as well as simplifying processes and / or implementing more intuitive interfaces. This disclosure also envisions further beneficial uses of data in hardware-based processes.
[0067] This disclosure envisions that, in some embodiments, data used by hardware-based processes may be used to include publicly available data. To protect user privacy, data may be anonymized, aggregated, and / or otherwise processed to remove or, where possible, limit any individual identifiers (e.g., restricting the sharing of gaze data). As discussed herein, entities that collect, share, and / or otherwise utilize such data should obtain user consent before collecting such data and / or provide transparency in the collection process. Furthermore, this disclosure envisions that entities responsible for the use of data (including, but not limited to, data used in connection with hardware-based processes) should strive to comply with robust privacy policies and / or privacy measures.
[0068] For example, such entities may implement and consistently follow strategies and measures deemed to meet or exceed industry standards and regulatory requirements for developing and / or training hardware-based processes. In doing so, efforts should be made to ensure that all intellectual property and privacy considerations are maintained. Training should include measures to protect training data (such as personal information) by providing adequate safeguards against misuse or exploitation. Such strategies and measures should cover all phases of the rules-based process development, training, and use, including data collection, data preparation, model training, model evaluation, model deployment, and ongoing monitoring and maintenance. Transparency and measurability should always be maintained. Such policies should be readily accessible to users and should be updated as data collection and / or use change. User data should be collected for the entity's lawful and reasonable use and not shared or sold outside of these lawful uses. Furthermore, such collection and sharing should be conducted in a transparent manner to users and / or with their informed consent. Additionally, such entities should consider taking any necessary steps to defend and safeguard access to such data and ensure that others with access to the data comply with their privacy policies and processes. Furthermore, such entities should be subject to appropriate third-party assessments for transparency purposes to demonstrate their compliance with widely accepted privacy policies and practices. Additionally, policies and / or practices should be appropriate for the specific types of data collected and / or accessed (e.g., user gaze data) and tailored to specific use cases and applicable laws and standards, including jurisdiction-specific considerations.
[0069] In some implementations, the hardware-based process can utilize a model that can be trained (e.g., in a supervised or unsupervised learning manner) using a variety of training data, including data collected using the user device. This use of user-collected data (e.g., user gaze data) can be limited to operations on the user device. For example, model training can be performed locally on the user device, so that no part of the data is transmitted to another device. In other implementations, model training can be performed using one or more other devices besides the user device (e.g., a server), but in a privacy-preserving manner, such as through multi-party computation, encrypted by secretly sharing data, or other means, ensuring that user data is not leaked to these other devices.
[0070] In some implementations, the trained model can be centrally stored on the user device or on multiple devices, as is the case in federated learning. This distributed storage can also be done in a privacy-preserving manner, for example, through encryption, where each piece of data is fragmented so that the data cannot be reassembled or used by a device alone (i.e., only in conjunction with another device) or only by the user device. In this way, user or device behavior patterns are not leaked, while leveraging the increased computing resources of other devices to train and execute the ML model. Therefore, user-collected data can be protected. In some specific implementations, data from multiple devices can be combined in a privacy-preserving manner to train the ML model.
[0071] In some implementations, this disclosure envisions that data used for hardware-based processes can be kept strictly separate from the platform in which the hardware-based processes are deployed and / or used to interact with users and / or process data. In such implementations, data used for offline training of the hardware-based processes can be maintained in a secure data repository with restricted access and / or not retained for longer than necessary for the training objectives. In some implementations, the hardware-based processes can utilize local memory caching to temporarily store data during a user session. Local memory caching can be used to improve the performance of the hardware-based processes. However, to protect user privacy, data stored in the local memory cache can be erased after the user session ends. Any temporary cached data used for online learning or inference can be quickly erased after processing. All data collection, transmission, and / or storage should use industry-standard encryption and / or secure communication.
[0072] In some implementations, as noted above, techniques such as federated learning, differential privacy, secure hardware components, homomorphic encryption and / or multi-party computation, and others, can be used to further protect personal information data during the training and / or use of hardware-based processes. Hardware-based processes should be monitored for changes in the underlying data distribution, such as concept drift or data skew that may degrade the performance of hardware-based processes over time.
[0073] In some implementations, a combination of offline and online training is used to train the hardware-based process. Offline training can use a carefully selected dataset to establish baseline model performance, while online training allows the hardware-based process to continuously adapt and / or improve. This disclosure recognizes the importance of maintaining strict data management measures throughout this process to ensure user privacy is protected.
[0074] In some implementations, hardware-based processes can be designed with safeguards to maintain adherence to the original intended purpose even when rule-based processes are adapted based on new data. Any significant changes to data collection and / or application used in hardware-based processes can (and in some cases should) be transparently communicated to affected stakeholders and / or include obtaining user consent regarding changes to the manner in which user data is collected and / or utilized.
[0075] Regardless of the foregoing, this disclosure also contemplates implementation schemes for users to selectively restrict and / or block the use and / or access to data. That is, this disclosure contemplates that hardware and / or software elements may be provided to prevent or block access to data. For example, with respect to some services, the inventive technology should be configured to allow users to opt-in or opt-out to participate in data collection at any time during or after service registration. In another example, the inventive technology should be configured to allow users to opt out of providing certain data used for training rule-based processes and / or as input during the inference phase of such systems. In yet another example, the inventive technology should be configured to allow users to choose to limit the length of time their data is maintained or to completely prohibit their data from being used by rule-based processes. In addition to providing opt-in and opt-out options, this disclosure also contemplates providing notifications related to access to or use of personal information. For example, users may be notified when their data is input into a hardware-based process for training or inference purposes, and / or alerted when a hardware-based process generates output or makes a decision based on the user's data.
[0076] This disclosure recognizes that hardware-based processes should be combined with explicit limitations and / or oversight to mitigate risks that may exist even when such systems are designed, developed, and / or operated in accordance with industry best practices and standards. For example, outputs may be generated that could be considered erroneous, harmful, offensive, and / or biased; such outputs may not necessarily reflect the opinions or positions of the entity that developed or deployed these systems. Furthermore, in some cases, references to or non-references to third-party products and / or services in these outputs should not be construed as an endorsement or association by the entity providing the hardware-based process with the third-party products and / or services. The generated content may be filtered to remove potentially inappropriate or dangerous material before being presented to users, while maintaining the ability for human oversight and / or to cover or correct erroneous or undesirable outputs as a safeguard.
[0077] This disclosure further envisions that users of hardware-based processes should avoid using the service in any way that infringes upon, misappropriates, or violates the rights of any party. Furthermore, hardware-based processes should not be used for any illegal or unlawful activities, nor should they be used to develop any applications or use cases that will commit or assist in the committing of crimes or other torts, illegal or unlawful acts, including misinformation, false information, false statements (such as deepfakes), deception, impersonation, and propaganda. Hardware-based processes should not violate, misappropriate, or infringe upon any party's copyright, trademark, privacy and portrait rights, trade secrets, patents, or other proprietary or statutory rights, and should be appropriately attributed to the content as needed. Furthermore, hardware-based processes should not interfere with any security, digital signature, digital rights management, content protection, verification, or authentication mechanisms. Hardware-based processes should not misrepresent machine-generated output as human-generated output.
[0078] Figure 6 This is a flowchart representation of only a small portion of an exemplary method 600 for magnifying a source image based on a region of interest (ROI) according to some specific implementations, the ROI being determined based on user gaze, user head pose, and / or camera calibration information. In some implementations, method 600 is performed by a device, such as a mobile device, desktop computer, laptop computer, HMD, or server device. In some implementations, the device has a screen for displaying images and / or a screen for viewing stereoscopic images, such as a head-mounted display (HMD, such as device 105 of Figure 1, for example). In some implementations, method 600 is performed by processing logic components, including hardware, firmware, software, or combinations thereof. In some implementations, method 600 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). Each block in method 600 can be implemented and executed in any order.
[0079] At box 602, method 600 obtains source content at a first resolution to be presented to the user via a display of a device such as an HMD. For example, the source image can be obtained from source content 308, as relative to... Figure 3 As described.
[0080] In some implementations, the source content may include an image distorted due to wide-angle lens distortion imposed to optimize transmission and encoding. For example, the source content may include images as relative to... Figure 2 The content described is a distorted image 205a in distorted space 215.
[0081] In some specific implementations, the source content includes immersive media.
[0082] In some specific implementations, the source content includes transparent content.
[0083] At box 604, method 600 identifies the user's gaze direction and head pose based on sensor data obtained via one or more sensors. For example, gaze direction vector 304 and head pose 302 can be tracked in real time to keep content aligned with the user, such as relative to... Figure 3 As described.
[0084] In some implementations, information associated with a user's gaze direction can be excluded from access by systems outside the electronic device. For example, pipeline 400 can be configured to protect user privacy by ensuring that external (third-party) applications do not have direct access to gaze information during the magnification process, as relative to... Figure 4 As described.
[0085] At box 606, method 600 predicts a region of interest corresponding to the region of the source content based on the identified user gaze direction and the identified user head pose. For example, as relative to Figure 2 As described, a portion 207 of the distorted image 205a can be identified as the region of interest.
[0086] In some implementations, the user's gaze direction and head pose can be used to identify the location of the user's view relative to at least one previous frame of the source content, and the region of interest can be predicted for the next frame of the source content presented after at least one previous frame.
[0087] In some implementations, at least one previous frame may include content.
[0088] In some implementations, at least one preceding frame may be a blank frame that has no content.
[0089] In some specific implementations, the predicted region of interest can be further based on calibration information associated with one or more sensors (e.g., such as...). Figure 3 The calibration data 306 is illustrated in the figure.
[0090] In some specific implementations, the prediction of regions of interest can be further based on the context of the source content.
[0091] In some specific implementations, the region of the source content may include at least a portion of at least one object.
[0092] At box 608, method 600 enlarges the region of the source content to a second resolution greater than the first resolution, as described relative to Figure 1.
[0093] At box 610, method 600 generates combined content by combining a magnified region with a second resolution with the remaining region of the source content having a first resolution. For example, the (magnified) portion 207 of the distorted image 205a (e.g., corresponding to the region of interest) can be mapped into user view space 220 for combination with the unmagnified portion 212 (e.g., the remaining portion of the distorted image 205a), thereby creating combined content mapped into view space 220 to generate a rendered image 205b including a high-resolution portion 209 and a low-resolution portion 217, as relative to... Figure 2 As described.
[0094] In some implementations, the magnified and remaining areas of the source content are mapped separately into the user view space to generate combined content.
[0095] In some specific implementations, combining the magnified area with the remaining area of the source content can include the area between the magnified area and the remaining area, such as relative to... Figure 3 As described.
[0096] At box 612, method 600 presents a view of the combined content to the user via a display.
[0097] Figure 7 This is a block diagram of example device 700. Device 700 illustrates an exemplary device configuration for electronic devices 105 and 110 of FIG1. Although certain specific features are illustrated, those skilled in the art will understand from this disclosure that various other features are not illustrated for the sake of brevity and to avoid obscuring more relevant aspects of the specific embodiments disclosed herein. Therefore, as a non-limiting example, in some specific implementations, device 700 includes one or more processing units 702 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 706, one or more communication interfaces 708 (e.g., USB, FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.14x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE, SPI, I2C and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 710, output devices (e.g., one or more displays) 712, one or more internal and / or external image sensor systems 714, memory 720, and one or more communication buses 704 for interconnecting these components and various other components.
[0098] In some embodiments, one or more communication buses 704 include circuitry that interconnects system components and controls communication between system components. In some embodiments, one or more I / O devices and sensors 706 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, one or more depth sensors (e.g., structured light, time-of-flight, etc.), one or more cameras (e.g., an inward-facing camera and an outward-facing camera of an HMD), one or more infrared sensors, one or more thermal sensors, etc.
[0099] In some embodiments, one or more displays 712 are configured to present a view of a physical environment, a graphical environment, an extended reality environment, etc., to a user. In some embodiments, one or more displays 712 are configured to present content to a user (determined based on the user's determined user / object location within the physical environment). In some embodiments, one or more displays 712 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emission display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 712 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. In one example, device 700 includes a single display. In another example, device 700 includes displays for each of the user's eyes.
[0100] In some embodiments, one or more image sensor systems 714 are configured to acquire image data corresponding to at least a portion of the physical environment 100. For example, one or more image sensor systems 714 include one or more RGB cameras, monochrome cameras, IR cameras, depth cameras, event-based cameras, etc. (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor). In various embodiments, one or more image sensor systems 714 also include an illumination source emitting light, such as a flash. In various embodiments, one or more image sensor systems 714 also include an on-camera image signal processor (ISP) configured to perform multiple processing operations on the image data.
[0101] In some embodiments, sensor data may be acquired by devices (e.g., devices 105 and 110 of Figure 1) during a scan of a room in a physical environment. The sensor data may include a 3D point cloud and a sequence of 2D images corresponding to views of the room captured during the scan. In some embodiments, the sensor data includes image data (e.g., from an RGB camera), depth data (e.g., depth images from a depth camera), ambient light sensor data (e.g., from an ambient light sensor), and / or motion data from one or more motion sensors (e.g., accelerometers, gyroscopes, IMUs, etc.). In some embodiments, the sensor data includes visual inertial odometry (VIO) data determined based on the image data. The 3D point cloud can provide semantic information about one or more elements of the room. The 3D point cloud can provide information about the location and appearance of surface portions within the physical environment. In some embodiments, the 3D point cloud is acquired over time (e.g., during a scan of the room) and can be updated, with updated versions of the 3D point cloud obtained over time. For example, when the 3D representation is updated / adjusted over time (e.g., when a user scans a room), the 3D representation can be obtained (and analyzed / processed).
[0102] In some embodiments, sensor data may be positioning information, and some embodiments include VIO (Vehicle Identification and Odometry) to determine equivalent odometry information to estimate travel distance using sequential camera images (e.g., light intensity image data) and motion data (e.g., acquired from an IMU / motion sensor). Alternatively, some embodiments of this disclosure may include a Simultaneous Localization and Mapping (SLAM) system (e.g., a positioning sensor). This SLAM system may include a GPS-independent, multi-dimensional (e.g., 3D) laser scanning and range measurement system that provides real-time simultaneous localization and mapping. This SLAM system can generate and manage highly accurate point cloud data produced by reflections from laser scans of objects in the environment. Accurately tracking the movement of any points in the point cloud over time allows the SLAM system to use points in the point cloud as reference points for its location, maintaining an accurate understanding of its position and orientation as it travels through the environment.
[0103] In some embodiments, device 700 includes an eye-tracking system for detecting eye positioning and eye movement (e.g., eye gaze detection). For example, the eye-tracking system may include one or more infrared (IR) light-emitting diodes (LEDs), an eye-tracking camera (e.g., a near-infrared (NIR) camera), and an illumination source (e.g., an NIR light source) that emits light (e.g., NIR light) towards the user's eyes. Furthermore, the illumination source of device 700 may emit NIR light to illuminate the user's eyes, and the NIR camera may capture images of the user's eyes. In some embodiments, the images captured by the eye-tracking system may be analyzed to detect the user's eye positioning and movement, or to detect other information about the eyes such as pupil dilation or pupil diameter. Furthermore, the gaze point estimated from the eye-tracking images may enable gaze-based interaction with content displayed on a near-eye display of device 700.
[0104] Memory 720 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 720 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 720 may optionally include one or more storage devices remotely located to one or more processing units 702. Memory 720 includes a non-transitory computer-readable storage medium.
[0105] In some embodiments, memory 720 or a non-transitory computer-readable storage medium of memory 720 stores an optional operating system 730 and one or more instruction sets 740. Operating system 730 includes procedures for handling various basic system services and for performing hardware-related tasks. In some embodiments, instruction set 740 includes executable software defined by binary information stored in charge. In some embodiments, instruction set 740 is software executable by one or more processing units 702 to implement one or more of the techniques described herein.
[0106] Instruction set 740 includes expanded content instruction set 742 and combined content instruction set 744. Instruction set 740 can be represented as a single software executable file or multiple software executable files.
[0107] The magnification instruction set 742 is configured with instructions that can be executed by the processor to predict the region of interest corresponding to the region of the source content based on the identified user gaze direction and the identified user head pose, and in response, to magnify the region of the source content to a resolution larger than that of the remaining source content.
[0108] The Combination Content Instruction Set 744 is configured with instructions that can be executed by the processor to generate combined content by combining the enlarged region with the remaining region of the source content with a lower resolution.
[0109] Although instruction set 740 is shown as residing on a single device, it should be understood that in other specific implementations, any combination of elements may reside on separate computing devices. Furthermore, Figure 7 This is intended more as a functional description of various features present in a particular specific implementation than as a structural diagram of the specific implementation described herein. As will be appreciated by those skilled in the art, the items shown individually can be combined, and some items can be separated. The actual number of instruction sets and how features are allocated therein will vary depending on the specific implementation and may depend in part on the specific combination of hardware, software, and / or firmware chosen for that particular implementation.
[0110] Those skilled in the art will understand that well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure more relevant aspects of the specific embodiments of the examples described herein. Furthermore, other effective aspects and / or variations do not include all the details in the specific details described herein. Therefore, several details are described to provide a thorough understanding of the exemplary aspects illustrated in the accompanying drawings. Moreover, the drawings only illustrate some exemplary embodiments of this disclosure and should not be considered limiting.
[0111] While this specification contains numerous specific implementation details, these details should not be construed as limiting the scope of any invention or potentially claimed content, but rather as descriptions of features specific to particular embodiments of a particular invention. Certain features described in the context of different embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, while certain features may be described above as functioning in certain combinations and even initially claimed in this manner, one or more features of a claimed combination may be removed from that combination in some cases, and the claimed combination may involve sub-combinations or variations thereof.
[0112] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in a sequential order or the specific order shown, or requiring all illustrated operations to achieve the desired result. In some situations, multitasking and parallel processing may be advantageous. Furthermore, the partitioning of the various system components in the above embodiments should not be construed as requiring such partitioning in all embodiments, and it should be understood that the described program components and systems may generally be integrated together in a single software product or packaged into multiple software products.
[0113] Therefore, specific embodiments of the subject matter have been described. Other embodiments are also within the scope of the following claims. In some cases, the actions described in the claims can be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequence shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.
[0114] The embodiments of the subject matter and operation described in this specification may be implemented in digital electronic circuits or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents) or in a combination thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, such as one or more modules of computer program instructions encoded on a computer storage medium for execution by or control of the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing device. The computer storage medium may be or be included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in artificially generated propagating signals. The computer storage medium may also be or be included in one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).
[0115] The term "data processing apparatus" encompasses all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, systems-on-a-chip, or many or combinations of the foregoing. The apparatus may include special-purpose logic circuitry (e.g., FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)). In addition to hardware, the apparatus may include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or combinations thereof. The apparatus and execution environment can implement a variety of different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures. Unless otherwise specifically stated, it should be understood that throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," and "identifying" refer to the actions or processes of computing devices, such as one or more computers or similar electronic computing devices, that manipulate or convert data represented as physical electronic or magnetic quantities within the memory, registers, or other information storage devices, transmission devices, or display devices of a computing platform.
[0116] The one or more systems discussed herein are not limited to any particular hardware architecture or configuration. A computing device may include any suitable arrangement of components that provide results conditioned on one or more inputs. Suitable computing devices include computer systems based on multi-purpose microprocessors that access stored software that programs or configures the computing system from a general-purpose computing device to a special-purpose computing device that implements one or more specific implementations of this subject. The teachings contained herein may be implemented in the software used for programming or configuring the computing device using any suitable programming, scripting, or other type of language or combination of languages.
[0117] Specific implementations of the methods disclosed herein can be performed in the operation of such computing devices. The order of the boxes presented in the examples above can be varied; for example, the boxes can be reordered, combined, and / or divided into sub-blocks. Some boxes or processes can be executed in parallel. The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0118] The use of "applies to" or "configured to" in this document implies open and inclusive language, which does not exclude applicability to or configuration for performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, as processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated. The headings, lists, and numbering included herein are for illustrative purposes only and are not intended to be restrictive.
[0119] It will also be understood that while terms such as "first," "second," etc., may be used in this document to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first node may be called a second node, and similarly, a second node may be called a first node, changing the meaning of the description, provided that all occurrences of "first node" are consistently renamed and all occurrences of "second node" are consistently renamed. First nodes and second nodes are both nodes, but they are not the same node.
[0120] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and in the appended claims, the singular forms “a,” “an,” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will be further understood that the term “comprising,” when used in this specification, specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0121] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" can be interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when it is detected that the prerequisite is true" or "in response to detection" that the prerequisite is true, depending on the context.
Claims
1. A method, the method comprising: In electronic devices that have a processor and one or more displays: Source content at a first resolution is obtained and presented to the user via the one or more displays; Identify the user's gaze direction and head posture based on sensor data obtained through one or more sensors; The region of interest corresponding to the region of the source content is predicted based on the identified user gaze direction and the identified user head posture. Enlarge the region of the source content to a second resolution greater than the first resolution; Combined content is generated by combining a magnified region with the second resolution with the remaining region of the source content having the first resolution. as well as A view of the combined content is presented to the user via the one or more displays.
2. The method of claim 1, wherein the magnified region and the remaining region of the source content are individually mapped to the user view space for generating the combined content.
3. The method of claim 1, wherein the source content comprises an image distorted due to wide-angle lens distortion applied to optimize transmission and encoding.
4. The method of claim 1, wherein the user's gaze direction and the user's head pose identify the location of the user's view relative to at least one previous frame of the source content, and wherein the region of interest is predicted for the next frame of the source content presented after the at least one previous frame.
5. The method of claim 4, wherein the at least one previous frame includes content.
6. The method of claim 4, wherein the at least one previous frame is a blank frame without any content.
7. The method of claim 1, wherein the prediction of the region of interest is further based on calibration information associated with the one or more sensors.
8. The method of claim 1, wherein the prediction of the region of interest is further based on the context of the source content.
9. The method of claim 1, wherein combining the magnified region with the remaining region of the source content includes blending the region between the magnified region and the remaining region.
10. The method of claim 1, wherein information associated with the user's gaze direction is excluded from access by systems outside the electronic device.
11. The method of claim 1, wherein the region of the source content comprises at least a portion of at least one object.
12. The method of claim 1, wherein the source content includes immersive media.
13. The method of claim 1, wherein the source content includes transparent content.
14. A non-transitory computer-readable storage medium storing program instructions, the program instructions being executable via one or more processors to perform operations, the operations including: At an electronic device having the one or more processors and one or more displays: Source content at a first resolution is obtained and presented to the user via the one or more displays; Identify the user's gaze direction and head posture based on sensor data obtained through one or more sensors; The region of interest corresponding to the region of the source content is predicted based on the identified user gaze direction and the identified user head posture. Enlarge the region of the source content to a second resolution greater than the first resolution; Combined content is generated by combining a magnified region with the second resolution with the remaining region of the source content having the first resolution. as well as A view of the combined content is presented to the user via the one or more displays.
15. An electronic device, the electronic device comprising: One or more displays; Non-transitory computer-readable storage medium; and One or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium includes program instructions that, when executed on the one or more processors, cause the electronic device to perform operations including: Source content at a first resolution is obtained and presented to the user via the one or more displays; Identify the user's gaze direction and head posture based on sensor data obtained through one or more sensors; The region of interest corresponding to the region of the source content is predicted based on the identified user gaze direction and the identified user head posture. Enlarge the region of the source content to a second resolution greater than the first resolution; Combined content is generated by combining a magnified region having the second resolution with the remaining region of the source content having the first resolution; and A view of the combined content is presented to the user via the one or more displays.
16. The electronic device of claim 15, wherein the magnified region and the remaining region of the source content are individually mapped to a user view space for generating the combined content.
17. The electronic device of claim 15, wherein the source content includes an image distorted due to wide-angle lens distortion applied to optimize transmission and encoding.
18. The electronic device of claim 15, wherein the user's gaze direction and the user's head pose identify the location of the user's view relative to at least one previous frame of the source content, and wherein the region of interest is predicted for the next frame of the source content presented after the at least one previous frame.
19. The electronic device of claim 18, wherein the at least one previous frame includes content.
20. The electronic device of claim 18, wherein the at least one previous frame is a blank frame without any content.