Context-based selection of viewing angle correction operations

By selecting appropriate perspective correction operations based on context in the computing system, the distance perception, orientation and hand-eye coordination problems caused by HMD when presenting an XR environment are solved, and better XR experience and system resource optimization are achieved.

CN119948438APending Publication Date: 2025-05-06APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068140.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-23
Filing Date
2023-09-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When presenting an extended reality (XR) environment using a head mounted device (HMD), impaired distance perception, disorientation, and poor hand-eye coordination are caused due to different positioning of the eyes, display, and camera in space.

Method used

By performing a viewing angle correction operation in the computing system, image data of the physical environment is acquired using an image sensor, and an appropriate viewing angle correction operation set is selected based on the current state of the user, the executed application program and the state of the physical environment, and the corrected image data is generated and presented.

Benefits of technology

Improved distance perception, orientation and hand-eye coordination in the XR experience, reduced the possibility of motor vertigo, and optimized image quality and system resource use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948438A_ABST
    Figure CN119948438A_ABST
Patent Text Reader

Abstract

In some implementations, a method includes: acquiring image data associated with a physical environment; obtaining first contextual information including at least one of first user information associated with a current state of a user of the computing system, first application information associated with a first application executed by the computing system, and first environment information associated with a current state of the physical environment; selecting a first set of view correction operations based at least in part on the first contextual information; generating first corrected image data by performing the first set of view correction operations on the image data; and causing the first corrected image data to be presented.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 409,327, filed on September 23, 2022, which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure relates generally to perspective correction, and in particular to systems, methods, and devices associated with context-based selection of perspective correction operations. Background Art

[0004] In various embodiments, an extended reality (XR) environment is presented by a head mounted device (HMD). Various HMDs include a scene camera that captures an image of the physical environment (e.g., a scene) in which a user is present and a display that displays the image to the user. In some cases, the image or a portion thereof may be combined with one or more virtual objects to present an XR experience to the user. In other cases, the HMD may operate in a pass-through mode in which an image or a portion thereof is presented to the user without adding virtual objects. Ideally, the image of the physical environment presented to the user is substantially similar to what the user would see if the HMD were not present. However, due to the different positioning of the eyes, display, and camera in space, this may not occur, resulting in impaired distance perception, disorientation, and poor hand-eye coordination. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] In order that the present disclosure may be understood by those skilled in the art, a more particular description will be given with reference to some exemplary implementations, some of which are illustrated in the accompanying drawings.

[0006] Figure 1 is a block diagram of an exemplary operating environment according to some implementations.

[0007] Figure 2 Example scenarios related to capturing images of a physical environment and displaying the captured images are shown according to some implementations.

[0008] Figure 3 An image of a physical environment captured by an image sensor from a particular perspective according to some implementations.

[0009] Figure 4 It is implemented according to some specific Figure 3 A top-down perspective view of the physical environment.

[0010] Figure 5A is a block diagram of an exemplary input processing architecture according to some specific implementations.

[0011] Figure 5B According to some specific implementations Figure 5A An exemplary data structure related to the input processing architecture in .

[0012] Fig. 6A is a block diagram of an exemplary content delivery architecture according to some implementations.

[0013] Figure 6B According to some specific implementation Fig. 6A Block diagram of the perspective correction logic components associated with the content delivery architecture in FIG.

[0014] Figure 6C According to some specific implementation Fig. 6A Another block diagram of perspective correction logic components associated with the content delivery architecture in FIG.

[0015] Fig. 7A Various perspective correction operations are shown according to some implementations.

[0016] Figure 7B Point of view (POV) position correction operations are shown according to some implementations.

[0017] Figure 7C Various perspective correction scenarios are shown according to some implementations.

[0018] Fig. 8A and Figure 8B A flowchart representation of a method for context-based selection of perspective correction operations according to some implementations is shown.

[0019] Fig. 9 is a block diagram of an example controller according to some implementations.

[0020] Fig.10 is a block diagram of an exemplary electronic device according to some implementations.

[0021] As is common practice, the various features illustrated in the drawings may not be drawn to scale. Therefore, the sizes of the various features may be arbitrarily expanded or reduced for clarity. In addition, some of the drawings may not depict all of the components of a given system, method, or device. Finally, throughout the specification and drawings, similar reference numerals may be used to represent similar features. Summary of the invention

[0022] Various embodiments disclosed herein include devices, systems, and methods for context-based selection of perspective correction operations. In some embodiments, the method is performed at a computing system including a non-volatile memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more image sensors via a communication interface. The method includes: acquiring image data associated with a physical environment via the one or more image sensors; acquiring first context information, the first context information including at least one of first user information associated with a current state of a user of the computing system, first application information associated with a first application executed by the computing system, and first environment information associated with the current state of the physical environment; selecting a first perspective correction operation set based at least in part on the first context information; generating first corrected image data by performing the first perspective correction operation set on the image data; and causing the first corrected image data to be presented via the display device.

[0023] According to some specific implementations, an electronic device includes one or more displays, one or more processors, non-volatile memory, and one or more programs; the one or more programs are stored in the non-volatile memory and are configured to be executed by one or more processors, and the one or more programs include instructions for executing or causing the execution of any of the methods described herein. According to some specific implementations, a non-volatile computer-readable storage medium stores instructions that, when executed by one or more processors of the device, cause the device to execute or cause the execution of any of the methods described herein. According to some specific implementations, a device includes: one or more displays, one or more processors, non-volatile memory, and a device for executing or causing the execution of any of the methods described herein.

[0024] According to some specific implementations, a computing system includes one or more processors, non-volatile memory, an interface for communicating with a display device and one or more input devices, and one or more programs; the one or more programs are stored in the non-volatile memory and are configured to be executed by one or more processors, and the one or more programs include instructions for performing or causing the execution of the operations of any of the methods described herein. According to some specific implementations, a non-volatile computer-readable storage medium stores instructions that, when executed by one or more processors of a computing system having an interface for communicating with a display device and one or more input devices, cause the computing system to perform or cause the execution of the operations of any of the methods described herein. According to some specific implementations, a computing system includes one or more processors, non-volatile memory, an interface for communicating with a display device and one or more input devices, and a component for performing or causing the execution of the operations of any of the methods described herein. DETAILED DESCRIPTION

[0025] Many details are described in order to provide a thorough understanding of the example implementations shown in the accompanying drawings. However, the accompanying drawings only illustrate some example aspects of the present disclosure and should not be considered limiting. One of ordinary skill in the art will appreciate that other effective aspects and / or variations do not include all of the specific details described herein. In addition, well-known systems, methods, components, devices, and circuits are not described in detail so as not to obscure more relevant aspects of the example implementations described herein.

[0026] A physical environment refers to a physical place that people can sense and / or interact with without the help of an electronic device. A physical environment may include physical features, such as a physical surface or a physical object. For example, a physical environment corresponds to a physical park including physical trees, physical buildings, and physical people. People can directly sense and / or interact with the physical environment, such as through vision, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people sense and / or interact with via an electronic device. For example, an XR environment may include augmented reality (AR) content, mixed reality (MR) content, virtual reality (VR) content, and the like. In the case of an XR system, a subset of a person's physical movements or a representation thereof is tracked, and in response, one or more characteristics of one or more virtual objects simulated in the XR environment are adjusted in a manner that complies with at least one physical law. For another example, an XR system may detect the movement of an electronic device (e.g., a mobile phone, a tablet computer, a laptop computer, a head-mounted device, etc.) that presents an XR environment, and in response, adjust the graphics content and sound field presented to a person by the electronic device in a manner similar to how such views and sounds would change in a physical environment. In some cases (eg, for accessibility reasons), the XR system may adjust characteristics of graphical content in the XR environment in response to representations of physical movement (eg, voice commands).

[0027] There are many different types of electronic systems that enable people to sense and / or interact with various XR environments. Examples include head-mounted systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be placed on people's eyes (e.g., similar to contact lenses), headphones / earpieces, speaker arrays, input systems (e.g., wearable or handheld controllers with or without tactile feedback), smartphones, tablet computers, and desktop / laptop computers. A head-mounted system may have an integrated opaque display and one or more speakers. Alternatively, a head-mounted system may be configured to accept an external opaque display (e.g., a smartphone). A head-mounted system may incorporate one or more imaging sensors for capturing images or video of a physical environment, and / or one or more microphones for capturing audio of a physical environment. Instead of an opaque display, a head-mounted system may have a transparent or translucent display. A transparent or translucent display may have a medium through which light representing an image is directed to a person's eyes. The display may utilize digital light projection, OLED, LED, μLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a hologram medium, an optical combiner, an optical reflector, or any combination thereof. In some implementations, a transparent or translucent display may be configured to selectively become opaque. A projection-based system may employ retinal projection technology that projects a graphic image onto a person's retina. The projection system may also be configured to project a virtual object into a physical environment, such as as a hologram or onto a physical surface.

[0028] As described above, in an HMD with a display and a scene camera, the image of the physical environment presented to the user on the display may not always reflect what the user sees if the HMD does not exist due to the different positions of the eyes, display, and camera in space. In various cases, this leads to, for example, poor distance perception, disorientation of the user, and poor hand-eye coordination when interacting with the physical environment.

[0029] Various perspective correction operations (also referred to herein as "POV correction modes / operations") may be used to enhance the comfort and / or aesthetics of the XR experience, such as depth clamping, POV position correction, depth smoothing, hole filling, time warping, etc. However, each of these perspective operation modes is associated with various tradeoffs, such as potential motion sickness, resource / power consumption, image quality, etc.

[0030] As an example, full POV position correction (e.g., correction for each of the X, Y, and Z transition offsets between the user and the camera perspective) may reduce potential motion sickness but may introduce image artifacts. Thus, in this example, full POV position correction may be suitable for an immersive video playback experience when the user is in motion. Conversely, when the user is stationary, partial POV position correction (e.g., correction for the X or X+Z transition offsets between the user and the camera perspective) or no POV position correction may be more appropriate and, in turn, also save computational resources / power as opposed to full POV position correction.

[0031] Thus, as described herein, the computing system selects a perspective correction operation set (and sets values ​​for adjustable parameters associated therewith) based on contextual information associated with at least one of: (A) a state of the user, (B) a current application executed by the computing system, and (C) a state of the physical environment. In some implementations, the computing system may also consider the accuracy of depth information associated with the physical environment, user preferences, user history, etc. when selecting the perspective correction operation set.

[0032] Figure 1 1 is a block diagram of an exemplary operating environment 100 according to some implementations. Although relevant features are shown, those of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. To this end, as a non-limiting example, the operating environment 100 includes a controller 110 and an electronic device 120.

[0033] In some implementations, the controller 110 is configured to manage and coordinate the user's XR experience (sometimes referred to herein as an "XR environment" or "virtual environment" or "graphical environment"). In some implementations, the controller 110 includes a suitable combination of software, firmware, and / or hardware. Fig. 9 Controller 110 is described in more detail. In some implementations, controller 110 is a computing device located locally or remotely relative to physical environment 105. For example, controller 110 is a local server located within physical environment 105. As another example, controller 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside physical environment 105. In some implementations, controller 110 is communicatively coupled to electronic device 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). As another example, controller 110 is included in a housing of electronic device 120. In some implementations, the functionality of controller 110 is provided by and / or combined with electronic device 120.

[0034] In some implementations, the electronic device 120 is configured to provide an XR experience to the user. In some implementations, the electronic device 120 includes a suitable combination of software, firmware, and / or hardware. According to some implementations, the electronic device 120 presents XR content (also sometimes referred to herein as "graphic content" or "virtual content") to the user via a display 122 while the user is physically present within a physical environment 105, which includes a table 107 within a field of view 111 of the electronic device 120. Thus, in some implementations, the user holds the electronic device 120 in one or both of his / her hands. In some implementations, when providing XR content, the electronic device 120 is configured to display an XR object (e.g., an XR cylinder 109) and implement video pass-through of the physical environment 105 (e.g., including a representation 117 of the table 107) on the display 122. The following description is relative to Fig.10 The electronic device 120 is described in more detail. According to some implementations, the electronic device 120 provides an XR experience to the user while the user is virtually and / or physically present within the physical environment 105.

[0035] In some implementations, the user wears the electronic device 120 on his / her head. For example, in some implementations, the electronic device includes a head-mounted system (HMS), a head-mounted device (HMD), or a head-mounted housing (HME). Thus, the electronic device 120 includes one or more displays provided for displaying XR content. For example, in various implementations, the electronic device 120 surrounds the user's field of view. In some implementations, the electronic device 120 is a handheld device (such as a smart phone or a tablet computer) configured to present XR content, and the user no longer wears the electronic device 120 but holds the device with the display facing the user's field of view and the camera facing the physical environment 105. In some implementations, the handheld device may be placed in a housing that can be worn on the user's head. In some implementations, the electronic device 120 is replaced with an XR cabin, housing, or room configured to present XR content, in which the user no longer wears or holds the electronic device 120.

[0036] Figure 2 An exemplary scenario 200 is shown involving capturing an image of an environment and displaying the captured image according to some implementations. A user wears a device (e.g., Figure 1The image sensor 230 captures an image of the physical environment, and the display 210 displays the image of the physical environment to the user's eye 220. The image sensor 230 has a viewing angle that is vertically offset from the user's viewing angle (e.g., where the user's eye 220 is located) by a vertical offset 241. Additionally, the viewing angle of the image sensor 230 is longitudinally offset from the user's viewing angle by a longitudinal offset 242. Additionally, in various implementations, the viewing angle of the image sensor 230 is laterally offset from the user's viewing angle by a lateral offset (e.g., toward Figure 2 within or offset outside of the page).

[0037] Figure 3 3 is an image 300 of a physical environment 301 captured by an image sensor from a specific perspective according to some specific implementations. The physical environment 301 includes a structure 310 having a first surface 311 closer to the image sensor, a second surface 312 farther from the image sensor, and a third surface 313 connecting the first surface 311 and the second surface 312. The first surface 311 has letters A, B, and C painted thereon, the third surface 313 has letters D painted thereon, and the second surface 312 has letters E, F, and G painted thereon.

[0038] From a particular viewing angle, image 300 includes all of the letters painted on structure 310. However, from other viewing angles, the captured image may not include all of the letters painted on structure 310.

[0039] Figure 4 yes Figure 3 301. The physical environment 301 includes a structure 310 and a wearable HMD 420 (e.g. Figure 1 4. The electronic device 120 in FIG. 4 is a user 410 of FIG. 410. The user 410 has a left eye 411a at a left eye position, thereby providing a left eye perspective. The user 410 has a right eye 411b at a right eye position, thereby providing a right eye perspective. The HMD 420 includes a left image sensor 421a at a left image sensor position that provides a left image sensor perspective. The HMD 420 includes a right image sensor 421b at a right image sensor position that provides a right image sensor perspective. Since the left eye 411a of the user 410 and the left image sensor 421a of the HMD 420 are at different positions, they each provide a different perspective of the physical environment.

[0040] Figure 5A5 is a block diagram of an input processing architecture 500 according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. To this end, as a non-limiting example, the input processing architecture 500 includes a computing system such as Figure 1 and Figure 2 The controller 110 is shown; Figure 1 and Figure 3 The electronic device 120 as shown; and / or suitable combinations thereof.

[0041] like Figure 5A As shown, one or more local sensors 502 of the controller 110, the electronic device 120 and / or a combination thereof obtain information related to Figure 1 The local sensor data 503 may include local sensor data 503 associated with the physical environment 105 in the physical environment 105. For example, the local sensor data 503 may include images or streams thereof of the user and / or his / her eyes, images or streams thereof of the physical environment 105, simultaneous localization and mapping (SLAM) information for the physical environment 105 and the position of the electronic device 120 or the user relative to the physical environment 105, ambient lighting information for the physical environment 105, ambient audio information for the physical environment 105, acoustic information for the physical environment 105, dimensional information for the physical environment 105, semantic labels for objects within the physical environment 105, etc. In some implementations, the local sensor data 503 may include unprocessed or post-processed information.

[0042] Similarly, if Figure 5A As shown, one or more remote sensors 504 associated with optional remote input devices, control devices 130, etc. within the physical environment 105 obtain information related to Figure 1 105. For example, the remote sensor data 505 includes images or streams thereof of the user and / or his / her eyes, images or streams thereof of the physical environment 105, SLAM information for the physical environment 105 and the position of the electronic device 120 or the user relative to the physical environment 105, ambient lighting information for the physical environment 105, ambient audio information for the physical environment 105, acoustic information for the physical environment 105, dimensional information for the physical environment 105, semantic labels for objects within the physical environment 105, etc. In some implementations, the remote sensor data 505 includes unprocessed or post-processed information.

[0043] exist Figure 5AIn the embodiment of the present invention, one or more depth sensors 506 of the controller 110, the electronic device 120, and / or a combination thereof obtain depth data 507 associated with the physical environment 105. For example, the depth data 507 includes depth values, a depth map, a depth grid, etc. for the physical environment 105. In some implementations, the depth data 507 includes unprocessed or post-processed information.

[0044] According to some specific implementations, the privacy architecture 508 ingests local sensor data 503, remote sensor data 505, and / or depth data 507. In some specific implementations, the privacy architecture 508 includes one or more privacy filters associated with user information and / or identification information. In some specific implementations, the privacy architecture 508 includes an opt-in feature in which the electronic device 120 informs the user which user information and / or identification information is being monitored and how the user information and / or identification information will be used. In some specific implementations, the privacy architecture 508 selectively prevents and / or limits the input processing architecture 500 or a portion thereof from obtaining and / or sending user information. To this end, the privacy architecture 508 receives user preferences and / or selections from the user in response to prompting the user to make user preferences and / or selections. In some specific implementations, the privacy architecture 508 prevents the input processing architecture 500 from obtaining and / or sending user information unless and until the privacy architecture 508 obtains informed consent from the user. In some specific implementations, the privacy architecture 508 anonymizes (e.g., scrambles, obfuscates, encrypts, etc.) certain types of user information. For example, privacy framework 508 receives user input specifying which types of user information privacy framework 508 anonymizes. As another example, privacy framework 508 independently of the user specifies (e.g., automatically) to anonymize certain types of user information that may include sensitive and / or identifying information.

[0045] According to some implementations, the motion state estimator 510 obtains the local sensor data 503 and the remote sensor data 505 after being processed by the privacy framework 508. In some implementations, the motion state estimator 510 obtains (e.g., receives, retrieves, or determines / generates) a motion state vector 511 based on the input data, and updates the motion state vector 511 over time.

[0046] Figure 5B An exemplary data structure for motion state vector 511 according to some specific implementations is shown. Figure 5BAs shown, the motion state vector 511 may correspond to an N-tuple representation vector or a representation tensor, which includes a timestamp 571 (e.g., the time when the motion state vector 511 was last updated), a motion state descriptor 572 for the electronic device 120 (e.g., stationary, in motion, car, ship, bus, train, airplane, etc.), a translational movement value 574 associated with the electronic device 120 (e.g., heading, velocity value, acceleration value, etc.), an angular movement value 576 associated with the electronic device 120 (e.g., angular velocity value, angular acceleration value, etc. for each of the pitch dimension, the roll dimension, and the yaw dimension), and / or miscellaneous information 578. A person of ordinary skill in the art will appreciate that the motion state descriptor 572 used for the electronic device 120 may be a time stamp, ... Figure 5B The data structure of the motion state vector 511 in is merely an example, and may include different information parts in various other specific implementations and may be constructed in various ways in various other specific implementations.

[0047] According to some implementations, the eye tracking engine 512 obtains the local sensor data 503 and the remote sensor data 505 after having been processed by the privacy framework 508. In some implementations, the eye tracking engine 512 determines / generates an eye tracking vector 513 associated with the user's gaze direction based on the input data, and updates the eye tracking vector 513 over time.

[0048] Figure 5B An exemplary data structure for eye tracking vector 513 is shown according to some specific implementations. Figure 5B As shown, the eye tracking vector 513 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 581 (e.g., the time when the eye tracking vector 513 was last updated), one or more angle values ​​582 (e.g., roll value, pitch value, and yaw value) for the user's current gaze direction, one or more translation values ​​584 (e.g., x value, y value, and z value relative to the physical environment 105, the entire world, etc.) for the user's current gaze direction, and / or miscellaneous information 586. A person of ordinary skill in the art will understand that Figure 5B The data structure of the eye tracking vector 513 in is merely an example and may include different information portions in various other implementations and may be constructed in a variety of ways in various other implementations.

[0049] For example, the gaze direction indicates a point in the physical environment 105 that the user is currently viewing (e.g., associated with an x-coordinate, a y-coordinate, and a z-coordinate relative to the physical environment 105 or the world as a whole), a physical object, or a region of interest (ROI). For another example, the gaze direction indicates a point in the XR environment that the user is currently viewing (e.g., associated with an x-coordinate, a y-coordinate, and a z-coordinate relative to the XR environment), an XR object, or a region of interest (ROI).

[0050] According to some implementations, the head / body pose tracking engine 514 obtains the local sensor data 503 and the remote sensor data 505 after being processed by the privacy framework 508. In some implementations, the head / body pose tracking engine 514 determines / generates a pose characterization vector 515 based on the input data and updates the pose characterization vector 515 over time.

[0051] Figure 5B An exemplary data structure for a posture representation vector 515 is shown according to some specific implementations. Figure 5B As shown, the posture representation vector 515 can correspond to an N-tuple representation vector or a representation tensor, which includes a timestamp 591 (e.g., the time when the posture representation vector 515 was last updated), a head posture descriptor 592A (e.g., up, down, neutral, etc.), a translation value 592B for the head posture, a rotation value 592C for the head posture, a body posture descriptor 594A (e.g., standing, sitting, prone, etc.), a translation value 594B for the body posture, a rotation value 594C for the body posture, and / or miscellaneous information 596. In some specific implementations, the posture representation vector 515 also includes information associated with finger / hand / limb tracking. Those of ordinary skill in the art will understand that Figure 5B The data structure of the posture representation vector 515 in is merely an example, and may include different information parts in various other specific implementations and may be constructed in various ways in various other specific implementations.

[0052] According to some implementations, the characterization engine 530 obtains the motion state vector 511, the eye tracking vector 513, and the gesture characterization vector 515. In some implementations, the characterization engine 530 obtains (e.g., receives, retrieves, or determines / generates) the user characterization vector 531 based on the motion state vector 511, the eye tracking vector 513, and the gesture characterization vector 515.

[0053] Figure 5B An exemplary data structure for a user characterization vector 531 according to some specific implementations is shown. Figure 5BAs shown, the user representation vector 531 can correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 5101 (e.g., the time when the user representation vector 531 was last updated), motion state information 5102 (e.g., the motion state descriptor 572), gaze direction information 5104 (e.g., a function of one or more angle values ​​582 and one or more translation values ​​584 within the eye tracking vector 513), head posture information 5106A (e.g., head posture descriptor 592A), body posture information 5106B (e.g., a function of body posture descriptor 594A within the posture representation vector 515), limb tracking information 5106C (e.g., a function of body posture descriptor 594A within the posture representation vector 515 associated with the user's limbs being tracked by the controller 110, the electronic device 120 and / or a combination thereof), location information 5108 (e.g., a home location (such as a kitchen or living room), a vehicle location (such as a car, an airplane, etc.), etc.) and / or miscellaneous information 5109.

[0054] According to some implementations, the environment analyzer 516 obtains the local sensor data 503 and the remote sensor data 505 after being processed by the privacy framework 508. In some implementations, the environment analyzer 516 determines / generates environment information 517 associated with the physical environment 105, and updates the environment information 517 over time. For example, the environment information 517 includes one or more of the following: a map of the physical environment 105 (e.g., including dimensions for the physical environment 105 and locations of objects therein), labels for the physical environment 105 (e.g., kitchen, bathroom, etc.), labels for one or more objects within the physical environment 105 (e.g., kitchen utensils, food, clothes, etc.), background frequency values, ambient audio information (e.g., volume, audio signatures / fingerprints for one or more ambient sounds, etc.), ambient light levels for the physical environment 105, etc.

[0055] According to some implementations, the depth evaluator 518 obtains the depth data 507 after it has been processed by the privacy framework 508. In some implementations, the depth evaluator 518 determines / generates a depth confidence value 519 for the depth data 507 or separate confidence values ​​for different portions of the depth data 507. For example, if the depth data 507 for the current time period has changed by more than a predefined tolerance relative to the depth data 507 for the previous time period, the depth data 507 for the current time period may be assigned a low confidence value. For another example, if the depth data 507 includes discontinuities or adjacent values ​​outside of a variance threshold, the depth data 507 may be assigned a low confidence value. As yet another example, if the depth data 507 for the previous time period causes comfort or rendering issues, the depth data 507 for the current time period may be assigned a low confidence value. In this example, comfort can be estimated based on various sensor data, such as eye fatigue (e.g., eye tracking data indicating blinking, squinting, etc.), gaze direction (e.g., indicating that the user is looking away), body posture (e.g., moving body posture, twisting body posture, etc.), head posture, etc.

[0056] Fig. 6A 6 is a block diagram of an exemplary content delivery architecture 600 according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. To this end, as a non-limiting example, the content delivery architecture 600 includes in a computing system having one or more processors and non-transitory memory, such as Figure 1 and Fig. 9 The controller 110 is shown; Figure 1 and Fig.10 The electronic device 120 as shown; and / or suitable combinations thereof.

[0057] In some implementations, the interaction processor 604 obtains (e.g., receives, retrieves, or detects) one or more user inputs 601, such as eye tracking input or gaze-based input, hand tracking input, touch input, voice input, etc. In various implementations, the interaction processor 604 determines appropriate modifications to the user interface or XR environment based on the one or more user inputs 601 (e.g., translating an XR object, rotating an XR object, modifying the appearance of an XR object, adding or removing an XR object, etc.).

[0058] In various specific implementations, the content manager 630 manages and updates the layout, settings, structure, etc. for the UI or XR environment based on one or more user inputs 601, including one or more of the VA, XR content, one or more UI elements associated with the XR content, etc. To this end, the content manager 630 includes a buffer 634, a content updater 636, and a feedback engine 638.

[0059] In some implementations, the buffer 634 includes XR content for one or more past instances and / or frames, rendered image frames, etc. In some implementations, the content updater 636 modifies the user interface or XR environment over time based on one or more other user inputs 601, etc. In some implementations, the feedback engine 638 generates sensory feedback (e.g., visual feedback (such as text or lighting changes), audio feedback, tactile feedback, etc.) associated with the user interface or XR environment based on one or more other user inputs 601, etc.

[0060] According to some specific implementations, reference Fig. 6A In the rendering engine 650 in the embodiment, the pose determiner 652 determines the current camera pose of the electronic device 120 and / or the user 605 relative to the XR environment and / or the physical environment 105 based at least in part on the pose representation vector 515. In some specific implementations, the renderer 654 renders XR content, etc. according to the current camera pose associated therewith.

[0061] According to some implementations, the perspective correction logic component 620A / 620B obtains (e.g., receives, retrieves, generates, captures, etc.) uncorrected image data 631 (e.g., an input image stream) from one or more image sensors 1014, including one or more images of the physical environment 105 from the current camera pose of the electronic device 120 and / or the user 605. In some implementations, the perspective correction logic component 620A / 620B generates corrected image data 633 (e.g., an output image stream) by performing one or more perspective correction operations on the uncorrected image data 631. FIG. 7A to FIG. 7C Various perspective correction operations are described in more detail. In various implementations, the perspective correction logic component 620A / 620B selects an initial (first) set of perspective correction operations based on one or more of the depth data 507, the environment information 517, the one or more depth confidence values ​​519, the user characterization vector 531, and the application data 611.

[0062] In some implementations, the perspective correction logic 620A / 620B acquires (e.g., receives, retrieves, generates, determines, etc.) the depth data 507, the environment information 517, one or more depth confidence values ​​519, the user characterization vector 531, and the application data 611. In some implementations, the perspective correction logic 620A / 620B detects a transition trigger based on a change to one of the environment information 517, the user characterization vector 531, and / or the application data 611, the change satisfying a significance threshold. In some implementations, the perspective correction logic 620A / 620B detects a transition trigger when the motion state information 5102 within the user characterization vector 531 is associated with a change satisfying a significance threshold (e.g., the user motion state changes from sitting to walking, or vice versa, or identifies that the user is performing other activities such as climbing stairs or exercising). In some implementations, the perspective correction logic 620A / 620B detects a transition trigger when the one or more depth confidence values ​​519 drops below the depth confidence threshold.

[0063] In various implementations, in response to detecting a transition trigger, the perspective correction logic component 620A / 620B modifies the one or more perspective correction operations. For example, in response to detecting a transition trigger, the perspective correction 620 may transition from performing a first (or initial) set of perspective correction operations to a second set of perspective correction operations that is different from the first set of perspective correction operations. In one example, the first set of perspective correction operations and the second set of perspective correction operations include at least one overlapping perspective correction operation. For another example, the first set of perspective correction operations and the second set of perspective correction operations include mutually exclusive perspective correction operations. For another example, in response to detecting a transition trigger, the perspective correction 620A / 620B may modify one or more adjustable parameters associated with the one or more perspective correction operations.

[0064] According to some implementations, the optional image processing architecture 662 obtains the corrected image data 633 from the perspective correction logic component 620A / 620B. In some implementations, the image processing architecture 662 also performs one or more image processing operations on the image stream, such as distortion, color correction, gamma correction, sharpening, noise reduction, white balance, etc. In other implementations, the optional image processing architecture 662 may perform the one or more image processing operations on the uncorrected image data 631 from one or more image sensors 1014. In some implementations, the optional compositor 664 synthesizes the rendered XR content with the processed image stream of the physical environment from the image processing architecture 662 to produce a rendered image frame of the XR environment. In various implementations, the renderer 670 presents the rendered image frame of the XR environment to the user 604 via one or more displays 1012. It will be understood by those of ordinary skill in the art that the optional image processing architecture 662 and the optional compositor 664 may not be applicable to a fully virtual environment (or optical pass-through scene).

[0065] Figure 6B 6 is a block diagram of a perspective correction logic component 620A according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. To this end, as a non-limiting example, the perspective correction logic component 620A is included in a computing system having one or more processors and non-transitory memory, such as Figure 1 and Fig. 9 The controller 110 is shown; Figure 1 and Fig.10 The electronic device 120 as shown; and / or suitable combinations thereof.

[0066] In some implementations, the perspective correction logic component 620A includes a transition detection logic component 622 and a perspective correction engine 640. Figure 6B As shown, transition detection logic component 622 obtains (e.g., receives, retrieves, generates, etc.) environment information 517, user characterization vector 531, and / or application data 611. According to some implementations, during initialization, perspective correction engine 640 obtains (e.g., receives, retrieves, generates, etc.) environment information 517, user characterization vector 531, and / or application data 611, and selects an initial (first) set of one or more perspective correction operations based thereon.

[0067] According to various implementations, when a change to at least one of the environmental information 517, the user characterization vector 531, and / or the application data 611 satisfies a significance threshold, the transition detection logic 622 detects a transition trigger 623. In some implementations, the transition detection logic 622 includes a buffer 627 that stores the environmental information 517, the user characterization vector 531, and / or the application data 611 for one or more previous time periods. Thus, the transition detection logic 622 can compare the environmental information 517, the user characterization vector 531, and / or the application data 611 for the current time period with the environmental information 517, the user characterization vector 531, and / or the application data 611 for one or more previous time periods to determine whether the change to at least one of the environmental information 517, the user characterization vector 531, and / or the application data 611 satisfies a significance threshold. For example, the significance threshold filters out small / insignificant changes in the context information. For example, the significance threshold corresponds to a deterministic value or a non-deterministic value.

[0068] In some implementations, the transition detection logic 622 provides context information 625 to the perspective correction engine 640, the context information including an indication of a change in at least one of the environment information 517, the user characterization vector 531, and / or the application data 611, the change satisfying a significance threshold. In some implementations, the transition detection logic 622 provides a transition trigger 623 to the perspective correction engine 640 to select a new (second) set of one or more perspective correction operations based on the context information 625.

[0069] like Figure 6B As shown, the perspective correction engine 640 includes an operation selector 642, an adjustable parameter modifier 644, an operation executor 646, and a transition executor 648. In some embodiments, in response to detecting or obtaining a transition trigger 623, the operation selector 642 selects a new (second) perspective correction operation set based on context information 625. In some embodiments, in response to detecting or obtaining a transition trigger 623, the adjustable parameter modifier 644 modifies the values ​​of one or more adjustable parameters associated with the initial (first) perspective correction operation set based on the context information 625. In some embodiments, in response to detecting or obtaining a transition trigger 623, the adjustable parameter modifier 644 sets the values ​​of one or more adjustable parameters associated with the new (second) perspective correction operation set based on the context information 625. Refer to the following Fig. 7A The adjustable parameters associated with the perspective correction operation are described in more detail.

[0070] exist Figure 6B, the perspective correction engine 640 obtains uncorrected image data 631 from one or more image sensors 1014. In some implementations, in response to detecting or obtaining a transition trigger 623, the operation performer 646 generates second corrected image data (e.g., corrected image data 633) by performing a new (second) set of perspective correction operations on the uncorrected image data 631 based on the depth data 507. In some other implementations, before detecting or obtaining the transition trigger 623, the operation performer 646 generates first corrected image data (e.g., corrected image data 633) by performing an initial (first) set of perspective correction operations on the uncorrected image data 631 based on the depth data 507.

[0071] According to various implementations, transition processor 648 implements a continuous or discrete transition from first corrected image data (e.g., the result of performing an initial (first) set of perspective correction operations on uncorrected image data 631) to second corrected image data (e.g., the result of performing a new (second) set of perspective correction operations on uncorrected image data 631). In some implementations, the transition from the first corrected image data to the second corrected image data can be accompanied by a fade-out animation of the first corrected image data and / or a fade-in animation of the second corrected image data. In some implementations, the computing system renders a blur effect to blur the transition from the first corrected image data to the second corrected image data.

[0072] In various implementations, the transition detection logic 622 samples the depth data 507, the environment information 517, the user characterization vectors 531, and / or the application data 611 at a first frame rate (e.g., once per second or a sampling frame rate), while the perspective correction engine 640 generates corrected image data 633 at a second frame rate (e.g., 90 times per second or a display frame rate).

[0073] In various implementations, the transition detection logic 622 stores a counter in a buffer 627 and, at each sampling period, increases the counter by one, decreases the counter by one, or leaves the counter unchanged. When the counter crosses a threshold, the transition detection logic 622 provides a transition trigger 623 to the perspective correction engine 640. For example, when the counter increases above a first threshold, the transition detection logic 622 provides the transition trigger 623 to the perspective correction engine 640 to select a first set of one or more perspective correction operations, and when the counter decreases below a second threshold, the transition detection logic 622 provides the transition trigger 623 to the perspective correction engine 640 to select a second set of one or more perspective correction operations.

[0074] In various implementations, the counter remains unchanged when the transition detection logic 622 detects that a person is nearby based on the semantic tags in the environment information 517 during the sampling period. Thus, the perspective correction engine 640 does not transition between different sets of one or more perspective operations when a person is nearby.

[0075] In various implementations, when the transition detection logic 622 detects the presence of a strong depth gradient based on the depth data 507 during the sampling period, the counter is increased, and when the transition detection logic 622 detects the presence of a weak depth gradient based on the depth data 507 during the sampling period, the counter is decreased. Therefore, the strength of the depth gradient is the main driving factor for the transition between different sets of one or more perspective operations. In various implementations, when the transition detection logic 622 detects that the strength of the depth gradient is neither strong nor weak during the sampling period, the transition detection logic 622 determines whether at least one object of a class of objects is present in the environment based on the semantic tags in the environment information 517. In various implementations, the class of objects includes tables, kitchen countertops, monitors, electronic devices, and similar objects. For example, objects such as monitors and tables are likely to occlude other objects and cause larger depth gradients, for which it may be challenging to perform perspective correction without artifacts. If the transition detection logic 622 determines that at least one of the class of objects is present in the environment, the counter is increased, and if the transition detection logic 622 determines that none of the class of objects is present in the environment, the counter is decreased. Therefore, the presence or absence of any of this class of objects is a secondary driver of transitions between different sets of one or more perspective operations.

[0076] In various specific implementations, in response to detecting or acquiring a transition trigger 623, the perspective correction engine 640 transitions between performing a first perspective correction operation set on the uncorrected image data 631 and performing a second perspective correction operation set on the uncorrected image data 631. However, the transition is performed over multiple frame periods. Therefore, between performing the first perspective correction operation set and performing the second perspective correction operation set, the perspective correction engine 640 performs multiple intermediate perspective correction operation sets on the uncorrected image data 631.

[0077] In various implementations, the transition between the first perspective operation set and the second perspective operation set occurs over a transition period. During each frame period of the transition period, the perspective correction engine 640 performs an intermediate perspective correction operation set that is a percentage between the first perspective operation set and the second perspective operation set. The percentage increases by an increment each frame period.

[0078] In various implementations, the increment is selected so that the transition is not noticeable to the user. In various implementations, the increment is a fixed value, and therefore the transition period is a fixed length. However, in various implementations, the increment is based on the depth data 507, the environment information 517, the user characterization vector 531, the application data 611, and / or the context information 625. In various implementations, these data sources are sampled by the perspective correction engine at a third frame rate, which can be any frame rate between the first frame rate (of the transition detection logic component 622) and the second frame rate (of the corrected image data 633). In various implementations, these data sources are received by the perspective correction engine 640 as asynchronous events.

[0079] In various implementations, the increment is based on the head posture information 5106A of the user characterization vector 531. For example, in various implementations, the faster the head posture changes, the larger the increment. Similarly, the faster the speed of the head posture changes, the larger the increment. In various implementations, the increment is based on the gaze direction information 5104 of the user characterization vector 531. For example, in various implementations, when a blink or glance is detected, the increment is large (up to, for example, 30%). In various implementations, the increment is based on the application data 611. For example, in various implementations, the increment is based on the distraction level of the application (which can be determined in part by the gaze direction information 5104). For example, during a gaming application with many moving objects, the increment may be larger than during a meditation application. For another example, if the user is reading a large amount of text in a virtual window, the increment may be larger than if the user is manipulating the position of a virtual object.

[0080] Figure 6C 6 is a block diagram of perspective correction logic 620B according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. Figure 6C Similar to and adapted from Figure 6B Therefore, in Figure 6B and Figure 6C Similar reference numbers are used between the , and for the sake of brevity, only the differences will be described herein. To this end, as a non-limiting example, the perspective correction logic component 620B is included in a computing system having one or more processors and non-transitory memory, such as Figure 1 and Fig. 9 The controller 110 is shown; Figure 1 and Fig.10 The electronic device 120 as shown; and / or suitable combinations thereof.

[0081] like Figure 6CAs shown, the transition detection logic component 622 obtains (e.g., receives, retrieves, generates, etc.) a depth confidence value 519 for the depth data 507 or separate confidence values ​​for different portions of the depth data 507. For example, if the depth data 507 for the current time period has changed by more than a predefined tolerance relative to the depth data 507 for the previous time period, the depth data 507 for the current time period may be assigned a low confidence value. For another example, if the depth data 507 includes discontinuities or adjacent values ​​outside of a variance threshold, the depth data 507 may be assigned a low confidence value. As yet another example, if the depth data 507 for the previous time period causes comfort or rendering issues, the depth data 507 for the current time period may be assigned a low confidence value.

[0082] According to various implementations, when the one or more depth confidence values ​​519 fall below the depth confidence threshold, the transition detection logic component 622 detects the transition trigger 623. In some implementations, in response to determining that the one or more depth confidence values ​​519 fall below the depth confidence threshold, the transition detection logic component 622 provides the transition trigger 623 to the perspective correction engine 640 to select a new (second) set of one or more perspective correction operations. As an example, when the one or more depth confidence values ​​519 fall below the depth confidence threshold, the operation selector 642 can disable some perspective correction operations that depend on the depth data 507 because the depth data 507 is currently inaccurate when the new (second) set of one or more perspective correction operations is selected.

[0083] Fig. 7A Various perspective correction operations are shown according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. To this end, Fig. 7A A perspective correction operation table 700 (also sometimes referred to herein as "table 700") and a perspective correction continuation set 725 (also sometimes referred to herein as "continuation set 725") are shown.

[0084] For example, table 700 shows various perspective correction operations and / or adjustable parameters associated therewith. Figure 6B During initialization, the perspective correction engine 640 selects an initial (first) perspective correction operation set from the table 700 and sets the adjustable parameters associated therewith based on the environment information 517, the user characterization vector 531 and / or the application data 611. According to some specific implementations, reference Figure 6BIn response to detecting the transition trigger 623 , the perspective correction engine 640 selects a new (second) perspective correction operation set from the table 700 and sets the adjustable parameters associated therewith based on the context information 625 .

[0085] like Fig. 7A As shown, table 700 includes an adjustable depth clamp distance parameter 702 for a depth clamp operation, the adjustable depth clamp distance parameter having various selectable values ​​(e.g., between 40 cm from the camera and a simulated infinite distance from the camera). Fig. 7A , table 700 also includes an adjustable depth clamp surface type parameter 704 for depth clamp operation, the adjustable depth clamp surface type parameter having different selectable surface types, such as flat, spherical, etc. Table 700 also includes an adjustable depth clamp surface orientation parameter 706 for depth clamp operation, the adjustable depth clamp surface orientation parameter having various selectable rotation values ​​for roll, pitch, and / or yaw of the clamped surface.

[0086] like Fig. 7A As shown, table 700 includes a POV position correction operation 708 for selecting an X translation offset value, a Y translation offset value, and / or a Z translation offset value to fully or partially account for the difference in viewing angle between the camera and the user. Figure 7B The POV position correction operation 708 is described in more detail.

[0087] like Fig. 7A As shown, table 700 includes a depth smoothing operation 710 with various selectable parameters, such as a selection between bilinear filtering or depth smoothing across edges during the depth smoothing operation 710, a smoothing kernel size for the depth smoothing operation 710, and a selection between prioritizing foreground or background stability during the depth smoothing operation 710. Fig. 7A In FIG. 7 , table 700 also includes an adjustable depth resolution parameter 712 having various selectable depth resolutions (eg, from 0 ppd to 3 ppd). Fig. 7A , table 700 also includes an adjustable depth frame rate parameter 714 having various selectable depth frame rates (eg, from 30 fps to 90 fps).

[0088] like Fig. 7A As shown, table 700 also includes a hole filling operation 716 with various optional hole filling techniques, such as no hole, black / shadow effect, and color dilation. Fig. 7AIn the embodiment, table 700 also includes a time warping operation 718 with various optional time warping techniques, such as six degrees of freedom (6DOF) time warping or three degrees of freedom (3DOF) time warping. As will be understood by those skilled in the art, table 700 may include other miscellaneous operations or adjustable parameters 720.

[0089] In various implementations, images from a scene camera are transformed so that they appear to be captured at the location of the user's eyes using a depth map that represents, for each pixel of the image, the distance from the camera to the object represented by the pixel. However, performing such a transformation takes time. In various implementations, performing the transformation includes generating a definition of the transformation and applying the transformation to the image. Therefore, when an image captured at a first time is transformed so that the image appears to have been captured at the location of the user's eyes at the first time, the image is displayed at a second time later than the first time. In addition, the user may have moved between the first time and the second time. Therefore, the transformed image is not the image that the user would see if the HMD was not present, but the image that the user would see at the first time if the HMD was not present. Therefore, in various implementations, the image captured at the first time is transformed according to the time warp 718 so that the image appears to be captured at the predicted location of the user's eyes at the second time when the transformed image is displayed. In some implementations, the time warp 718 takes into account 6DOF or 3DOF between the perspective of the scene camera and the eye.

[0090] like Fig. 7A As shown, continuation set 725 illustrates various tradeoffs between no perspective correction 726A and full perspective correction 726B. As an example, if no perspective correction operation is enabled (e.g., the "no perspective correction" 726A end of continuation set 725), the computing system will produce the least number of visual artifacts and also consume the least amount of power / resources. However, continuing with the example, if no perspective correction operation is enabled (e.g., the "no perspective correction" 726A end of continuation set 725), no comfort enhancement is enabled, which may result in poor hand-eye coordination and potential motion sickness in some contexts, but may result in acceptable hand-eye coordination and reduced likelihood of motion sickness in other user and environmental contexts.

[0091] As another example, if all perspective correction operations are enabled (e.g., "full perspective correction" 726B end of continuation set 725), the computing system will produce the greatest number of visual artifacts and also consume the greatest amount of power / resources. However, continuing with this example, if all perspective correction operations are enabled (e.g., "full perspective correction" 726B end of continuation set 725), the most comfort enhancements are enabled, which can result in improved hand-eye coordination and reduced likelihood of motion sickness. Thus, the computing system or a component thereof (e.g., Figure 6B and Figure 6C The perspective correction engine 640 in the embodiment selects a set of perspective correction operations based on the current user state, the current state of the physical environment, and / or the current foreground application in order to balance visual artifact generation and power / resource consumption with comfort and potential motion sickness.

[0092] As one example, when the computing system is presenting an immersive video experience 727, comfort and / or potential motion sickness may not be an important factor because the current user state indicates a sitting posture, and thus the computing system may perform fewer perspective correction operations. Conversely, as another example, when the computing system is presenting an XR gaming experience 729, comfort and / or potential motion sickness may be an important factor because the current user state indicates a standing / moving posture, and thus the computing system may perform more perspective correction operations.

[0093] Figure 7B A method 730 for performing the POV position correction operation 708 according to some implementations is shown. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. To this end, as a non-limiting example, the method 730 is performed by a computing system or a component thereof (e.g., Figure 6B The perspective correction engine 640 in the computing system or a component thereof has one or more processors and non-volatile memory, such as Figure 1 and Fig. 9 The controller 110 is shown; Figure 1 and Fig.10 The electronic device 120 shown; and / or a suitable combination thereof. In some implementations, the method 730 is performed by a processing logic component (including hardware, firmware, software, or a combination thereof). In some implementations, the method 730 is performed by a processor executing instructions (e.g., code) stored in a non-transitory computer-readable medium (e.g., memory).

[0094] like Figure 7B As shown, method 730 includes acquiring (e.g., receiving, retrieving, capturing, etc.) uncorrected image data 631. Figure 7BAs shown, method 730 also includes detecting, via one or more depth sensors (e.g., Fig.10 One or more depth sensors 506 of the electronic device 120 in the physical environment 105 acquire (e.g., receive, retrieve, capture, etc.) depth data 507 associated with the physical environment 105. In various implementations, the depth information includes a depth map for the physical environment. In various implementations, the depth map is based on an initial depth map in which the value of each pixel represents the depth from the object represented by the pixel. For example, the depth map can be a modified version of the initial depth map.

[0095] Thus, in various implementations, obtaining depth information of the physical environment includes determining depth values ​​for the camera's two-dimensional coordinate set (e.g., image space / plane) via interpolation using depth values ​​for positions surrounding the camera's two-dimensional coordinate set. In various implementations, the depth values ​​are determined using a three-dimensional model of the physical environment. For example, the depth values ​​may be determined using ray tracing from the camera's position through an image plane at a pixel position to a static object in the three-dimensional model. Thus, in various implementations, obtaining depth information of the physical environment includes determining depth values ​​for the camera's two-dimensional coordinate set based on the three-dimensional model of the physical environment.

[0096] In various implementations, the depth information of the physical environment is a smoothed depth map obtained by spatially filtering the initial depth map. In various implementations, the depth information of the physical environment is a clamped depth map, in which each pixel of the initial depth map having a value below a depth threshold is replaced by the depth threshold.

[0097] The method 730 also includes performing a POV position correction operation 708 on the uncorrected image data 631 by transforming the uncorrected image data 631 into corrected image data 633 based on the depth data 507. According to some implementations, the POV position correction operation 708 includes transforming a camera two-dimensional coordinate set (e.g., image space / plane) into a display two-dimensional coordinate set (e.g., display space / plane) based on the depth information. In various implementations, the transformation is based on a difference between the perspective of an image sensor that captures the image of the physical environment and the perspective of a user.

[0098] In various implementations, the display two-dimensional coordinate set is determined according to the following relationship, where x c and c is the camera 2D coordinate set, x d and d is the set of two-dimensional coordinates of the display, P c is the 4×4 view projection matrix of the image sensor representing the viewing angle of the image sensor, P dis the user’s 4×4 view projection matrix representing the user’s perspective, and d is the depth map value at the camera’s set of 2D coordinates:

[0099]

[0100] In various implementations, method 730 also includes determining an input three-dimensional coordinate set in the physical environment by triangulating the display two-dimensional coordinate set and the second display two-dimensional coordinate set. In various implementations, the second display two-dimensional coordinate set is obtained in a manner similar to the display two-dimensional coordinate set for a second camera plane or a second image sensor (e.g., for a second eye of a user), where the display two-dimensional coordinate set is determined for a first eye of the user. For example, in various implementations, the device projects the physical three-dimensional coordinate set to a second image plane to obtain a second camera two-dimensional coordinate set, and transforms it using the depth information to generate a second display two-dimensional coordinate set.

[0101] Figure 7C Various perspective correction scenarios are shown according to some implementations. Although relevant features are shown, one of ordinary skill in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure more relevant aspects of the example implementations disclosed herein. To this end, Figure 7C Table 750 is shown with scenes 751A, 751B, and 751C.

[0102] like Figure 7C As shown, scenario 751A corresponds to a situation in which the user is watching a movie. With respect to scenario 751A, power consumption is prioritized and potential motion sickness is not prioritized because the user is not moving. For example, referring to scenario 751A, the computing system or its components (e.g., Figure 6B The perspective correction engine 640 in the 3D image processing module selects the following perspective correction operations and / or sets the following values ​​for various adjustable parameters: full POV position correction operation (e.g., X translation offset, Y translation offset, and Z translation offset to align camera and user perspective), 3DOF time warp operation on pass-through image data, simulates depth clamp distance to infinity, and sets depth resolution to 0 ppd. Thus, the potential for motion sickness will be high, power usage will be low, image quality will be high, and hand-eye coordination will be low.

[0103] like Figure 7C As shown, scene 751B corresponds to a situation where the user is in a crowded meeting place (e.g., a train station, an airport, a stadium, etc.) where depth estimation may be unstable and background frequency may be high. Relative to scene 751B, reducing visual artifacts is a priority. For example, referring to scene 751B, the computing system or its components (e.g., Figure 6B The perspective correction engine 640 in the image processing unit 640 selects the following perspective correction operations and / or sets the following values ​​for various adjustable parameters: partial POV position correction operation (e.g., accounting for X translation offset between camera and user perspective), 6DOF time warp operation, clamp plane at 70 cm from camera, and sets depth resolution to 1.5 ppd. Thus, the likelihood for motion sickness will be medium, power usage will be high, image quality will be high, and hand-eye coordination will be medium.

[0104] like Figure 7C As shown, scene 751C corresponds to a situation in which the user is working at a desk or browsing around a store / museum where the user is at least partially in motion. With respect to scene 751C, potential motion sickness is prioritized. For example, with reference to scene 751C, the computing system or its components (e.g., Figure 6B The perspective correction engine 640 in the 3D image processing unit 640 selects the following perspective correction operations and / or sets the following values ​​for various adjustable parameters: partial POV position correction operation (e.g., taking into account X and Z translation offsets between the camera and the user's perspective), 6DOF time warp operation, clamping plane at 70 cm from the camera, and setting the depth resolution to 1.5 ppd. Thus, the likelihood for motion sickness will be low, power usage will be high, image quality will be medium, and hand-eye coordination will be high.

[0105] Fig. 8A and Figure 8B 800 is a flowchart representation of a method 800 for context-based selection of perspective correction operations according to some implementations. In various implementations, the method 800 is performed at a computing system including non-transitory memory and one or more processors, wherein the computing system communicates with the user via a communication interface (e.g., Figure 1 and Fig. 9 The controller 110, Figure 1 and 10 The method 800 is a method for controlling a display device and an image sensor. ...

[0106] As represented by block 802, method 800 includes acquiring (e.g., receiving, retrieving, capturing, etc.) image data associated with a physical environment via the one or more image sensors. For example, the one or more image sensors correspond to a scene / outward-facing camera. As an example, referring to Figure 6B and Figure 6C , the computing system or a component thereof (e.g., the perspective correction engine 640) obtains uncorrected image data 631 from one or more image sensors 1014.

[0107] As represented by block 804, method 800 includes acquiring (e.g., receiving, retrieving, capturing, generating, etc.) first context information, the first context information including at least one of first user information associated with a current state of a user of the computing system, first application information associated with a first application executed by the computing system, and first environment information associated with a current state of a physical environment. In some implementations, the first context information is associated with a first time or time period. Thus, in various implementations, the computing system can continuously or periodically monitor and update the context information over time as the context information changes. As an example, referring to Figure 6B , the computing system or a component thereof (eg, transition detection logic component 622 ) acquires (eg, receives, retrieves, generates, etc.) environmental information 517 , user characterization vector 531 , and / or application data 611 .

[0108] In some implementations, the first user information includes at least one of head posture information, body posture information, eye tracking information, and user motion information that characterizes a current state of a user of the computing system. Figure 5A , Figure 5B and Fig. 6A The user characterization vector 531 is described in further detail.

[0109] In some specific implementations, the first application information includes at least an identifier associated with the first application and state information associated with the first application. For example, the first application information corresponds to the reference Fig. 6A and Figure 6B The application data 611 is described in further detail.

[0110] In some implementations, the first environmental information includes at least one of a label for the physical environment representing a current state of the physical environment, a background frequency value, a label for one or more objects within the physical environment, ambient audio information, and an ambient light level. For example, the first environmental information corresponds to a reference Figure 5A and Fig. 6A Environmental information 517 is described in further detail.

[0111] As represented by block 806, method 800 includes selecting a first set of perspective correction operations based at least in part on the first context information. In some implementations, the computing system also determines parameter values ​​for each perspective correction operation in the first set of perspective correction operations. In some implementations, the computing system determines the first set of perspective correction operations by balancing the effects of the first set of perspective correction operations on various user experience metrics (such as potential motion sickness, power usage, image quality, and hand-eye coordination) in accordance with the first context information. As an example, Fig. 7A The continuation set 725 in FIG. 7 shows various tradeoffs between “no perspective correction” 726A and “full perspective correction” 726B. For example, Figure 7C Table 750 in shows various perspective correction operations activated for different scenes 751A, 751B, and 751C.

[0112] As an example, see Figure 6B , the computing system or a component thereof (e.g., perspective correction engine 640) obtains (e.g., receives, retrieves, generates, etc.) environment information 517, user characterization vector 531, and / or application data 611, and selects an initial (first) set of one or more perspective correction operations based thereon during initialization. In some implementations, selecting the first set of perspective correction operations includes determining a value of an adjustable parameter associated with each perspective correction operation in the first set of perspective correction operations based at least in part on the first context information. For example, Fig. 7A Table 700 in shows various perspective correction operations and / or adjustable parameters associated therewith. In some implementations, the first perspective correction operation set includes at least one of depth clamping, point of view (POV) position correction, depth smoothing, hole filling, and time warping.

[0113] As an example, adjustable parameters for depth clamping operations include clamping distance, clamping surface type, clamping surface orientation, etc. As another example, adjustable parameters for POV position correction operations include full or partial X translation offset, Y translation offset, and / or Z translation offset to account for differences between user and camera perspectives. As yet another example, adjustable parameters for time warping operations include 6DOF time warping or 3DOF time warping.

[0114] As represented by block 808, method 800 includes generating first corrected image data by performing a first set of perspective correction operations on the image data. In some implementations, the computing system generates the first corrected image stream by performing the first set of perspective correction operations on the uncorrected image stream based on the current depth information. As an example, referring to Figure 6B, the computing system or a component thereof (e.g., operation performer 646) generates first corrected image data (e.g., corrected image data 633) by performing an initial (first) set of perspective correction operations on the uncorrected image data 631 based on the depth data 507. For example, the depth information includes a depth value for each pixel in the uncorrected image stream, a depth map for the physical environment, etc.

[0115] As represented by block 810, method 800 includes causing presentation of first corrected image data via a display device. As an example, referring to Fig. 6A , the computing system or a component thereof (e.g., the renderer 670 ) presents the corrected image data 633 (or a derivative thereof synthesized with the rendered XR content) via one or more displays 1012 .

[0116] According to some implementations, as represented by block 812, method 800 includes detecting a change from first context information to second context information, the second context information including at least one of second user information associated with a current state of a user of the computing system, second application information associated with a second application different from the first application executed by the computing system, and second environment information associated with a current state of a physical environment. In some implementations, the computing system generates a transition trigger in response to detecting a change in the context information that satisfies a significance threshold.

[0117] As an example, see Figure 6B , when a change to at least one of the environmental information 517, the user characterization vector 531, and / or the application data 611 satisfies a significance threshold, the computing system or a component thereof (e.g., the transition detection logic component 622) detects a transition trigger 623. In some implementations, the transition detection logic component 622 includes a buffer 627 that stores the environmental information 517, the user characterization vector 531, and / or the application data 611 for one or more previous time periods. Thus, the transition detection logic component 622 can compare the environmental information 517, the user characterization vector 531, and / or the application data 611 for the current time period with the environmental information 517, the user characterization vector 531, and / or the application data 611 for one or more previous time periods to determine whether the change to at least one of the environmental information 517, the user characterization vector 531, and / or the application data 611 satisfies a significance threshold. For example, the significance threshold filters out small / insignificant changes in the context information. For example, the significance threshold corresponds to a deterministic value or a non-deterministic value.

[0118] In some implementations, the first set of perspective correction operations and the second set of perspective correction operations include mutually exclusive perspective correction operations. In some implementations, the first set of perspective correction operations and the second set of perspective correction operations include at least one overlapping perspective correction operation. In some implementations, the at least one overlapping perspective correction operation includes an adjustable parameter having a first value in the first set of perspective correction operations, and wherein the at least one overlapping perspective correction operation includes an adjustable parameter having a second value different from the first value in the second set of perspective correction operations. In some implementations, the first value and the second value are deterministic values ​​(e.g., one of a plurality of predetermined values) or non-deterministic values ​​(e.g., dynamically calculated based on current context information).

[0119] In some implementations, the change from the first context information to the second context information corresponds to a transition from a first application executed by the computing system to a second application different from the first application executed by the computing system. For example, the transition from the first application to the second application corresponds to a transition from an immersive video playback application to an interactive gaming application, or vice versa.

[0120] In some implementations, a change from the first context information to the second context information corresponds to a change in the current state of the user associated with a transition from a first user motion state to a second user motion state different from the first user motion state. For example, the transition from the first user motion state to the second user motion state corresponds to a transition from sitting to walking, or vice versa.

[0121] In some implementations, a change from the first context information to the second context information corresponds to a change in the current state of the physical environment associated with a transition from a first ambient light level detected within the physical environment to a second ambient light level that is different from the first ambient light level detected within the physical environment. For example, the transition from the first ambient light level to the second ambient light level corresponds to a transition from indoors to outdoors, or vice versa.

[0122] In some implementations, a change from the first context information to the second context information corresponds to a change in the current state of the physical environment associated with a transition from a first background associated with a first frequency value within the physical environment to a second background associated with a second frequency value different from the first frequency value within the physical environment. For example, a transition from a first background associated with a first frequency value to a second background associated with a second frequency value corresponds to a transition from a first background associated with a blank wall to a second background associated with a cluttered collection of shelves, or vice versa.

[0123] According to some implementations, as represented by block 814, in response to detecting a change from the first context information to the second context information, method 800 includes: selecting a second set of perspective correction operations that is different from the first set of perspective correction operations performed on the image data; generating second corrected image data by performing the second set of perspective correction operations on the uncorrected image stream; and causing presentation of the second corrected image data via a display device. In some implementations, the computing system determines the second set of perspective correction operations by balancing the effects of the second set of perspective correction operations on various user experience metrics, such as potential motion sickness, power usage, image quality, and hand-eye coordination.

[0124] As an example, see Figure 6B In response to detecting or acquiring a transition trigger 623, the computing system or a component thereof (e.g., an operation selector 642) selects a new (second) set of perspective correction operations based on context information 625. According to some implementations, context information 625 includes an indication of a change in at least one of environmental information 517, user characterization vector 531, and / or application data 611, the change satisfying a significance threshold. Continuing with this example, referring to Figure 6B In response to detecting or acquiring the transition trigger 623, the computing system or a component thereof (e.g., the operation executor 646) generates second corrected image data (e.g., corrected image data 633) by performing a new (second) set of perspective correction operations on the uncorrected image data 631 based on the depth data 507. Continuing with the example, referring to Fig. 6A , the computing system or a component thereof (e.g., the renderer 670) presents the second corrected image data 633 (or a derivative thereof synthesized with the rendered XR content) via one or more displays 1012.

[0125] In some implementations, method 800 includes causing, via a display device, presentation of an animation associated with the first corrected image data before presenting the second corrected image data. In some implementations, the transition from the first corrected image data to the second corrected image data may be accompanied by a fade-out animation of the first corrected image data and / or a fade-in animation of the second corrected image data. In some implementations, the computing system presents a blur effect to blur the transition from the first corrected image data to the second corrected image data. In some implementations, method 800 includes stopping presentation of the first corrected image data before presenting the second corrected image data. As an example, refer to Figure 6BIn response to detecting or acquiring a transition trigger 623, the computing system or a component thereof (e.g., a transition processor 648) implements a continuous or discrete transition from first corrected image data (e.g., the result of performing an initial (first) set of perspective correction operations on the uncorrected image data 631) to second corrected image data (e.g., the result of performing a new (second) set of perspective correction operations on the uncorrected image data 631).

[0126] In various implementations, detecting a change from the first context information to the second context information includes storing a counter and increasing or decreasing the counter based on at least one of the second user information, the second application information, and the second environment information. In various implementations, the computing system detects the change from the first context information to the second context information in response to the counter breaking a threshold.

[0127] In various specific implementations, increasing or decreasing the counter includes increasing the counter based on a determination that the second environment information indicates a strong depth gradient of the current environment, and decreasing the counter based on a determination that the second environment information indicates a weak depth gradient of the current environment. In various specific implementations, increasing or decreasing the counter includes increasing the counter based on a determination that the second environment information indicates a medium depth gradient of the current environment and the current environment includes at least one object of a class of objects, and decreasing the counter based on a determination that the second environment information indicates a medium depth gradient of the current environment and the current environment does not include any of the class of objects.

[0128] In various implementations, presenting the second corrected image data occurs during a transition period after presenting the first corrected image data, and method 800 further includes selecting an intermediate perspective correction operation set at an intermediate time within the transition period that is a percentage between the first perspective correction operation set and the second perspective correction operation set. Method 800 also includes generating intermediate corrected image data by performing the intermediate perspective correction operation set on the uncorrected image stream, and causing presentation of the intermediate corrected image data via a display device.

[0129] In various implementations, the percentage is based on second user information indicating a change in the user's head posture. In various implementations, the percentage is based on second application information indicating a distraction level of the second application.

[0130] According to some specific implementations, the computing system is further communicatively coupled to one or more depth sensors via a communication interface, and method 800 also includes: obtaining depth information associated with the physical environment via the one or more depth sensors, wherein the first corrected image data is generated by performing a first perspective correction operation set on the image data based on the depth information. In some specific implementations, method 800 also includes: determining one or more confidence values ​​for the depth information, wherein the first perspective correction operation set is determined at least in part based on the first context information or the one or more confidence values ​​for the depth information. For example, the one or more confidence values ​​for the depth information include a pixel-by-pixel confidence value, an image-by-image confidence value, etc.

[0131] As an example, see Figure 6C , when one or more depth confidence values ​​519 drop below a depth confidence threshold, the computing system or a component thereof (e.g., transition detection logic component 622) detects a transition trigger 623. Continuing with this example, referring to Figure 6C , when one or more depth confidence values ​​519 drop below a depth confidence threshold, the computing system or a component thereof (e.g., operation selector 642) may disable some perspective correction operations that depend on the depth data 507 because the depth data 507 is currently inaccurate when a new (second) set of one or more perspective correction operations is selected.

[0132] Fig. 9 is a block diagram of an example of a controller 110 according to some implementations. While certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some embodiments, the controller 110 includes one or more processing units 902 (e.g., a microprocessor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a graphics processing unit (GPU), a central processing unit (CPU), a processing core, and / or the like), one or more input / output (I / O) devices 906, one or more communication interfaces 908 (e.g., a universal serial bus (USB), FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), infrared (IR), Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 910, a memory 920, and one or more communication buses 904 for interconnecting these components and various other components.

[0133] In some implementations, one or more communication buses 904 include circuits that interconnect system components and control communications between system components. In some implementations, one or more I / O devices 906 include at least one of a keyboard, a mouse, a touchpad, a joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0134] The memory 920 includes a high-speed random access memory, such as a dynamic random access memory (DRAM), a static random access memory (SRAM), a double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some specific implementations, the memory 920 includes a non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. The memory 920 optionally includes one or more storage devices remotely located from one or more processing units 902. The memory 920 includes a non-transitory computer-readable storage medium. In some specific implementations, the memory 920 or the non-transitory computer-readable storage medium of the memory 920 stores the following programs, modules, and data structures or their subsets, including an optional operating system 930, an XR experience module 940, an input processing architecture 500, a perspective correction logic component 620A / 620B, an interactive processor 604, a content manager 630, and a rendering engine 650.

[0135] The operating system 930 includes procedures for handling various basic system services and for performing hardware-related tasks. In some specific implementations, the XR experience module 940 is configured to manage and coordinate single or multiple XR experiences of one or more users (e.g., a single XR experience of one or more users, or multiple XR experiences of corresponding groups of one or more users). To this end, in various specific implementations, the XR experience module 940 includes a data acquisition unit 942, a tracking unit 944, a coordination unit 946, and a data sending unit 948.

[0136] In some specific implementations, the data acquisition unit 942 is configured to at least Figure 1 The electronic device 120 acquires data (e.g., presentation data, interaction data, sensor data, location data, etc.). To this end, in various specific implementations, the data acquisition unit 942 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0137] In some implementations, the tracking unit 944 is configured to map the physical environment 105 and to track at least the electronic device 120 relative to the physical environment 105. Figure 1The location / position of the physical environment 105. To this end, in various implementations, the tracking unit 944 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0138] In some implementations, the coordination unit 946 is configured to manage and coordinate the XR experience presented to the user by the electronic device 120. To this end, in various implementations, the coordination unit 946 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0139] In some implementations, the data transmission unit 948 is configured to transmit data (e.g., presentation data, location data, etc.) to at least the electronic device 120. To this end, in various implementations, the data transmission unit 948 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0140] In some implementations, the input processing architecture 500 is configured to process input data, as described above with reference to Figure 5A To this end, in various implementations, the input processing architecture 500 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0141] In some implementations, the perspective correction logic 620A / 620B is configured to generate corrected image data 633 by performing one or more perspective correction operations on the uncorrected image data 631, as described above with reference to FIG. 6A to FIG. 6C To this end, in various implementations, the perspective correction logic component 620A / 620B includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0142] In some implementations, the interaction processor 604 is configured to obtain one or more user inputs 601, such as eye tracking input or gaze-based input, hand tracking input, touch input, voice input, etc., as described above with reference to Fig. 6A To this end, in various implementations, the interaction processor 604 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0143] In some implementations, the content manager 630 is configured to manage and update the layout, settings, structure, etc. for the UI or XR environment based on one or more user inputs 601, including one or more of the VA, XR content, one or more UI elements associated with the XR content, etc., as described above with reference to Fig. 6ATo this end, in various implementations, the interaction processor 604 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0144] In some implementations, the rendering engine 650 is configured to render UI or XR content, as described above with reference to Fig. 6A As shown. Fig. 6A As shown, rendering engine 650 includes a pose determiner 652, a renderer 654, an optional image processing architecture 662, and an optional compositor 664. To this end, in various implementations, rendering engine 650 includes instructions and / or logic for such instructions as well as heuristics and metadata for such heuristics.

[0145] Although the data acquisition unit 942, tracking unit 944, coordination unit 946, data sending unit 948, input processing architecture 500, perspective correction logic component 620A / 620B, interaction processor 604, content manager 630 and rendering engine 650 are shown as residing on a single device (e.g., controller 110), it should be understood that in other specific embodiments, any combination of the data acquisition unit 942, tracking unit 944, coordination unit 946, data sending unit 948, input processing architecture 500, perspective correction logic component 620A / 620B, interaction processor 604, content manager 630 and rendering engine 650 may be located in a separate computing device.

[0146] also, Fig. 9 It serves more as a functional description of various features that may be present in a particular implementation, rather than as a structural schematic diagram of the implementation described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Fig. 9 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software and / or firmware selected for a specific specific implementation.

[0147] Fig.10is a block diagram of an example of an electronic device 120 according to some implementations. While certain specific features are illustrated, those skilled in the art will recognize from this disclosure that various other features are not illustrated for the sake of brevity and so as not to obscure more relevant aspects of the implementations disclosed herein. To this end, as a non-limiting example, in some specific implementations, the electronic device 120 includes one or more processing units 1002 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 1006, one or more communication interfaces 1008 (e.g., USB, FIREWIRE, THUNDERBOLT, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 1010, one or more displays 1012, one or more optional internal-facing and / or external-facing image sensors 1014, one or more optional depth sensors 1016, a memory 1020, and one or more communication buses 1004 for interconnecting these components and various other components.

[0148] In some implementations, the one or more communication buses 1004 include circuits that interconnect system components and control communications between system components. In some implementations, the one or more I / O devices and sensors 1006 include at least one of an inertial measurement unit (IMU), an accelerometer, a magnetometer, a gyroscope, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen sensor, a blood glucose sensor, etc.), one or more microphones, one or more speakers, a haptic engine, etc.

[0149] In some implementations, one or more displays 1012 are configured to provide a user interface or XR experience to the user. In some implementations, one or more displays 1012 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field effect transistor (OLET), organic light-emitting diode (OLED), surface conduction electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical system (MEMS) and / or similar display types. In some implementations, one or more displays 1012 correspond to diffraction, reflection, polarization, holographic and other waveguide displays. For example, the electronic device 120 includes a single display. As another example, the electronic device includes a display for each eye of the user. In some implementations, one or more displays 1012 can present MR and VR content.

[0150] In some implementations, the one or more image sensors 1014 are configured to acquire image data corresponding to at least a portion of the user's face (including the user's eyes) (and may be referred to as an eye tracking camera). In some implementations, the one or more image sensors 1014 are configured to face forward so as to acquire image data corresponding to the physical environment that the user would see when the electronic device 120 is not present (and may be referred to as a scene camera). The one or more optional image sensors 1014 may include one or more RGB cameras (e.g., with a complementary metal oxide semiconductor (CMOS) image sensor or a charge coupled device (CCD) image sensor), one or more infrared (IR) cameras, one or more event-based cameras, and the like. In some implementations, the one or more depth sensors 506 correspond to a structured light device, a time of flight device, and the like.

[0151] Memory 1020 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some specific implementations, memory 1020 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 1020 optionally includes one or more storage devices remotely located from one or more processing units 1002. Memory 1020 includes non-transitory computer-readable storage media. In some specific implementations, memory 1020 or non-transitory computer-readable storage media of memory 1020 stores the following programs, modules, and data structures or subsets thereof, including an optional operating system 1030 and an XR rendering module 1040.

[0152] The operating system 1030 includes processes for handling various basic system services and for performing hardware-related tasks. In some specific implementations, the XR rendering module 1040 is configured to present XR content to the user via one or more displays 1012. To this end, in various specific implementations, the XR rendering module 1040 includes a data acquisition unit 1042, a renderer 670, and a data sending unit 1048.

[0153] In some specific implementations, the data acquisition unit 1042 is configured to at least Figure 1 The controller 110 acquires data (eg, presentation data, interaction data, sensor data, location data, etc.) To this end, in various implementations, the data acquisition unit 1042 includes instructions and / or logic components for these instructions as well as heuristics and metadata for the heuristics.

[0154] In some implementations, the renderer 670 is configured to display the transformed image via one or more displays 1012. To this end, in various implementations, the renderer 670 includes instructions and / or logic for such instructions as well as heuristics and metadata for such heuristics.

[0155] In some implementations, the data sending unit 1048 is configured to send data (e.g., presentation data, location data, etc.) to at least the controller 110. In some implementations, the data sending unit 1048 is configured to send authentication credentials to the electronic device. To this end, in various implementations, the data sending unit 1048 includes instructions and / or logic components for these instructions and heuristics and metadata for the heuristics.

[0156] Although the data acquisition unit 1042, the renderer 670, and the data sending unit 1048 are illustrated as residing on a single device (e.g., the electronic device 120), it should be understood that in other implementations, any combination of the data acquisition unit 1042, the renderer 670, and the data sending unit 1048 may be located in separate computing devices.

[0157] also, Fig.10 It is more of a functional description of various features that may be present in a particular implementation, as opposed to a schematic diagram of the structure of the implementation described herein. As one of ordinary skill in the art will recognize, items shown separately may be combined, and some items may be separated. For example, Fig.10 Some functional modules shown separately in the figure may be implemented in a single module, and the various functions of a single functional block may be implemented by one or more functional blocks in various specific implementations. The actual number of modules and the division of specific functions and how the features are distributed among them will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software and / or firmware selected for a specific specific implementation.

[0158] Although various aspects of specific implementations within the scope of the appended claims are described above, it should be apparent that the various features of the above-mentioned specific implementations can be embodied in a variety of forms, and any specific structures and / or functions described above are merely illustrative. Based on this disclosure, it should be understood by those skilled in the art that the aspects described herein can be implemented independently of any other aspects, and two or more of these aspects can be combined in various ways. For example, a device and / or a method can be implemented using any number of aspects set forth herein. In addition, in addition to or different from one or more aspects set forth herein, such a device and / or such a method can be implemented using other structures and / or functions.

[0159] It will also be understood that, although the terms "first", "second", etc. may be used to describe various elements in this article, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element. For example, a first node may be referred to as a second node, and similarly, a second node may be referred to as a first node, which changes the meaning of the description, as long as all occurrences of the "first node" are consistently renamed and all occurrences of the "second node" are consistently renamed. Both the first node and the second node are nodes, but they are not the same node.

[0160] The terms used herein are only for describing specific implementations and are not intended to limit the claims. As used in the description of this specific implementation and the appended claims, the singular forms of "one", "a kind of" and "the" are intended to also cover the plural forms, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" used herein refers to and covers any and all possible combinations of one or more items in the associated listed items. It will also be understood that the term "comprising" when used in this specification specifies the presence of stated features, integers, steps, operations, elements and / or parts, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, parts, and / or their groupings.

[0161] As used herein, the term "if" may be interpreted to mean "when the antecedent is true" or "when the antecedent is true" or "in response to determining" or "upon determining" or "in response to detecting" that the antecedent is true, depending on the context. Similarly, the phrase "if it is determined that [the antecedent is true]" or "if [the antecedent is true]" or "when [the antecedent is true]" is interpreted to mean "upon determining that the antecedent is true" or "in response to determining" or "upon determining" that the antecedent is true or "when detecting that the antecedent is true" or "in response to detecting" that the antecedent is true, depending on the context.

Claims

1. A method, comprising: At a computing system including non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device and one or more image sensors via a communication interface: acquiring, via the one or more image sensors, image data associated with a physical environment; obtaining first context information, the first context information comprising at least one of first user information associated with a current state of a user of the computing system, first application information associated with a first application executed by the computing system, and first environment information associated with a current state of the physical environment; selecting a first set of perspective correction operations based at least in part on the first contextual information; generating first corrected image data by performing the first set of perspective correction operations on the image data; and The first corrected image data is caused to be presented via the display device.

2. The method of claim 1, wherein the first set of perspective correction operations comprises at least one of depth clamping, point-of-view (POV) position correction, depth smoothing, hole filling, and time warping.

3. The method according to any one of claims 1 to 2, wherein the first user information comprises at least one of head posture information, body posture information, eye tracking information, and user motion information characterizing the current state of the user of the computing system. 4 . The method according to claim 1 , wherein the first application information comprises at least an identifier associated with the first application and state information associated with the first application.

5. A method according to any one of claims 1 to 4, wherein the first environmental information includes at least one of a label for the physical environment, a background frequency value, a label for one or more objects within the physical environment, ambient audio information, and an ambient light level characterizing the current state of the physical environment.

6. The method of any one of claims 1 to 5, wherein selecting the first set of perspective correction operations comprises determining a value for an adjustable parameter associated with each perspective correction operation in the first set of perspective correction operations based at least in part on the first context information.

7. The method according to any one of claims 1 to 6, further comprising: detecting a change from the first context information to second context information, the second context information comprising at least one of second user information associated with the current state of a user of the computing system, second application information associated with a second application different from the first application executed by the computing system, and second environment information associated with the current state of the physical environment; as well as In response to detecting the change from the first context information to the second context information: selecting a second set of perspective correction operations different from the first set of perspective correction operations performed on the image data; generating second corrected image data by performing the second set of perspective correction operations on the uncorrected image stream; and The second corrected image data is caused to be presented via the display device.

8. The method of claim 7, wherein the first set of perspective correction operations and the second set of perspective correction operations comprise mutually exclusive perspective correction operations.

9. The method of claim 7, wherein the first set of perspective correction operations and the second set of perspective correction operations include at least one overlapping perspective correction operation.

10. The method of claim 9, wherein the at least one overlapping perspective correction operation includes an adjustable parameter having a first value in the first perspective correction operation set, and wherein the at least one overlapping perspective correction operation includes the adjustable parameter having a second value different from the first value in the second perspective correction operation set.

11. A method according to any one of claims 7 to 10, wherein the change from the first context information to the second context information corresponds to a transition from the first application executed by the computing system to a second application different from the first application executed by the computing system.

12. A method according to any one of claims 7 to 10, wherein the change from the first context information to the second context information corresponds to a change in the current state of the user associated with a transition from a first user motion state to a second user motion state different from the first user motion state.

13. A method according to any one of claims 7 to 10, wherein the change from the first context information to the second context information corresponds to a change in the current state of the physical environment associated with a transition from a first ambient light level detected within the physical environment to a second ambient light level different from the first ambient light level detected within the physical environment.

14. A method according to any one of claims 7 to 10, wherein the change from the first context information to the second context information corresponds to a change in the current state of the physical environment associated with a transition from a first background associated with a first frequency value within the physical environment to a second background associated with a second frequency value different from the first frequency value within the physical environment.

15. The method according to any one of claims 7 to 14, further comprising: An animation associated with the first corrected image data is caused to be presented via the display device prior to presenting the second corrected image data.

16. The method according to any one of claims 7 to 14, further comprising: Presentation of the first corrected image data is stopped before presentation of the second corrected image data.

17. The method according to any one of claims 7 to 16, wherein detecting the change from the first context information to the second context information comprises: Storage counter; increasing or decreasing the counter based on at least one of the second user information, the second application information, and the second environment information; as well as The change from the first context information to the second context information is detected in response to the counter breaking a threshold.

18. The method of claim 17, wherein increasing or decreasing the counter comprises: incrementing the counter based on determining that the second environment information indicates a strong depth gradient of the current environment; as well as The counter is decremented based on determining that the second environment information indicates a weak depth gradient of the current environment.

19. The method of claim 18, wherein increasing or decreasing the counter comprises: incrementing the counter based on determining that the second environment information indicates a medium depth gradient of the current environment and the current environment includes at least one object of a class of objects; as well as The counter is decremented based on determining that the second environment information indicates a medium depth gradient of the current environment and the current environment does not include any object of the class of objects.

20. The method according to any one of claims 7 to 19, wherein presenting the second corrected image data occurs during a transition period after presenting the first corrected image data, the method further comprising, at an intermediate time within the transition period: selecting an intermediate perspective correction operation set, the intermediate perspective correction operation set being a percentage between the first perspective correction operation set and the second perspective correction operation set; generating intermediate corrected image data by performing a set of intermediate perspective correction operations on the uncorrected image stream; and The intermediate corrected image data is caused to be presented via the display device. The method of claim 20 , wherein the percentage is based on the second user information indicating a change in a head posture of the user.

22. The method of claim 20 or 21, wherein the percentage is based on the second application information indicating a distraction level of the second application.

23. The method of any one of claims 1 to 22, wherein the computing system is further communicatively coupled to one or more depth sensors via the communication interface, and the method further comprises: Depth information associated with the physical environment is acquired via the one or more depth sensors, wherein the first corrected image data is generated by performing the first set of perspective correction operations on the image data based on the depth information.

24. The method according to claim 23, further comprising: One or more confidence values ​​for the depth information are determined, wherein the first set of perspective correction operations is determined based at least in part on the first context information or the one or more confidence values ​​for the depth information.

25. A computing system, the computing system comprising: one or more processors; non-transitory memory; an interface for communicating with a display device and one or more image sensors; and One or more programs, the one or more programs being stored in the non-volatile memory, the one or more programs, when executed by the one or more processors, causing the computing system to perform any of the methods of claims 1 to 24.

26. A non-volatile memory storing one or more programs that, when executed by one or more processors of a computing system having an interface for communicating with a display device and one or more image sensors, cause the computing system to perform any of the methods described in claims 1 to 24.

27. A computing system, the computing system comprising: one or more processors; non-transitory memory; an interface for communicating with a display device, one or more audio output devices, and one or more image sensors; and Means for causing the computing system to perform any of the methods of claims 1 to 24.