Techniques for displaying and capturing images
By using a camera to capture image streams in an extended reality system and applying transformation operations to generate a set of transformed images, the problem of image inaccuracy caused by viewpoint differences is solved, thus improving the user experience.
Patent Information
- Application Number
- CN202410678001.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-10
- Filing Date
- 2024-05-29
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-05-29
AI Technical Summary
In extended reality systems, the viewpoint difference between the user and the camera is not corrected, causing the image displayed in pass-through mode to not accurately reflect the user's physical environment, thus affecting the user experience.
A set of cameras captures an image stream, a first transformation operation is applied to generate a transformed image set, which is then displayed on a monitor. Upon receiving a capture request, a stereo output image set is generated. The image set is then processed through different second transformation operations to generate a stereo output image.
By correcting viewpoint discrepancies, the accuracy of image display is improved, user discomfort is reduced, and the user experience of extended reality systems is enhanced.
Smart Images

Figure CN119071651B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application is a non-provisional of U.S. Provisional Patent Application No. 63 / 470,081, filed May 31, 2023, and claims the benefit of the U.S. Provisional Patent Application under 35 U.S.C. 119(e), the contents of which are incorporated herein by reference as if fully disclosed herein. TECHNICAL FIELD
[0003] The described implementations generally relate to separately processing image streams for real-time display and media capture events. BACKGROUND
[0004] Extended reality systems can be used to generate a partially or fully simulated environment (e.g., a virtual reality environment, a mixed reality environment, etc.) in which virtual content can replace or augment the physical world. The simulated environment can provide a user with an engaging experience and be used in gaming, personal communication, virtual travel, healthcare, and many other contexts. In some cases, the simulated environment can include information captured from the user’s environment. For example, a pass-through mode of an extended reality system can use one or more displays of the extended reality system to display images of the user’s physical environment. This can allow the user to perceive their physical environment via the displays of the extended reality system.
[0005] Images of the user’s environment can be captured using one or more cameras of the extended reality system. Depending on the positioning of the cameras in the extended reality system, a difference in viewpoint between the user and the cameras, if not corrected, can cause the images displayed in pass-through mode to not accurately reflect the user’s physical environment. This can cause the user to be disoriented or can negatively impact the user experience during pass-through mode. SUMMARY
[0006] The implementations described herein relate to systems, devices, and methods for performing separate image processing. Some implementations relate to a method comprising: capturing, using a set of cameras of a device, a stream of images; generating, from the stream of images, a first set of transformed images by applying a first transform operation to images in the stream of images; and displaying, at a set of displays of the device, the first set of transformed images. The method further comprises: receiving a capture request; and in response to receiving the capture request, generating a set of stereoscopic output images. Generating the set of stereoscopic output images comprises: selecting a set of images from the stream of images; generating, from the set of images, a second set of transformed images by applying a second transform operation to the set of images that is different from the first transform operation; and using the second set of transformed images to generate the set of stereoscopic output images.
[0007] In some variations of these methods, generating the set of stereoscopic output images includes generating a fused stereoscopic output image from two or more images of the set of images. Additionally or alternatively, the set of stereoscopic output images comprises a stereoscopic video. In some of these variations, generating the set of stereoscopic output images includes performing a video stabilization operation on the second set of transformed images.
[0008] In some cases, the first transformation operation is a perspective correction operation. Additionally or alternatively, the second transformation operation can be an image rectification operation. Generating the set of stereoscopic output images can include generating metadata associated with the set of stereoscopic output images. The metadata can include field of view information for the set of cameras, pose information for the set of stereoscopic output images, and / or a set of default disparity values for the set of stereoscopic output images. In some cases, generating the metadata associated with the set of stereoscopic output images includes selecting the set of default disparity values based on a scene captured by the set of stereoscopic output images. In some variations, virtual content is added to the first set of transformed images.
[0009] Some embodiments relate to a device comprising a set of cameras, a set of displays, a memory, and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions that cause the one or more processors to perform any of the above methods. Similarly, other embodiments relate to a non-transitory computer readable medium comprising instructions that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising the steps of any of the above methods.
[0010] Other embodiments relate to a method comprising capturing, using a first camera and a second camera of a device, a first set of image pairs of a scene. Each image pair of the first set of image pairs includes a first image captured by the first camera and a second image captured by the second camera. The method includes generating, from the first set of image pairs, a first set of transformed image pairs, wherein generating the first set of transformed image pairs includes, for each image pair of the first set of image pairs, generating a first perspective corrected image from the first image based on a difference between a viewpoint of the first camera and a viewpoint of a user and generating a second perspective corrected image from the first image based on a difference between a viewpoint of the second camera and the viewpoint of the user. The method further includes displaying, on a set of displays of the device, the first set of transformed image pairs.
[0011] Additionally, the method includes receiving a capture request, selecting a second set of image pairs from the first set of image pairs in response to receiving the capture request, and generating a second set of warped image pairs from the second set of image pairs. This includes, for each image pair in the first set of image pairs: generating a first de-warped image from the first image; generating a second de-warped image from the second first image; and aligning the first de-warped image and the second de-warped image. The method includes generating a set of output images using the second set of warped image pairs.
[0012] In some variations of these methods, generating the set of output images includes generating a fused stereoscopic output image from two or more image pairs in the set of warped image pairs. Additionally or alternatively, the set of stereoscopic output images includes a stereoscopic video formed from the second set of warped image pairs. In some of these variations, generating the set of stereoscopic output images includes performing a video stabilization operation on the second set of warped image pairs.
[0013] Generating the set of stereoscopic output images can include generating metadata associated with the set of stereoscopic output images. The metadata can include field of view information for at least one of the first camera or the second camera, pose information for the set of output images, and / or a set of default disparity values for a set of stereoscopic images in the stereoscopic output images. In some cases, generating the metadata associated with the set of stereoscopic output images includes selecting the set of default disparity values based on a scene captured by the set of output images. In some variations, virtual content is added to the first set of warped image pairs.
[0014] Some embodiments relate to a device including a first camera, a second camera, a set of displays, a memory, and one or more processors operatively coupled to the memory, where the one or more processors are configured to execute instructions that cause the one or more processors to perform any of the above methods. Similarly, other embodiments relate to a non-transitory computer readable medium including instructions that, when executed by at least one computing device, cause the at least one computing device to perform operations including the steps of any of the above methods.
[0015] In addition to the exemplary aspects and embodiments described above, further aspects and embodiments will become apparent from the following description and the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0016] The present disclosure will be readily understood by those skilled in the art from the following detailed description in conjunction with the accompanying drawings, wherein like reference numerals designate like structures throughout the drawings:
[0017] Figure 1 A block diagram of an exemplary electronic device that can be used in the extended reality systems described herein is shown.
[0018] Figure 2 An example operating environment including variations of the electronic device as described herein is shown.
[0019] Figure 3 An example of an image stream is shown.
[0020] Figure 4 A process is depicted that uses a display image processing pipeline and an output image processing pipeline to separately process images in an image stream.
[0021] Figure 5 An example of a display image processing pipeline as described herein is depicted.
[0022] Figure 6 An example of an output image processing pipeline as described herein is depicted.
[0023] Figure 7 An example of an image generation unit that can be used as part of an output image processing pipeline described herein is depicted.
[0024] It is to be understood that the ratios and dimensions (relative or absolute) of the various features and elements (and collections and groupings thereof), as well as the relationships between the boundaries, spacings, and positions thereof, presented in the drawings are provided merely to facilitate an understanding of the various embodiments described herein and as such can not necessarily be presented or shown as drawn to scale and are not intended to indicate any preference or requirement as to the illustrated embodiments in order to exclude embodiments incorporating techniques described in connection therewith. DETAILED DESCRIPTION
[0025] Reference will now be made in detail to representative embodiments illustrated in the accompanying drawings. It should be understood that the following description is not intended to limit the embodiments to one preferred embodiment. On the contrary, it is intended to cover alternatives, modifications, and equivalents that can be included within the spirit and scope of the described embodiments as defined by the appended claims.
[0026] The embodiments disclosed herein relate to devices, systems, and methods for separately processing an image stream to generate a first set of images for display on a set of displays and a second set of images that are stored or transmitted as part of a media capture event. Specifically, a first set of images from the image stream can be used to generate a first set of transformed images, and these transformed images can be displayed on a set of displays. A second set of images can be selected from the image stream in response to a capture request associated with a media capture event, and the second set of images can be used to generate a second set of transformed images. A set of output images can be generated from the second set of transformed images, and the set of output images can be stored or transmitted for later viewing.
[0027] Reference is made below to Figures 1 to 7These and other implementations will be discussed here. However, those skilled in the art will readily understand that the detailed descriptions of these figures given herein are for illustrative purposes only and should not be construed as limiting.
[0028] The devices and methods described herein can be used as part of an extended reality system in which an extended reality environment is generated and displayed to a user. Various terms are used herein to describe the various extended reality systems and associated extended reality environments described herein. For example, as used herein, a “physical environment” is part of the physical / real world surrounding a user, which the user can perceive and interact with without the assistance of the extended reality system described herein. For example, a physical environment may include a room or outdoor space within a building, and any person, animal, or object (collectively referred to herein as a “real-world object”), such as plants, furniture, books, etc.
[0029] As used herein, an "extended reality environment" refers to a fully or partially simulated environment that a user can perceive and / or interact with using an extended reality system as described herein. In some cases, an extended reality environment can be a virtual reality environment, which is a fully simulated environment in which the user's physical environment is completely replaced by virtual content within the virtual reality environment. Virtual reality environments may be independent of the user's physical environment and therefore allow the user to perceive themselves in different simulated locations (e.g., standing on a beach when they are actually standing in a room of a building). Virtual reality environments may include virtual objects that the user can interact with (e.g., simulated objects that can be perceived by the user but do not actually exist in the physical environment).
[0030] In other cases, an extended reality environment can be a mixed reality environment, that is, a fully or partially simulated environment in which virtual content can be presented along with a portion of the user's physical environment. Specifically, a mixed reality environment may include a reproduction and / or modified representation of one or more portions of the user's physical environment surrounding the extended reality system. In this way, the user can perceive their physical environment (directly or indirectly) through the mixed reality environment while still perceiving the virtual content.
[0031] As used herein, a “recreation” of a portion of a physical environment refers to a portion of an extended reality environment that recreates, within the extended reality environment, the portion of the physical environment. For example, an extended reality system can include one or more cameras that are capable of capturing images of the physical environment. The extended reality system can present a portion of the images to the user by displaying the images via an opaque display such that the user views the physical environment indirectly via the displayed images. Additionally, in some cases, the extended reality environment is displayed using foveated rendering in which different levels of fidelity (e.g., image resolution) are used to render different portions of the extended reality environment depending on the user’s gaze direction. In these cases, for the purposes of this application, a recreated portion that is rendered using these foveated rendering techniques at a lower fidelity is still considered to be a recreation.
[0032] As used herein, a “modified representation” of a portion of a physical environment refers to a portion of an extended reality environment that is derived from the physical environment but intentionally obscures one or more aspects of the physical environment. Whereas indirect recreation attempts to replicate a portion of a user’s physical environment within the extended reality environment, a modified representation intentionally alters one or more visual aspects of a portion of the user’s physical environment (e.g., using one or more visual effects such as artificial blurring). In this way, a modified representation of a portion of the user’s physical environment can allow the user to perceive certain aspects of the portion of the physical environment while obscuring other aspects. In the example of artificial blurring, the user can still be able to perceive the general shape and placement of real-world objects within the modified representation, but can not be able to perceive visual details of those objects that are visible in the physical environment. In cases where the extended reality environment is displayed using foveated rendering, portions of the modified representation in peripheral regions of the extended reality environment (relative to the user’s gaze) can be rendered using foveated rendering techniques at a lower fidelity.
[0033] A recreation and / or modified representation of a user’s physical environment can be used in a variety of extended reality environments. For example, an extended reality system can be configured to operate in a “pass-through” mode during which the extended reality system generates and presents an extended reality environment that includes a recreation of a portion of the user’s physical environment. This allows the user to indirectly view a portion of their physical environment even though one or more components of the extended reality system (e.g., a display) can interfere with the user’s ability to directly view the same portion of their physical environment. When the extended reality system is operating in pass-through mode, some or all of the extended reality environment can include a recreation of the user’s physical environment. In some cases, in addition to the recreation of the user’s physical environment, one or more portions of the extended reality environment can include virtual content and / or a modified representation of the user’s physical environment. Additionally or alternatively, virtual content (e.g., graphical elements of a graphical user interface, virtual objects) can be overlaid on the recreated portion.
[0034] Generally, the extended reality systems described herein include electronic devices that are capable of capturing images and displaying images to a user. Figure 1 A block diagram depicting an example electronic device 100, which can be part of an extended reality system as described herein, is depicted. The electronic device 100 can be configured to capture, process, and display images that are part of an extended reality environment. In some implementations, the electronic device 100 is a handheld electronic device (e.g., a smartphone or a tablet) that is configured to present an extended reality environment to a user. In some of these implementations, the handheld electronic device can be temporarily attached to an accessory that allows the user to wear the handheld electronic device. In other implementations, the electronic device 100 is a head-mounted device (HMD) that is worn by the user.
[0035] In some embodiments, the electronic device 100 has a bus 102 that operatively couples an I / O section 104 with one or more computer processors 106 and memory 108, and includes circuitry that interconnects and controls communications between the components of the electronic device 100. The I / O section 104 includes various system components that can aid in the operation of the electronic device 100. The electronic device 100 includes a set of displays 110 and a set of cameras 112. The set of cameras 112 can capture images of a user’s physical environment (e.g., a scene), and these images can be used when generating an extended reality environment as described herein (e.g., as part of a pass-through mode). The set of displays 110 can display the extended reality environment so that the user can view the extended reality environment, allowing the user to perceive their physical environment via the electronic device 100. The set of displays 110 can include a single display, or can include multiple displays (e.g., one display for each eye of the user). Each display in the set of displays can utilize any suitable display technology, and can include, for example, a liquid crystal display (LCD), an organic light-emitting diode (OLED) display, a light-emitting diode (LED) display, a quantum dot light-emitting diode (QLED) display, etc.
[0036] The memory 108 of the electronic device 100 can include one or more non-transitory computer-readable storage devices. These non-transitory computer-readable storage devices can be used to store computer-executable instructions that, when executed by the one or more computer processors 106, can cause the computer processors to perform the processes described herein (the various image capture and processing techniques described below). Additionally, the non-transitory computer-readable storage devices can be used to store images that are captured as part of a media capture event as described herein.
[0037] A computer-readable storage device can be any medium that can tangibly contain or store computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device, such as one or more processors 106. In some examples, the storage device is a transitory computer-readable storage medium. In some examples, the storage device is a non-transitory computer-readable storage medium. Non-transitory computer-readable storage devices can include, but are not limited to, magnetic, optical, and / or semiconductor storages. Examples of such storage include magnetic disks, optical discs based on CD, DVD, or Blu-ray technologies, and persistent solid-state memory such as flash, solid-state drives, and the like.
[0038] The one or more processors 106 can include, for example, a processor, microprocessor, programmable logic array (PLA), programmable array logic (PAL), generic array logic (GAL), complex programmable logic device (CPLD), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or any other programmable logic device (PLD) that can be configured to execute the operating system and applications of the electronic device 100 and facilitate the capture, processing, display, and storage of images, as described herein.
[0039] Accordingly, any of the processes described herein can be stored as instructions on a non-transitory computer-readable storage device, such that a processor can utilize the instructions to perform various steps of the processes described herein. Similarly, the devices described herein include a memory (e.g., the memory 108) and one or more processors (e.g., the processors 106) operatively coupled to the memory. The one or more processors can receive instructions from the memory and be configured to execute the instructions to perform various steps of the processes described herein. Any of the processes described herein can be performed using the devices described herein as a method of capturing and displaying images.
[0040] In some implementations, the electronic device 100 includes a set of depth sensors 114 (which can include a single depth sensor or multiple depth sensors), each depth sensor of the set of depth sensors configured to compute depth information for a portion of an environment in front of the electronic device 100. In particular, each depth sensor of the set of depth sensors 114 can compute depth information within a coverage field (e.g., a widest lateral range for which the depth sensor is able to provide depth information). In some cases, the coverage field of one depth sensor of the set of depth sensors 114 can at least partially overlap with a field of view (e.g., a spatial range of a scene that a camera is able to capture using an image sensor of the camera) of at least one camera of the set of cameras 112, thus allowing the depth sensor to compute depth information associated with the field of view of the set of cameras 112.
[0041] Information from the depth sensors can be used to calculate distances between the depth sensors and various points in the environment surrounding the electronic device 100. In some cases, depth information from the set of depth sensors 114 can be used to generate a depth map, as described herein. Depth information can be calculated in any suitable manner. In one non-limiting example, a depth sensor can utilize stereo imaging, in which two images are taken from different positions, and the distance (disparity) between corresponding pixels in the two images can be used to calculate depth information. In another example, a depth sensor can utilize structured light imaging, whereby the depth sensor can image a scene while projecting a known illumination pattern (typically using infrared illumination) toward the scene, and then can see how the pattern is distorted by the scene to calculate depth information. In yet another example, a depth sensor can utilize time-of-flight sensing, which calculates depth based on the amount of time it takes for (typically infrared) light emitted from the depth sensor to return from the scene. Time-of-flight depth sensors can utilize direct time-of-flight or indirect time-of-flight, and can illuminate an entire field of coverage at a given time, or can illuminate only a subset of the field of coverage at a given time (e.g., via one or more spots, stripes, or other patterns that can be fixed or can be scanned across the field of coverage). In cases where the depth sensors utilize infrared illumination, the infrared illumination can be used for a range of environmental conditions without being perceived by a user.
[0042] In some implementations, the electronic device 100 includes an eye tracker 116. The eye tracker 116 can be configured to determine the position of a user’s eyes relative to the electronic device 100 (or particular components thereof). The eye tracker 116 can include any suitable hardware for recognizing and locating a user’s eyes, such as one or more cameras, depth sensors, combinations thereof, and the like. In some cases, the eye tracker 116 can also be configured to detect the position / direction of a user’s gaze. It will be appreciated that the eye tracker 116 can include a single module configured to determine the position of both of a user’s eyes, or can include multiple units, each configured to determine the position of a corresponding eye of the user.
[0043] Additionally or alternatively, the electronic device 100 can include a set of sensors 118 capable of determining the motion and / or orientation of the electronic device 100 (or particular components thereof). For example, Figure 1 The illustrated set of sensors 118 includes an accelerometer 120 and a gyroscope 122. Information from the set of sensors 118 can be used to calculate pose information for the electronic device 100 (or particular components thereof). This can allow the electronic device 100 to track its position and orientation as it moves. Additionally or alternatively, information captured by the set of cameras 112 and / or the set of depth sensors 114 can be used to calculate pose information for the electronic device 100 (or particular components thereof), e.g., by using simultaneous localization and mapping (SLAM) techniques.
[0044] Additionally, the electronic device 100 can include a set of input mechanisms 124 (e.g., a touch screen, soft keys, a keyboard, a virtual keyboard, buttons, knobs, a joystick, switches, a dial, combinations thereof, and the like) that a user can manipulate to interact with and provide input to the extended reality system. The electronic device 100 can include a communication unit 126 to receive application and operating system data using Wi-Fi, Bluetooth, near-field communication (NFC), cellular, and / or other wireless communication techniques. The electronic device 100 is not limited to Figure 1 the components and configurations of the electronic device 100, but can include other or additional components in multiple configurations according to need.
[0045] Figure 2 An operating environment 200 is shown that includes an electronic device 202, which can be part of an extended reality system to perform the processes described herein. The electronic device 202 can be configured in any manner described above with respect to the electronic device 100. Figure 1 For example, the electronic device 202 can include a set of cameras including a first camera 112a and a second camera 112b, and a set of displays including a first display 110a and a second display 110b. Additionally, the electronic device 202 is shown to include a depth sensor 114 and an eye tracker 116, such as described herein.
[0046] The first camera 112a has a first field of view 230a and is positioned to capture images of a first portion of the physical environment surrounding the electronic device 202. Similarly, the second camera 112b has a second field of view 230b and is positioned to capture images of a second portion of the physical environment surrounding the electronic device 202. The first field of view 230a and the second field of view 230b can at least partially overlap, such that the first camera 112a and the second camera 112b can simultaneously capture a common portion of the physical environment surrounding the electronic device 202. For example, Figure 2 The operating environment 200 shown in FIG. 2 includes a real-world object 206 that is positioned in both the first field of view 230a and the second field of view 230b. Thus, when the electronic device 202 is positioned as shown in FIG. 2, the images captured by both the first camera 112a and the second camera 112b will include the real-world object 206. Figure 2
[0047] As described in greater detail herein, images captured by the first and second cameras 112a, 112b can be displayed via the set of displays as part of an extended reality environment to provide a user with a representation of their physical environment (e.g., as part of a pass-through mode). For example, images captured by the first camera 112a can be displayed on the first display 110a (e.g., as part of an extended reality environment) and viewable by the user’s first eye 204a, such as when the electronic device 202 is worn on or otherwise held near the user’s head. Similarly, images captured by the second camera 112b can be displayed on the second display 110b (e.g., as part of an extended reality environment) and viewable by the user’s second eye 204b.
[0048] Depending on the relative positioning between components of the electronic device 202 and the relative positioning between the electronic device 202 and the user’s eyes 204a, 204b, the images captured by the first and second cameras 112a, 112b can provide a viewpoint of the operating environment 200 that is different from the case where the user directly views their physical environment without the electronic device 202. In particular, the set of cameras can be spatially offset from the set of displays in one or more directions. For example, the imaging plane of the first camera 112a can be separated from the first display 110a by a distance 220a along a first (e.g., horizontal) direction, and the imaging plane of the second camera 112b can be separated from the second display 110b by a distance 220b along the first direction. In some cases, the imaging planes of the first and second cameras 112a, 112b can also be rotated relative to the first and second displays 110a, 110b, respectively.
[0049] Additionally, the first and second eyes 204a, 204b can be separated from the first and second displays 110a, 110b by a corresponding distance (e.g., a first eye-to-display distance 222a between the first display 110a and the first eye 204a and a second eye-to-display distance 222b between the second display 110b and the second eye 204b). The first camera 112a is separated from the first display 110a by a first reference point (denoted by the dashed line in FIG. 1) and the second camera 112b is separated from the second display 110b by a second reference point (denoted by the dashed line in FIG. 1). Figure 2Object 206 in the diagram represents the first scene distance 232a, and the second camera 112b is separated from the second scene distance 232b by a second reference point (which may be the same point as or different from the first reference point). The offset (spatial and / or rotation) between the first camera 112a and the first display 110a, the distance 222a from the first eye to the display, and the first scene distance 232a can collectively characterize the difference between the viewpoint of the first eye 204a and the viewpoint of the first camera 112a relative to the first reference point. Similarly, the offset between the second camera 112b and the second display 110b, the distance 222b from the second eye to the display, and the second scene distance 232b can collectively characterize the difference between the viewpoint of the second eye 204b and the viewpoint of the second camera 112b relative to the second reference point.
[0050] Therefore, it might be desirable to adjust images captured by a set of cameras to account for these viewpoint differences before these images are displayed (e.g., to generate perspective-corrected images), which could reduce user discomfort when viewing these images. However, such viewpoint adjustments may not be desirable when a user wishes to use a set of cameras to capture images or videos (e.g., to save for later viewing or to share with others). These captured images may later be viewed by different devices (which may have a different configuration than device 202), viewed from different viewpoints (e.g., viewed by different users or when the user is at a different distance from the device), and / or viewed in a specific manner within an extended reality environment (e.g., within a viewing window or range within the extended reality environment). Presenting perspective-corrected images in these cases can lead to a poor user experience, and further adjusting these images to fit the new viewing environment can undesirably introduce or amplify artifacts.
[0051] The process described herein includes obtaining data from one or more cameras (e.g., ...). Figure 1 and Figure 2 Electronic devices 100 and 202 (one or more cameras in a set of cameras) acquire a collection of images and process these images using multiple separate image processing pipelines. Specifically, a first image processing pipeline generates images that can be displayed on a set of displays, while a second image processing pipeline generates a collection of output images that can be stored for later viewing or shared with others. The first and second image processing pipelines can receive the same images as input but will utilize different image processing techniques tailored to their respective outputs.
[0052] At least some of the images in the set of images are captured during a media capture event. Specifically, a set of cameras can be used to capture images during one or more photography modes (e.g., a photo mode that can capture still images, a video mode that can capture video, a panorama mode that can capture panoramic photos, a portrait mode that can capture still photos with applied bokeh, etc.). Generally, during these modes, the device can display (e.g., via a set of displays of the electronic device 100, 202) a camera user interface that displays a “live preview.” The live preview is a stream of images captured by the set of cameras and presented in real-time, and can represent a field of view (which can be a subset of the field of view of the cameras) that will be captured when the camera initiates a media capture event. In other words, the live preview allows the user to see which portion of the scene is currently being imaged and to decide when to capture a photo or video. Figure 1 and Figure 2 The live preview can be presented to the user as part of an extended reality environment, where the extended reality environment is operating in pass-through mode. In these cases, at least a portion of the extended reality environment includes a reproduction of the user’s physical environment, thereby allowing the user to perceive their physical environment and to understand which portion of the scene will be captured during a media capture event. The extended reality environment can include one or more graphical user elements or other features (e.g., a transition, such as a transition between the reproduction of the physical environment and virtual content or a modified representation of the physical environment, that indicates the boundaries of the image to be captured by the device) to help guide the user in capturing the image.
[0053] The live preview can be presented to the user as part of an extended reality environment, where the extended reality environment is operating in pass-through mode. In these cases, at least a portion of the extended reality environment includes a reproduction of the user’s physical environment, thereby allowing the user to perceive their physical environment and to understand which portion of the scene will be captured during a media capture event. The extended reality environment can include one or more graphical user elements or other features (e.g., a transition, such as a transition between the reproduction of the physical environment and virtual content or a modified representation of the physical environment, that indicates the boundaries of the image to be captured by the device) to help guide the user in capturing the image.
[0054] When the extended reality system initiates a media capture event, the set of cameras will capture images and generate media (e.g., capture and generate still photos when in photo mode, or capture and generate video when in video mode) depending on the current photography mode, which can then be stored locally on the device 100 (e.g., in a non-transitory computer-readable storage device that is part of the memory 108), transmitted to a remote server for storage, or transmitted to another system (e.g., as part of a messaging or streaming application). It should be understood that in some cases, there can be a frame buffer of images captured by the set of cameras, and some images collected prior to the initiation of the media capture event can be used to generate the captured media that is stored or transmitted.
[0055] Media capture events can be initiated when an electronic device, as described herein, receives a capture request. A capture request can be received under certain predetermined conditions (e.g., a software application running on the device can automatically request a group of cameras to initiate a media capture event when certain criteria are met, with appropriate user permission), or when a user gives a command to initiate a media capture event by interacting with shutter controls on a user interface, pressing a designated button on the electronic device, giving a voice command, etc.
[0056] The process described herein can be performed as part of a media capture event, and can utilize a first image processing pipeline to generate a live preview image displayed to the user, and utilize a second image processing pipeline to generate the captured media to be stored or transmitted. For example, Figure 3 An exemplary image stream 300 is shown, comprising a sequence of images of a scene, the image sequence being generated by a set of cameras of a device (e.g., Figure 1 and Figure 2 The image stream 300 is captured by a set of cameras (electronic devices 100, 202). In an example where the set of cameras includes two cameras, the image stream 300 includes images captured by the first camera (e.g., ...). Figure 2 The first camera image stream 300a of the electronic device 202 (first camera 112a) and the second camera (e.g., Figure 2 The second camera image stream 300b of the electronic device 202 (second camera 112b) captures the image 304.
[0057] A set of cameras may continuously capture an image stream 300 during operation of an electronic device, or may initiate the capture of an image stream 300 in response to certain conditions (e.g., the device entering pass-through mode). When the image stream 300 includes multiple camera image streams (e.g., a first camera image stream 300a and a second camera image stream 300b), images from each camera stream (such as images captured simultaneously from different cameras) may be grouped and processed together. For example, the image stream 300 may include a series of image pairs, wherein each image pair includes a first image captured by a first camera (e.g., an image from the first camera image stream 300a) and a second image captured by a second camera (e.g., an image from the second camera image stream 300b).
[0058] For example, Figure 3Image stream 300 is shown to include a total of sixteen images (eight images 302 labeled LI - L8 in first camera image stream 300a and eight images 304 labeled Rl - R8 in second camera image stream 300b), and thus will have eight image pairs (the first pair of images includes images LI and Rl, the second pair of images includes images L2 and R2, and so on). The set of cameras can continue to capture image stream 300 for as long as any media capture event that can be desired needs to be performed. A predetermined number of images (or image pairs) from image stream 300 can be temporarily stored in a buffer at a given time (e.g., with the oldest images in the buffer being replaced with new images), allowing a particular number of most recently captured images to be available to the device at any given time during the capture of image stream 300.
[0059] The processes described herein can select a first set of images (or image pairs) 306 from image stream 300, and will use the first set of images 306 to generate a set of live preview images, as described in greater detail below. It should be understood that when the processes as described herein are used to select a set of images for additional image processing (e.g., to generate images therefrom), the processes need not wait for all images to be selected before initiating the image processing. Using the first set of images 306 as a non-limiting example, the device can begin generating live preview images from images LI and Rl before other images from the first set 306 have been captured, such as images L7, L8, R7, and R8. In other words, the first image processing pipeline can continuously generate live preview images.
[0060] As Figure 3As shown, a capture request 308 can be received during capture of the image stream (e.g., as shown here during capture of images L4 and R4). In response to receiving the capture request 308, a second set of images 310 can be selected as part of the media capture event. The second set of images 310 can be used (e.g., by a second image processing pipeline) to generate a set of output images that form captured media that is stored or transmitted as discussed herein. In some cases, as previously discussed, the second set of images 310 can include one or more images or image pairs (e.g., images L3 and R3) that were captured prior to the capture request 308. The number of images 310 selected to be in the second set 310 can be predetermined (e.g., certain photography modes can select a particular number of images associated with each capture request 308) or dynamically determined (e.g., in a video mode, the number of images in the second set 310 can depend on when a termination request is received that causes the media capture event to stop, allowing the user to capture videos of different lengths). In some variations, receipt of the capture request 308 can optionally prompt the set of cameras to temporarily change one or more operational parameters while capturing images. For example, in response to receiving the capture request 308, the set of cameras can change the frame rate at which images are captured, or can intentionally change the exposure time of one or more images (e.g., the image pair including images L5 and R5 are shown as being captured with a longer exposure time than surrounding image pairs).
[0061] Some or all of the images used to generate output images as part of the media capture event (e.g., the second set of images 310) can also be used to generate real-time preview images that are displayed to the user (e.g., via the first set of images 306). Thus, different image processing pipelines can receive the same images as input, but will use different image processing techniques to generate multiple different images. For example, Figure 4 A variation of the process 400 is depicted that uses multiple image processing pipelines to generate a first set of transformed images for display to a user (e.g., as part of a real-time preview in an extended reality environment) and a second set of transformed images for generating one or more output images as part of a media capture event. The blocks performed as part of the process 400 can be performed as a method, or can be stored as instructions on a non-transitory computer-readable storage device, such that a processor can utilize the instructions to perform the various blocks of the process described herein. Similarly, a device described herein includes a memory (e.g., the memory 108) and one or more processors (e.g., the processor 106) operatively coupled to the memory, where the one or more processors are configured to execute instructions that cause the one or more processors to perform the blocks of the process 400.
[0062] In particular, the process 400 uses a set of cameras 404 to capture an image stream 402. In some cases, the set of cameras 404 includes a first camera 404a and a second camera 404b, and the image stream 402 includes images 402a captured by the first camera 404a and images 402b captured by the second camera 404b. In these cases, the image stream 402 can form a series of image pairs, where each image pair includes a first image captured by the first camera 404a and a second image captured by the second camera 404b.
[0063] The process 400 includes a display image processing pipeline 406 configured to generate a first set of transformed images 408 from images in the image stream 402, such as, for example, the first set of images 306 from the image stream 300 of FIG. 3. Figure 3 As part of generating the first set of transformed images 408, the display image processing pipeline 406 can apply a first transform operation to the images in the image stream 402. In some cases, this first transform operation includes a perspective correction operation in which the images are transformed based on a difference between a viewpoint of one of the set of cameras 404 and a viewpoint of the user. In these cases, the display image processing pipeline 406 can receive and utilize viewpoint data 410 representing information about the viewpoints of the set of cameras 404 and the user. Perspective correction operations will be described in greater detail herein with respect to FIG. 4B. Figure 5
[0064] For each image in the image stream 402 input to the display image processing pipeline 406, the display image processing pipeline 406 can generate a corresponding transformed image in the first set of transformed images 408. For example, images 402a captured by the first camera 404a can be used to generate a first subset of these transformed images 408a, and images 402b captured by the second camera 404b can be used to generate a second subset of transformed images 408a. Thus, the first set of transformed images 408 can include a first set of pairs of transformed images, where each pair of images in the image stream 402 is used to generate a corresponding pair of transformed images in the first set of pairs of transformed images.
[0065] At step 412, the first set of transformed images is displayed on a set of displays (e.g., the set of displays 110 of the electronic device 100 of FIG. 1). Figure 1 The first set of transformed images can be displayed as part of an extended reality environment (e.g., in pass-through mode), which can provide a real-time preview such as discussed herein. In cases where the set of displays includes a first display and a second display (such as, for example, the first display 110a and the second display 110b of the electronic device 100 of FIG. 1), the first set of transformed images can be displayed on the first display and the second display. Figure 2 In the case where the set of displays includes a first display 110a and a second display 110b of the electronic device 202, the first subset 408a of transformed images 408 can be displayed on the first display, and the second subset 408b of transformed images 408 can be displayed on the second display. In this way, the first eye of the user can view images captured by the first camera (via the first display), and the second eye of the user can view images captured by the second camera (via the second display). In the case where the set of displays includes a single display, each pair of transformed images can be displayed on the single display at the same time. For example, a first transformed image of a pair of images can be displayed on a first portion of the display (where the image can be viewed by the first eye of the user), while a second transformed image of the pair of images can be displayed on a second portion of the display (where the image can be viewed by the second eye of the user). In some cases, the display image processing pipeline 406 can add virtual content to the transformed images, such that the virtual content can be viewed by the user when the transformed images are displayed to the user at block 412.
[0066] Additionally, a capture request can be received during the capture of the image stream 402. At block 414, a set of images 416 is selected from the image stream 300 (e.g., from the first set of images 402a and the second set of images 402b) in response to the receipt of the capture request. The set of selected images 416 includes one or more images captured after the receipt of the capture request, and in some cases can also include one or more images captured before the receipt of the capture request, such as described above with respect to block 408. Figure 3 Figure 3 The set of selected images 416 can include a first subset of images 416a captured by the first camera 404a (e.g., selected from the images 402a) and a second subset of images 416b captured by the second camera 404b (e.g., selected from the images 402b). Thus, the set of selected images 416 can include a selected set of pairs of images, where each pair of images includes a first image captured by the first camera 404a and a second image captured by the second camera 404b.
[0067] The process 400 can utilize an output image processing pipeline 418 configured to generate a set of output images 420 using the set of selected images 416. In some cases, the set of output images 420 includes a first subset of output images 420a generated from a first selected subset of images 416a and a second subset of output images 420b generated from a second subset of images 416b. Thus, the set of output images 420 can include a set of output image pairs. These output image pairs can include one or more stereoscopic images and / or stereoscopic videos. In stereoscopic imaging, two images taken from different perspectives can allow a user to perceive depth when presented to different eyes of the user. Thus, a stereoscopic image represents a single pair of images that can later be presented to a user as a three-dimensional still image, while a stereoscopic video represents a series of pairs of images that can later be presented to a user as a three-dimensional video sequentially. Depending on the photography mode, the set of output images 420 can include a single stereoscopic image, multiple stereoscopic images (e.g., generated as part of a burst mode), only stereoscopic videos, or a combination of stereoscopic images and stereoscopic videos. In cases where the set of output images 420 includes both stereoscopic images and stereoscopic videos, the stereoscopic images can be presented to a user as three-dimensional images, but stereoscopic video clips associated with the stereoscopic images can also be played.
[0068] As part of generating the set of output images 420, the output image processing pipeline 418 can generate a second set of transformed images (not shown) from the set of selected images 416. The second set of transformed images can be used to generate the output images 420. In some cases, the output image processing pipeline 418 applies a second transform operation to the images in the image stream 402 that is different from the first transform operation applied by the display image processing pipeline 406. In this way, while the display image processing pipeline and the output image processing pipeline can sometimes receive the same images for processing, they will separately process these input images (e.g., in parallel) to provide different outputs. In some variations, the second transform operation includes an image correction operation, such as described in more detail with respect to Figure 6
[0069] In some cases, the output image processing pipeline 418 is also configured to generate metadata 422 associated with the set of output images 420. This metadata (examples of which will be described in more detail below with respect to Figure 7 the set of output images 420 can be stored together. In addition to the set of selected images 416, the output image processing pipeline 418 can also receive context information 424. The context information 424 can include information about the device (or components thereof, such as a set of cameras) and / or the scene at the time the image stream 404 was captured.
[0070] For example, in some variations, context information 424 includes a depth map 426. The depth map 426 includes a pixel matrix, where each pixel corresponds to a specific location in the scene imaged by a set of cameras, and has a corresponding depth value representing the distance between the device (or a portion thereof, such as a single camera in a set of cameras) and the specific location in the scene. In some cases, the device includes a depth sensor (e.g., Figure 1 and Figure 2 In the case of the depth sensor 114 of the electronic devices 100, 202, the depth map 426 can be derived from the depth information generated by the depth sensor. Alternatively, the depth map 426 can be derived from image analysis of images (or multiple images) captured by one or more cameras of the device (such as one or more images in image stream 402). In these cases, monocular or stereo depth estimation techniques can be used to analyze the images to estimate depth information for different parts of the scene. It should also be understood that the output image processing pipeline 418 can receive multiple depth maps, such as a first depth map of the first camera 404a (which represents the scene distance relative to the first camera 404a) and a second depth map of the second camera 404b (which represents the scene distance relative to the second camera 404b).
[0071] Alternatively or additionally, context information 424 may include information about the content of the scene captured by the selected set of images 416. For example, context information 424 may include a scene information map 428. Specifically, scene information map 428 includes a pixel matrix, where each pixel corresponds to a corresponding location in the scene and has a corresponding identifier indicating the type of object present at that location (and optionally, a confidence value associated with that identifier). For example, scene information map 428 may be configured to identify locations in a scene that include faces, such that each pixel of the scene information map indicates whether a face is present at the corresponding location in the scene, or includes a value (such as a normalized value between 0 and 1) representing the confidence that a face is present at the corresponding location in the scene. In some variations, context information 424 may include multiple scene information maps (each associated with a different type of object), or may include a single scene information map 428 that can indicate the presence of multiple types of objects (e.g., people, animals, trees, etc.) in the scene.
[0072] In particular, the scene information map 428 (or maps) can provide an understanding of the content of the scene, which can be used in generating the set of output images 420 and / or the metadata 422. For example, the output image processing pipeline can be configured to apply different image processing techniques to people in the scene than to other objects present in the scene. Additionally or alternatively, the scene information map 428 (or maps) can be used to select a default parallax value that is included in the metadata 422 and that can later be used (e.g., when displaying the output images) to control the parallax at which the output images are displayed (e.g., in a viewing area within an extended reality environment).
[0073] The scene information map 428 can be generated in any suitable manner. For example, the set of selected images 416 and / or some or all of the depth map 426 can be analyzed, such as using image segmentation and / or image classification techniques, as will be readily appreciated by one of ordinary skill in the art. These techniques can be used to find a set of predetermined objects within the scene, the selection of which can be set based on user preferences, the photography mode that was active when the set of selected images 416 was captured, and / or the desired effect selected by the user that will be applied to generate the set of output images 420. In some variations, the scene information map 428 can be based at least in part on user input. For example, the user can identify a region of interest of the scene (e.g., by determining the user’s gaze when it is related to the scene, by determining the user’s gestures when it is related to the scene, other user input, combinations thereof, etc.), and can identify one or more objects in the region of interest to generate the scene information map 428.
[0074] Additionally or alternatively, the contextual information 424 can include pose information 430 associated with the device (or one or more components thereof). The pose information 430 can include the position and / or orientation of the device during capture of the set of selected images 416. This can be used to generate similar pose information associated with the set of output images 420. The pose information associated with the set of output images 420 can be stored as metadata 422, and can later be used (e.g., when displaying the output images) to control the position and / or orientation at which the output images are displayed (e.g., in an extended reality environment).
[0075] At block 432, the set of output images 420 (including such metadata 422 if the set of output images 420 is associated with metadata 422) can be stored in a non-transitory computer-readable storage device (e.g., as a file, in a database, etc.). The set of output images 420 can be stored in association with the metadata 422, and the metadata 422 can be stored in association with the set of selected images 416 and / or the depth map 426. Figure 1The set of output images 420 can be stored in the memory 108 of the electronic device 100 (e.g., as part of the memory 108 of the electronic device 100). Additionally or alternatively, the set of output images 420 can be transmitted to a different device (e.g., as part of an email, message, live stream, etc., or for remote storage). As part of storing and / or transmitting the set of output images 420, the set of output images 420 (including such metadata 422, if the set of output images 420 is associated with metadata 422) can be encoded or otherwise written in a specified format. As a non-limiting example, stereoscopic images can be written as high efficiency image file formats, such as high efficiency image container (HEIC) files. As another non-limiting example, stereoscopic video can be written as high efficiency video coding (HEVC) files, such as multi-view high efficiency video coding (MV-HEVC) files.
[0076] By separately processing the image stream 402 between the display image processing pipeline 406 and the output image processing pipeline 418, these image processing pipelines can be customized to generate different types of images. Thus, this process can generate the first set of transformed images 408 at block 412 in a manner that is customized for real-time display of these images (e.g., as part of a live preview and / or pass-through mode in an extended reality environment). These images can be discarded after they are presented to the user, or can optionally be stored in a non-transitory computer-readable storage device (e.g., as part of the memory 108 of the electronic device 100) as part of block 432. Additionally, the process 400 can generate the set of output images 420, which can be customized for different intended viewing experiences. For example, the set of output images 420 can be intended to be viewed on a different device (with the same or different configuration as the device used to perform the process 400) and / or as part of a particular user experience with different viewing requirements than those of block 412. Figure 1
[0077] It should be understood that various blocks of the process 400 can occur simultaneously and continuously. For example, certain images in the first set of transformed images 408 can be displayed at block 412 while the display image processing pipeline 406 is generating other images in the first set of transformed images 408. This can also occur while the first camera 404a and the second camera 404b are capturing images to be added to the image stream 402 and while the output image processing pipeline 418 is generating the output images 420, as discussed herein.
[0078] Figure 5 A variation of the display image processing pipeline operation 406 of the process 400 is shown. As shown, the display image processing pipeline 406 includes a perspective correction operation 500 that generates a set of perspective-corrected images from a set of images received by the display image processing pipeline 406 (e.g., the images in the image stream 402). In general, the perspective correction operation takes an image captured from one viewpoint (“source viewpoint”) and transforms the image so that it appears to have been captured from a different viewpoint (“target viewpoint”). In particular, the perspective correction operation can utilize the viewpoint of the user as the target viewpoint and transform the image accordingly so that it appears to have been captured from the viewpoint of the user. In this way, when the perspective-corrected images are displayed to the user at block 412 of the process 400, the objects presented in these images can appear at the same height and distance as real-world objects in the physical environment of the user. This can improve the accuracy with which the physical environment of the user is reproduced in the extended reality environment as part of a pass-through and / or live preview mode.
[0079] The perspective correction operation 500 can utilize the viewpoint data 410 to determine the source viewpoint (e.g., the viewpoint of the camera used to capture the image) and the target field of view (e.g., the viewpoint of the user or the viewpoint of a particular eye of the user). For example, the viewpoint data 410 can include position information that includes one or more distances and / or angles (such as, for example, the first distance 220a and the second distance 220b described with respect to Figure 2 In some cases, the offset between the set of cameras and the set of displays can be known a priori (e.g., when the set of cameras and the set of displays have a fixed spatial relationship) or can be dynamically determined by the device (e.g., when one or more of the set of cameras and / or the set of displays are movable within the device).
[0080] The viewpoint data 410 can include information about the spatial relationship between the user and the device (or components thereof). This can include a single set of values representing the relative positioning of the user’s two eyes with respect to the device, or can include different sets of values for each eye (such as, for example, the first eye-to-display distance 222a and the second eye-to-display distance 222b described with respect to Figure 2 The relative positioning between the user and the device can be determined, for example, using an eye tracker such as the eye tracker 116 of the electronic device 100, 202 described with respect to Figure 1 and Figure 2 The viewpoint data 410 can include information about the scene imaged by the set of cameras (such as the distance from the device to a particular object in the scene). This can include, for example, as determined by a depth sensor (e.g., the depth sensor 118 of the electronic device 100, 202 described with respect to Figure 1 and Figure 2 The perspective correction operation 500 can utilize the viewpoint data 410 to determine the source viewpoint (e.g., the viewpoint of the camera used to capture the image) and the target field of view (e.g., the viewpoint of the user or the viewpoint of a particular eye of the user). For example, the viewpoint data 410 can include position information that includes one or more distances and / or angles (such as, for example, the first distance 220a and the second distance 220b described with respect to Figure 2 In some cases, the offset between the set of cameras and the set of displays can be known a priori (e.g., when the set of cameras and the set of displays have a fixed spatial relationship) or can be dynamically determined by the device (e.g., when one or more of the set of cameras and / or the set of displays are movable within the device).
[0080] The viewpoint data 410 can include information about the spatial relationship between the user and the device (or components thereof). This can include a single set of values representing the relative positioning of the user’s two eyes with respect to the device, or can include different sets of values for each eye (such as, for example, the first eye-to-display distance 222a and the second eye-to-display distance 222b described with respect to Figure 2 The relative positioning between the user and the device can be determined, for example, using an eye tracker such as the eye tracker 116 of the electronic device 100, 202 described with respect to Figure 1 and Figure 2 The viewpoint data 410 can include information about the scene imaged by the set of cameras (such as the distance from the device to a particular object in the scene). This can include, for example, as determined by a depth sensor (e.g., the depth sensor 118 of the electronic device 100, 202 described with respect to Figure 1 and Figure 2 The perspective correction operation 500 can utilize the viewpoint data 410 to determine the source viewpoint (e.g., the viewpoint of the camera used to capture the image) and the target field of view (e.g., the viewpoint of the user or the viewpoint of a particular eye of the user). For example, the viewpoint data 410 can include position information that includes one or more distances and / or angles (such as, for example, the first distance 220a and the second distance 220b described with respect to Figure 2 In some cases, the offset between the set of cameras and the set of displays can be known a priori (e.g., when the set of cameras and the set of displays have a fixed spatial relationship) or can be dynamically determined by the device (e.g., when one or more of the set of cameras and / or the set of displays are movable within the device).
[0080] The viewpoint data 410 can include information about the spatial relationship between the user and the device (or components thereof). This can include a single set of values representing the relative positioning of the user’s two eyes with respect to the device, or can include different sets of values for each eye (such as, for example, the first eye-to-display distance 222a and the second eye-to-display distance 222b described with respect to Figure 2 The relative positioning between the user and the device can be determined, for example, using an eye tracker such as the eye tracker 116 of the electronic device 100, 202 described with respect to Figure 1 and Figure 2 The viewpoint data 410 can include information about the scene imaged by the set of cameras (such as the distance from the device to a particular object in the scene). This can include, for example, as determined by a depth sensor (e.g., the depth sensor 118 of the electronic device 100, 202 described with respect to Figure 1 and Figure 2distance information determined by one or more cameras of a set of cameras, by a combination thereof, etc.
[0081] Collectively, the viewpoint data 410 can be used to determine a difference between a source viewpoint and a target viewpoint for a particular image. The perspective correction operation 500 can use the difference between the source viewpoint and the target viewpoint to apply a transformation (e.g., a homography transformation or a warping function) to an input image (e.g., an image of the image stream 402) to generate a perspective corrected image. Specifically, for each pixel of the input image at a pixel location in the untransformed space, a new pixel location of the perspective corrected image is determined in the transformed space of the transformed image. Examples of perspective correction operations are discussed in greater detail in U.S. Patent No. 10,832,427 Bl entitled “Scene camera retargeting” and U.S. Patent No. 11,373,271 Bl entitled “Adaptive image warping based on object and distance information,” the contents of which are hereby incorporated by reference in their entirety and attached herewith as appendices.
[0082] In the case where the first set of images includes a set of image pairs, the perspective correction operation 500 includes a first perspective correction operation 500a and a second perspective correction operation 500b. For each image pair, the first perspective correction operation 500a generates a first perspective corrected image from a first image of the image pair. Similarly, the second perspective correction operation 500b generates a second perspective corrected image from a second image of the image pair. Specifically, the first perspective correction operation 500a can be applied to the images 402a captured by the first camera 404a and can generate a perspective corrected image based on a difference between a viewpoint of the first camera 404a and a viewpoint of the user (such as, for example, a viewpoint of the first eye of the user). Similarly, the second perspective correction operation 500b can be applied to the images 402b captured by the second camera 404b and can generate a perspective corrected image based on a difference between a viewpoint of the second camera 404b and a viewpoint of the user (such as, for example, a viewpoint of the second eye of the user).
[0083] The perspective-corrected images generated by the perspective correction operation 500 can form a first set of transformed images 408. In some cases, additional processing steps can be applied to these images. For example, at block 502, the display image processing pipeline 406 can add virtual content to the first set of transformed images 408, such as one or more graphical elements of a graphical user interface, virtual objects, etc. This can allow the process to generate an extended reality environment with virtual content, and can allow a user to experience this virtual content when the first set of transformed images 408 is displayed to the user at block 412 of the process 500. It should be understood that, as part of generating the first set of transformed images 408, the display image processing pipeline 406 can perform additional processing steps, such as cropping a portion of an image.
[0084] Figure 6 A variation of the output image processing pipeline 418 of the process 400 is shown. As shown, the output image processing pipeline 418 includes a transform unit 600 that performs a transform operation on images input into the output image processing pipeline 418 to generate transformed images. Thus, the output image processing pipeline can be configured to generate a set of transformed images from the set of selected images 416 (e.g., a second set of transformed images that is different from the first set of transformed images 408 generated by the display image processing pipeline 406). An image generation unit 602 generates a set of output images 420 from the set of transformed images. The transform operation performed by the transform unit 600 can be different from the perspective correction operation 500 utilized in the display image processing pipeline operation 406. For example, in some variations, the transform unit 600 applies an image rectification operation that transforms images to project the images into a common target image plane. In Figure 6 In examples, the transform unit 600 includes a de-warping operation 604 and an alignment operation 606 that collectively perform the image rectification operation. In some cases, a set of cameras can have fisheye lenses or other similar optics designed to increase the field of view of the cameras. While images captured by these cameras can be beneficial to reproduce a three-dimensional environment (e.g., as part of an extended reality environment as described herein), the images often include distortions when viewed as two-dimensional images (e.g., object proportions can change and straight lines in the physical environment can appear curved in the captured images).
[0085] Accordingly, the de-warp operation 604 can generate de-warped images to account for lens distortion, as will be readily appreciated by one of ordinary skill in the art. In other words, the images will be de-warped from their physical lens projection to a straight line projection, such that straight lines in three-dimensional space are mapped to straight lines in the two-dimensional images. For each image in the selected set of images 416 received by the output image processing pipeline 418, a first de-warp operation 604a can generate a first de-warped image from a first image in the image pair (e.g., one of the first subset of images 416a captured by the first camera 404a), and a second de-warp operation 604a can generate a second de-warped image from a second image in the image pair (e.g., one of the second subset of images 416b captured by the second camera 404b). In cases where the first camera 404a and the second camera 404b have different configurations (e.g., different fields of view and / or different lens designs), the first de-warp operation 604a and the second de-warp operation 604b can be customized to account for the differences between these cameras.
[0086] For each image pair in the selected set of images 416, the alignment operation 606 can align the first de-warped image and the second de-warped image generated by the de-warp operation 604 such that epipolar lines of these de-warped images are horizontally aligned. In this way, the transformation unit 600 can generate a transformed image pair for each image pair it receives, which simulates having a common image plane (even though they can be captured by cameras having different image planes). These transformed image pairs can be used by the image generation unit 602 to generate the set of output images 420.
[0087] Figure 7 A variation of the image generation unit 602 is shown, which is configured to generate the set of output images 420 as part of the process 400. Optionally, the image generation unit 602 can also be configured to generate metadata 422 associated with the set of output images 420. As shown, the image generation unit 602 includes a still image processing unit 700 configured to generate a set of still images, a video processing unit 702 configured to generate a video, and a metadata processing unit 704 to generate metadata associated with the output images. For a given media capture event, the output image processing pipeline 418 will use some or all of these units in generating the set of output images 420, depending on the photography mode. Figure 7
[0088] The still image processing unit 700 is configured to receive a first set of transformed images 701 (which can be some or all of the transformed images generated by the transformation unit 600) and can apply an image fusion operation 706 to generate one or more fused images 708 from two or more images in the first set of transformed images 701. In cases where the transformed images 701 include a set of transformed image pairs, the image fusion operation 706 generates one or more fused image pairs, which can form one or more fused stereoscopic output images. In these cases, a first fused image 708a in the fused image pair can be generated by applying a first image fusion operation 706a to a first subset of the transformed images 701 (e.g., all of which are generated from images 416a captured by the first camera 404a). A second fused image 708b in the fused image pair can be generated by applying a second image fusion operation 706b to a second subset of the transformed images 701 (e.g., all of which are generated from images 416b captured by the second camera 404b). In general, an image fusion operation includes selecting a reference image from a set of selected images and using information from other selected images to modify some or all of its pixel values. The image fusion operation 706 can generate a fused output image with reduced noise and / or higher dynamic range, and can utilize any suitable image fusion technique that would be readily understood by one of ordinary skill in the art.
[0089] The video processing unit 702 is configured to receive a second set of transformed images 703 (which can be some or all of the transformed images generated by the transformation unit 600) and can perform a video stabilization operation 710 on the second set of transformed images 703 to generate a stabilized video 712. In cases where the second set of transformed images 703 includes a set of image pairs, the stabilized video 712 can be a stereoscopic video including a first set of images 712a (e.g., generated from images 416a captured by the first camera 404a) and a second set of images 712b (e.g., generated from images 416b captured by the second camera 404b). Thus, the video processing unit 702 can output a stereoscopic video. In cases where the set of output images 420 includes both fused output images and stereoscopic video, the video processing unit 702 can optionally utilize one or more fused images 708 as part of the video stabilization operation 710 (e.g., frames of the stereoscopic video can be stabilized relative to a fused stereoscopic image). The video stabilization operation 710 can utilize any suitable electronic image stabilization technique that would be readily understood by one of ordinary skill in the art. In some cases, the video stabilization operation 710 can receive contextual information 424 (e.g., scene information, pose information, etc.), such as described above, which can aid in performing the video stabilization.
[0090] In some variations, the image generation unit 602 can optionally include an additional content unit 730 configured to add additional content to still images generated by the still image processing unit 700 and / or video generated by the video processing unit 702. For example, in some cases, the additional content unit 730 can apply virtual content 732 (e.g., one or more virtual objects) and / or other effects 734 (e.g., filters, style transfer operations, etc.) to output images generated by the image generation unit 602.
[0091] The metadata processing unit 704 is configured to generate metadata 422 associated with the set of output images 420. Depending on the type of metadata 422 being generated, the metadata processing unit 704 can receive a third set of images 705 (which can include some or all of the transformed images generated by the transformation unit 600, the output images generated by the transformation unit 600, and / or any of the images captured or generated as part of the process 400). Additionally or alternatively, the metadata processing unit 704 can receive contextual information 424, such as described herein, which can be used to generate the metadata 422.
[0092] In some variations, the metadata processing unit 704 is configured to generate field of view information 714 for use by the set of cameras to capture the set of selected images 416 used to generate the set of output images 420. This field of view information 714 can be fixed (e.g., in cases where the field of view of the set of cameras is fixed), or dynamically determined (e.g., in cases where one or more of the set of cameras has optical zoom capabilities to provide a variable field of view). Additionally or alternatively, the metadata processing unit 704 is configured to generate baseline information 716 representing one or more separation distances between cameras in the set of cameras (e.g., a separation distance between the first camera 404a and the second camera 404b). Again, this information can be fixed (e.g., in cases where the set of cameras are in fixed positions relative to one another) or dynamically determined (e.g., where one or more of the set of cameras can move with a device that houses the set of cameras).
[0093] In some cases, the metadata processing unit 718 is configured to generate pose information 718 associated with the set of output images. In particular, the pose information 718 can describe the pose of the output images, and thereby describe the relative position and / or orientation of the scene captured by the set of output images 420. This information can include or otherwise be derived from the pose information 430 associated with the underlying images (e.g., the set of selected images 416) used to generate the output images 420.
[0094] In some cases, the metadata processing unit 718 is configured to set a default disparity value 720 for a set of stereoscopic output images. The default disparity value for a given stereoscopic image represents a recommended disparity (e.g., the distance at which the images in the stereoscopic image are separated) at which to present the images in the stereoscopic image when viewing the stereoscopic image. Changing the disparity at which the stereoscopic image is presented will affect the way in which a user perceives the depth of the scene, and thus can be used to emphasize different portions of the scene captured by the set of output images 420.
[0095] In some cases, it can be desirable to select a set of default disparity values based on the scene captured by the set of output images 420. In these cases, the context information 424 such as the scene information map 428, the depth map 426, and / or the pose information 430 can be used to select the set of default disparity values. This context information 424 can be associated with the set of output images 420 (e.g., generated from images or information captured contemporaneously with the selected images 416 used to generate the set of output images 420). This context information 424 can allow the metadata processing unit 704 to determine one or more aspects of the scene captured by the set of output images 420 (e.g., the presence and / or location of one or more objects in the scene), and select default disparity values based on these determinations.
[0096] The metadata 422 generated by the metadata processing unit 704 can include information that is applicable to all of the images in the set of output images and / or information that is specific to particular images or image pairs. For example, the metadata 422 can include pose information 718 for an image or images in the output images 420, but can include a single set of baseline information 716 that applies to all of the output images 420. Similarly, the set of default disparity values 720 can include a single value that applies to each image pair in the set of output images 420, or can include different values for different image pairs (e.g., a first stereoscopic image can have a first default disparity value and a second stereoscopic image can have a second default disparity value).
[0097] For the purposes of illustration, the foregoing descriptions use specific naming to provide a thorough understanding of the described embodiments. However, it will be apparent to one skilled in the art that specific details are not required in order to practice the described embodiments. Thus, the foregoing descriptions of specific embodiments are presented for purposes of illustration and description. They are not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Many modifications and variations are possible in view of the above teachings. The described embodiments are to be considered in a descriptive sense only and not for purposes of limitation.
[0098] Implementation examples are described in the following numbered clauses:
[0099] 1. A device, comprising:
[0100] a set of cameras;
[0101] a set of displays;
[0102] a memory; and
[0103] one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to:
[0104] capture a stream of images using the set of cameras;
[0105] generate a first set of transformed images from the stream of images by applying a first transform operation to images in the stream of images;
[0106] display the first set of transformed images using the set of displays;
[0107] receive a capture request; and
[0108] in response to receiving the capture request, generate a set of stereoscopic output images, wherein generating the set of stereoscopic output images comprises:
[0109] select a set of images from the stream of images;
[0110] generate a second set of transformed images from the set of images by applying a second transform operation to the set of images that is different from the first transform operation;
[0111] and
[0112] generate the set of stereoscopic output images using the second set of transformed images.
[0113] 2. The device of clause 1, wherein generating the set of stereoscopic output images comprises:
[0114] generating a fused stereoscopic output image from two or more image pairs in the set of images.
[0115] 3. The device of clause 1, wherein the set of stereoscopic output images comprises a stereoscopic video.
[0116] 4. The device of clause 3, wherein generating the set of stereoscopic output images comprises:
[0117] performing a video stabilization operation on the second set of transformed images.
[0118] 5. The device of any of clauses 1-4, wherein the first transform operation is a perspective correction operation.
[0119] 6. The device of any of clauses 1-5, wherein the second transform operation is an image correction operation.
[0120] 7. The device of any of clauses 1-6, wherein generating the set of stereoscopic output images comprises generating metadata associated with the set of stereoscopic output images.
[0121] 8. The device of clause 7, wherein the metadata comprises field of view information for the set of cameras.
[0122] 9. The device of clause 7 or clause 8, wherein the metadata comprises pose information for the set of stereoscopic output images.
[0123] 10. The device of any of clauses 7-9, wherein the metadata comprises a set of default disparity values for the set of stereoscopic output images.
[0124] 11. The device of clause 10, wherein generating metadata associated with the set of stereoscopic output images comprises selecting the set of default disparity values based on a scene captured by the set of stereoscopic output images.
[0125] 12. A method comprising:
[0126] capturing a stream of images using a set of cameras of a device;
[0127] generating a first set of transformed images from the stream of images by applying a first transform operation to images in the stream of images;
[0128] displaying the first set of transformed images at a set of displays of a device;
[0129] receiving a capture request; and
[0130] in response to receiving the capture request, generating a set of stereoscopic output images, wherein
[0131] generating the set of stereoscopic output images comprises:
[0132] selecting a set of images from the stream of images;
[0133] generating a second set of transformed images from the set of images by applying a second transform operation to the set of images that is different from the first transform operation; and
[0134] generating the set of stereoscopic output images using the second set of transformed images.
[0135] 13. The method of clause 12, wherein generating the set of stereoscopic output images comprises:
[0136] generating a fused stereoscopic output image from two or more of the set of images.
[0137] 14. The method of clause 12, wherein the set of stereoscopic output images comprises a stereoscopic video.
[0138] 15. The method of clause 14, wherein generating the set of stereoscopic output images comprises:
[0139] performing a video stabilization operation on the second set of transformed images.
[0140] 16. The method of any of clauses 12-15, wherein the first transformation operation is a perspective correction operation.
[0141] 17. The method of any of clauses 12-16, wherein the second transformation operation is an image rectification operation.
[0142] 18. The method of any of clauses 12-17, wherein generating the set of stereoscopic output images comprises generating metadata associated with the set of stereoscopic output images.
[0143] 19. The method of clause 18, wherein the metadata comprises field of view information for the set of cameras.
[0144] 20. The method of clause 18 or clause 19, wherein the metadata comprises pose information for the set of stereoscopic output images.
[0145] 21. The method of any of clauses 18-20, wherein the metadata comprises a set of default disparity values for the set of stereoscopic output images.
[0146] 22. The method of clause 21, wherein generating metadata associated with the set of stereoscopic output images comprises selecting the set of default disparity values based on a scene captured by the set of stereoscopic output images.
[0147] 23. The method of any of clauses 12-22, the method comprising adding virtual content to the first set of transformed images.
[0148] 24. A non-transitory computer-readable medium comprising instructions that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising the steps of any of clauses 12-23.
[0149] 25. A device, the device comprising:
[0150] a first camera;
[0151] a second camera;
[0152] a set of displays;
[0153] a memory; and
[0154] one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or
[0155] instructions causing the one or more processors to:
[0156] capture, using the first camera and the second camera, a first set of pairs of images of a scene, wherein each pair of images of the first set of pairs of images includes:
[0157] a first image captured by the first camera; and a second image captured by the second camera;
[0158] generate, from the first set of pairs of images, a first set of transformed pairs of images, wherein generating the first set of transformed pairs of images includes, for each pair of images of the first set of pairs of images:
[0159] generate, from the first image, a first perspective-corrected image based on a difference between a viewpoint of the first camera and a viewpoint of a user; and generate, from the second image, a second perspective-corrected image based on a difference between a viewpoint of the second camera and the viewpoint of the user;
[0160] display, on the set of displays, the first set of transformed pairs of images;
[0161] receive a capture request;
[0162] in response to receiving the capture request, select, from the first set of pairs of images, a second set of pairs of images;
[0163] generate, from the second set of pairs of images, a second set of transformed pairs of images, wherein generating the second set of transformed pairs of images includes, for each pair of images of the first set of pairs of images:
[0164] generate, from the first image, a first de-warp image;
[0165] generate, from the second image, a second de-warp image; and align the first de-warp image and the second de-warp image;
[0166] and
[0167] generate, using the second set of transformed pairs of images, a set of output images.
[0168] 26. The device of clause 25, wherein generating the set of output images comprises:
[0169] generating a fused stereoscopic output image from two or more image pairs in the second set of transformed image pairs.
[0170] 27. The device of clause 25, wherein the set of output images comprises a stereoscopic video formed from the second set of transformed image pairs.
[0171] 28. The device of clause 27, wherein generating the set of output images comprises:
[0172] performing a video stabilization operation on the second set of transformed image pairs.
[0173] 29. The device of any of clauses 25-28, wherein generating the set of output images comprises generating metadata associated with the set of output images.
[0174] 30. The device of clause 29, wherein the metadata comprises field of view information for at least one of the first camera or the second camera.
[0175] 31. The device of clause 29 or clause 30, wherein the metadata comprises pose information for the set of output images.
[0176] 32. The device of any of clauses 29-31, wherein:
[0177] the set of output images comprises a set of stereoscopic images; and
[0178] the metadata comprises a set of default disparity values for the set of stereoscopic output images.
[0179] 33. The device of clause 32, wherein generating metadata associated with the set of output images comprises selecting the set of default disparity values based on the scene captured by the set of output images.
[0180] 34. The device of any of clauses 25-33, wherein the processor is configured to add virtual content to the first set of transformed image pairs.
[0181] 35. A method comprising:
[0182] capturing, using a first camera and a second camera of a device, a first set of image pairs of a scene, wherein each image pair in the first set of image pairs comprises:
[0183] a first image captured by the first camera and a second image captured by the second camera.
[0184] a first image captured by the first camera; and
[0185] a second image captured by the second camera;
[0186] generating a first set of transformed image pairs from the first set of image pairs, wherein generating the first set of transformed image pairs comprises, for each image pair in the first set of image pairs,
[0187] each image pair:
[0188] generating a first perspective-corrected image from the first image based on a difference between a viewpoint of the first camera and a viewpoint of a user; and
[0189] generating a second perspective-corrected image from the first image based on a difference between a viewpoint of the second camera and the viewpoint of the user;
[0190] displaying the first set of transformed image pairs on a set of displays of the device;
[0191] receiving a capture request;
[0192] in response to receiving the capture request, selecting a second set of image pairs from the first set of image pairs;
[0193] generating a second set of transformed image pairs from the second set of image pairs comprises, for each image pair in the second set of image pairs,
[0194] each image pair in the first set of image pairs:
[0195] generating a first de-warp image from the first image;
[0196] generating a second de-warp image from the second image; and
[0197] aligning the first de-warp image and the second de-warp image; and using the second set of transformed image pairs to generate a set of output images.
[0198] 36. The method of clause 35, wherein generating the set of output images comprises:
[0199] generating a fused stereoscopic output image from two or more image pairs in the second set of transformed image pairs.
[0200] 37. The method of clause 35, wherein the set of output images comprises a stereoscopic video formed from the second set of transformed image pairs.
[0201] 38. The method of clause 37, wherein generating the set of output images comprises:
[0202] performing a video stabilization operation on the second set of pairs of transformed images.
[0203] 39. The method of any of clauses 35-38, wherein generating the set of output images comprises generating metadata associated with the set of output images.
[0204] 40. The method of clause 39, wherein the metadata comprises field of view information for at least one of the first camera or the second camera.
[0205] 41. The method of clause 39 or clause 40, wherein the metadata comprises pose information for the set of output images.
[0206] 42. The method of any of clauses 39-41, wherein:
[0207] the set of output images comprises a set of stereoscopic images; and
[0208] the metadata comprises a set of default disparity values for the set of stereoscopic output images.
[0209] 43. The method of clause 42, wherein generating metadata associated with the set of output images comprises selecting the set of default disparity values based on the scene captured by the set of output images.
[0210] 44. The method of any of clauses 35-43, comprising adding virtual content to the first set of pairs of transformed images.
[0211] 45. A non-transitory computer-readable medium comprising instructions that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising the steps of any of clauses 35-44.
Claims
1. A device, the device comprising: a set of cameras; a set of displays; a memory; and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions causing the one or more processors to: capture a stream of images using the set of cameras; generate a first set of transformed images from the stream of images by applying a first transform operation to images in the stream of images; display the first set of transformed images using the set of displays; receive a capture request to capture a set of stereoscopic output images; and in response to receiving the capture request, generate the set of stereoscopic output images, wherein generating the set of stereoscopic output images comprises: selecting a set of images from the stream of images; generating a second set of transformed images from the set of images by applying a second transform operation to the set of images that is different from the first transform operation; and generating the set of stereoscopic output images using the second set of transformed images, wherein: the first set of transformed images and the second set of transformed images are generated from the same images.
2. The device of claim 1, wherein generating the set of stereoscopic output images comprises: generating a fused stereoscopic output image from two or more image pairs in the set of images.
3. The device of claim 1, wherein the set of stereoscopic output images comprises a stereoscopic video.
4. The device of claim 3, wherein generating the set of stereoscopic output images comprises: performing a video stabilization operation on the second set of transformed images.
5. The device of claim 1, wherein the first transform operation is a perspective correction operation.
6. The device of claim 1, wherein the second transform operation is an image rectification operation. generating metadata associated with the set of stereoscopic output images.
8. The device of claim 7, wherein the metadata comprises field of view information for the set of cameras.
7. The device of claim 1, wherein generating the set of stereoscopic output images comprises:
9. The device of claim 7, wherein the metadata comprises pose information for the set of stereoscopic output images.
10. The device of claim 7, wherein the metadata comprises a set of default disparity values for the set of stereoscopic output images. selecting the set of default disparity values based on a scene captured by the set of stereoscopic output images.
12. A method, comprising:
11. The device of claim 10, wherein generating metadata associated with the set of stereoscopic output images comprises: capturing a stream of images using a set of cameras of a device; generating a first set of transformed images from the stream of images by applying a first transform operation to images in the stream of images; displaying the first set of transformed images at a set of displays of a device; receiving a capture request to capture a set of stereoscopic output images; and in response to receiving the capture request, generating the set of stereoscopic output images, wherein generating the set of stereoscopic output images comprises: selecting a set of images from the stream of images; generating a second set of transformed images from the set of images by applying a second transform operation to the set of images that is different from the first transform operation; and generating the set of stereoscopic output images using the second set of transformed images, wherein: generating a set of stereoscopic output images using the second set of transformed images, wherein: the first set of transformed images and the second set of transformed images are generated from the same images.
13. The method of claim 12, wherein generating the set of stereoscopic output images comprises: generating a fused stereoscopic output image from two or more pairs of images in the set of images.
14. The method of claim 12, wherein the set of stereoscopic output images comprises a stereoscopic video.
15. The method of claim 14, wherein generating the set of stereoscopic output images comprises: performing a video stabilization operation on the second set of transformed images.
16. The method of claim 12, wherein the first transformation operation is a perspective correction operation.
17. The method of claim 12, wherein the second transformation operation is an image rectification operation.
18. The method of claim 12, wherein generating the set of stereoscopic output images comprises: generating metadata associated with the set of stereoscopic output images.
19. The method of claim 18, wherein the metadata comprises field of view information for the set of cameras.
20. The method of claim 18, wherein the metadata comprises pose information for the set of stereoscopic output images.
21. The method of claim 18, wherein the metadata comprises a set of default disparity values for the set of stereoscopic output images.
22. The method of claim 21, wherein generating metadata associated with the set of stereoscopic output images comprises: selecting the set of default disparity values based on a scene captured by the set of stereoscopic output images.
23. The method of claim 12, comprising adding virtual content to the first set of transformed images.
Citation Information
Patent Citations
Scene camera retargeting
US10832427B1
Adaptive image warping based on object and distance information
US11373271B1
Spherical virtual reality camera
WO2017118662A1