Technology for displaying and capturing images
Separate image processing pipelines in extended reality systems correct perspective differences to enhance user comfort and image capture quality, ensuring accurate display and storage of physical environments.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- APPLE INC
- Filing Date
- 2024-05-30
- Publication Date
- 2026-05-08
AI Technical Summary
In extended reality systems, the perspective difference between the user and the camera can cause the displayed image to inaccurately reflect the user's physical environment, leading to user discomfort and adverse impacts on the experience during pass-through modes.
Separate image processing pipelines are used to generate perspective-corrected images for display and stereo output images for capture, including transformations and metadata generation to align camera viewpoints with user perspectives, and optionally adding virtual content.
This approach enhances user comfort by accurately displaying the physical environment while allowing for high-quality image capture and storage, maintaining user experience across different viewing environments.
Smart Images

Figure 0007855639000001 
Figure 0007855639000002 
Figure 0007855639000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application is non - provisional and claims the benefit under 35 U.S.C. 119(e) of U.S. Provisional Patent Application No. 63 / 470,081, filed on May 31, 2023, the content of which is incorporated herein by reference as if fully disclosed herein.
[0002] The described embodiments generally relate to separately processing image streams for real - time display and media capture events.
Background Art
[0003] An extended reality system can be used to generate a partially or fully simulated environment (e.g., a virtual reality environment, a mixed reality environment, etc.) in which virtual content can replace or extend the physical world. The simulated environment can provide an engaging experience for the user and is used in games, personal communication, virtual travel, healthcare, and many other contexts. In some cases, the simulated environment may include information captured from the user's environment. For example, the pass - through mode of an extended reality system may display an image of the user's physical environment using one or more displays of the extended reality system. Thereby, the user can perceive their physical environment through the display(s) of the extended reality system.
[0004] An image of the user's environment may be captured using one or more cameras of the extended reality system. Depending on the camera placement in the extended reality system, the perspective difference between the user and the camera, if not corrected, can cause the image displayed in the pass - through mode not to accurately reflect the user's physical environment. This can cause discomfort to the user or have an adverse impact on the user experience during the pass - through mode.
Summary of the Invention
[0005] The embodiments described herein relate to systems, devices, and methods for performing separate image processing. Some embodiments relate to a method that includes the steps of capturing an image stream using a set of cameras of the device, generating a first set of transformed images from the image stream by applying a first transformation operation to the images in the image stream, and displaying the first set of transformed images on a set of displays of the device. The method further includes the steps of receiving a capture request and generating a set of stereo output images in response to receiving the capture request. The step of generating a set of stereo output images includes the steps of selecting a set of images from the image stream, generating a second set of transformed images from the set of images by applying a second transformation operation different from the first transformation operation to the set of images, and generating a set of stereo output images using the second set of transformed images.
[0006] In some variations of these methods, the step of generating a set of stereo output images includes the step of generating a fused stereo output image from two or more pairs of images in the set of images. Additionally or alternatively, the set of stereo output images includes a stereo video. In some of these variations, the step of generating a set of stereo output images includes the step of performing a video stabilization operation on a second set of transformed images.
[0007] In some cases, the first transformation operation is a perspective correction operation. Alternatively, the second transformation operation may be an image correction operation. The step of generating a set of stereo output images may include a step of generating metadata associated with the set of stereo output images. The metadata may include field-of-view information for the set of cameras, pose information for the set of stereo output images, and / or a set of default parallax values for the set of stereo output images. In some cases, the step of generating metadata associated with the set of stereo output images includes a step of selecting a set of default parallax values based on the scene captured by the set of stereo output images. In some variations, virtual content is added to the first set of transformed images.
[0008] Some embodiments relate to a device comprising a set of cameras, a set of displays, memory, and one or more processors operably coupled to the memory, wherein the one or more processors are configured to execute instructions causing one or more processors to perform any of the above methods. Similarly, yet another embodiment relates to a non-temporary computer-readable medium containing instructions, which, when executed by at least one computing device, cause at least one computing device to perform an operation including any of the steps of the above methods.
[0009] Further embodiments relate to a method that includes the step of capturing a first set of image pairs of a scene using a first camera and a second camera of the device. Each image pair in the first set of image pairs includes a first image captured by the first camera and a second image captured by the second camera. The method includes the step of generating a first set of transformed image pairs from the first set of image pairs, the step of generating a first perspective-corrected image from the first image based on the difference between the viewpoint of the first camera and the viewpoint of the user, and for each image pair in the first set of image pairs, the step of generating a second perspective-corrected image from the first image based on the difference between the viewpoint of the second camera and the viewpoint of the user. The method further includes the step of displaying the first set of transformed image pairs on a set of displays of the device.
[0010] Furthermore, the method includes the steps of receiving a capture request, selecting a second set of image pairs from a first set of image pairs in response to receiving the capture request, and generating a second set of transformed image pairs from the second set of image pairs. This includes, for each image pair in the first set of image pairs, generating a first dewarped image from a first image, generating a second dewarped image from a second first image, and aligning the first and second dewarped images. The method also includes the step of generating a set of output images using the second set of transformed image pairs.
[0011] In some variations of these methods, the step of generating a set of output images includes the step of generating a fused stereo output image from two or more pairs of images from the set of transformed image pairs. Additionally or alternatively, the set of stereo output images includes a stereo video formed from a second set of transformed image pairs. In some of these variations, the step of generating a set of stereo output images includes the step of performing a video stabilization operation on the second set of transformed image pairs.
[0012] The step of generating a set of stereo output images may include the step of generating metadata associated with the set of stereo output images. The metadata may include field of view information from at least one of the first or second cameras, pose information for the set of output images, and / or a set of default disparity values for the set of stereo images of the stereo output images. In some cases, the step of generating metadata associated with the set of stereo output images includes the step of selecting a set of default disparity values based on the scene captured by the set of output images. In some variations, virtual content is added to the first set of transformed image pairs.
[0013] Some embodiments relate to a device comprising a first camera, a set of second displays, memory, and one or more processors operably coupled to the memory, wherein the one or more processors are configured to execute instructions causing one or more processors to perform any of the methods described above. Similarly, yet another embodiment relates to a non-temporary computer-readable medium containing instructions, which, when executed by at least one computing device, cause at least one computing device to perform an operation including any of the steps of the methods described above.
[0014] In addition to the exemplary embodiments and models described above, further embodiments and models will become apparent by referring to the drawings and considering the following description.
[0015] The disclosure will be easily understood by the following detailed description, along with the attached drawings in which similar reference numbers specify similar structural elements. [Brief explanation of the drawing]
[0016] [Figure 1] This specification shows a block diagram of an exemplary electronic device that may be used in the augmented reality system described herein.
[0017] [Figure 2] An exemplary operating environment is shown that includes a variant form of the electronic device described herein.
[0018] [Figure 3] An example of an image stream is shown.
[0019] [Figure 4] A process is shown of separately processing images of an image stream using a display image processing pipeline and an output image processing pipeline.
[0020] [Figure 5] An example of the display image processing pipeline described herein is shown.
[0021] [Figure 6] An example of the output image processing pipeline described herein is shown.
[0022] [Figure 7] An example of an image generation unit that can be used as part of the output image processing pipeline described herein is shown.
[0023] The proportions and dimensions (relative or absolute) of various features and elements (as well as their collections and groups), and the boundaries, separation points, and positional relationships presented between them, are provided in the accompanying figures merely to facilitate the understanding of the various embodiments described herein, and thus may not necessarily be presented or illustrated to scale, and it should be understood that there is no intention to indicate any preference or requirement for the illustrated embodiments, excluding the embodiments described with reference to them.
Mode for Carrying Out the Invention
[0024] Here, representative embodiments illustrated in the accompanying drawings will be described in detail. It should be understood that the following description is not intended to limit these embodiments to one preferred embodiment. On the contrary, the following description is intended to encompass alternative forms, modifications, and equivalents that can be included within the spirit and scope of the described embodiments as defined by the appended claims.
[0025] The embodiments disclosed herein are directed to devices, systems, and methods for separately processing image streams to generate a first set of images for display by a set of displays and a second set of images for storage or transfer as part of a media capture event. Specifically, a first set of transformed images can be generated using the first set of images from the image stream, and these transformed images can be displayed on the set of displays. The second set of images can be selected from the image stream in response to capture requests associated with the media capture event, and the second set of images can be used to generate a second set of transformed images. The set of output images may be generated from the second set of transformed images, and the set of output images may be stored or transmitted for later viewing.
[0026] In the following, these embodiments and other embodiments will be described with reference to FIGS. 1 to 7. However, those skilled in the art will readily understand that the forms for implementing the invention given in this specification with respect to these figures are for illustrative purposes only and should not be construed as limiting.
[0027] The devices and methods described herein may be used as part of an augmented reality system in which an augmented reality environment is generated and displayed to the user. Various terms are used herein to describe the various augmented reality systems and associated augmented reality environments described herein. For example, as used herein, “physical environment” is a part of the physical / real world around the user that the user can perceive and interact with without the assistance of the augmented reality system described herein. For example, the physical environment may include a room or outdoor space in a building, as well as any people, animals, or objects in that space (collectively referred to herein as “real-world objects”) such as plants, furniture, books, etc.
[0028] As used herein, “augmented reality environment” means a fully or partially simulated environment that a user can perceive and / or interact with using an augmented reality system as described herein. In some cases, an augmented reality environment may be a virtual reality environment, which refers to a fully simulated environment in which the user’s physical environment is completely replaced by virtual content within the virtual reality environment. A virtual reality environment may not depend on the user’s physical environment and may therefore allow the user to perceive being in a different simulated location (e.g., standing on a beach when actually standing in a room in a building). A virtual reality environment may include virtual objects with which the user can interact (e.g., simulated objects that can be perceived by the user but do not actually exist in the physical environment).
[0029] In other cases, the augmented reality environment may be a mixed reality environment, a fully or partially simulated environment, in which virtual content may be presented alongside parts of the user's physical environment. Specifically, the mixed reality environment may include a reproduction and / or modified representation of one or more parts of the user's physical environment surrounding the augmented reality system. In this way, the user may be able to perceive their physical environment (directly or indirectly) through the mixed reality environment while still perceiving the virtual content.
[0030] As used herein, “reproduction” of a part of the physical environment refers to a part of the augmented reality environment that reproduces that part of the physical environment within the augmented reality environment. For example, an augmented reality system may include one or more cameras capable of capturing images of the physical environment. The augmented reality system may present parts of these images to the user by displaying them through an opaque display so that the user indirectly views the physical environment through the displayed images. In addition, in some cases, the augmented reality environment is displayed using foveated rendering, and different parts of the augmented reality environment are rendered using different levels of fidelity (e.g., image resolution) depending on the direction of the user's line of sight. In these cases, parts of the reproduction rendered at lower fidelity using these foveated rendering techniques are still considered reproductions for the purposes of this application.
[0031] As used herein, a “modified representation” of a portion of the physical environment refers to a portion of the augmented reality environment that is derived from the physical environment but intentionally obscures one or more aspects of the physical environment. Indirect representation attempts to replicate a portion of the user’s physical environment within the augmented reality environment, but a modified representation intentionally alters one or more visual aspects of that portion of the user’s physical environment (e.g., by using one or more visual effects such as artificial bokeh). In this way, a modified representation of a portion of the user’s physical environment may allow the user to perceive certain aspects of that portion of the physical environment while obscuring others. In the case of artificial bokeh, the user may still perceive the overall shape and placement of real-world objects within the modified representation, but may not perceive the visual details of these objects that would otherwise be visible in the physical environment. In cases where the augmented reality environment is displayed using foveated rendering, portions of the modified representation that are in the peripheral areas of the augmented reality environment (relative to the user’s line of sight) may be rendered at lower fidelity using the foveated rendering technique.
[0032] Reproductions and / or modified representations of a user's physical environment can be used in various augmented reality environments. For example, an augmented reality system may be configured to operate in "pass-through" mode, during which the system generates and presents an augmented reality environment that includes a reproduction of a portion of the user's physical environment. This allows the user to indirectly view a portion of their physical environment even if one or more components of the augmented reality system (e.g., a display) could interfere with the user's ability to directly view the same portion of their physical environment. When the augmented reality system is operating in pass-through mode, part or all of the augmented reality environment may include a reproduction of the user's physical environment. In some cases, one or more portions of the augmented reality environment may include virtual content and / or a modified representation of the user's physical environment in addition to the reproduction. In addition, or alternatively, virtual content (e.g., graphical elements of a graphical user interface, virtual objects) may overlap the portion of the reproduction.
[0033] In general, the augmented reality systems described herein include electronic devices capable of capturing and displaying images to a user. Figure 1 shows a block diagram of an exemplary electronic device 100 that may be part of an augmented reality system described herein. The electronic device 100 may be configured to capture, process, and display images as part of an augmented reality environment. In some implementations, the electronic device 100 is a handheld electronic device (e.g., a smartphone or tablet) configured to present an augmented reality environment to a user. In some of these implementations, the handheld electronic device may be temporarily attached to an accessory that allows the handheld electronic device to be worn by the user. In other implementations, the electronic device 100 is a head-mounted device (HMD) worn by the user.
[0034] In some embodiments, the electronic device 100 has a bus 102 that operably connects an I / O section 104 to one or more computer processors 106 and memory 108, and includes circuits that interconnect and control communication between components of the electronic device 100. The I / O section 104 includes various system components that can assist the operation of the electronic device 100. The electronic device 100 includes a set of displays 110 and a set of cameras 112. The set of cameras 112 can capture images of the user's physical environment (e.g., a scene) and use these images when generating an augmented reality environment (e.g., as part of a pass-through mode), as described herein. The set of displays 110 can display the augmented reality environment so that the user can see it, thereby enabling the user to perceive their physical environment through the electronic device 100. The set of displays 110 may include a single display or multiple displays (e.g., one display for each eye of the user). Each display in the set of displays may utilize any suitable display technology, including, for example, liquid crystal displays (LCDs), organic light-emitting diode (OLED) displays, light-emitting diode (LED) displays, quantum dot light-emitting diode (QLED) displays, and the like.
[0035] The memory 108 of the electronic device 100 may include one or more non-temporary computer-readable storage devices. These non-temporary computer-readable storage devices may be used to store computer executable instructions, which, when executed by one or more computer processors 106, cause the computer processors to perform processes described herein (various image capture and processing techniques described below). In addition, non-temporary computer-readable storage devices may be used to store images captured as part of a media capture event, such as those described herein.
[0036] A computer-readable storage device can be any medium capable of tangibly containing or storing computer-executable instructions for use by or in connection with an instruction execution system, apparatus, or device (e.g., one or more processors 106). In some embodiments, the storage device is a temporary computer-readable storage medium. In some embodiments, the storage device is a non-temporary computer-readable storage medium. Non-temporary computer-readable storage devices may include, but are not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include magnetic disks, CDs, DVDs, or optical disks based on Blu-ray technology, as well as persistent solid-state memory such as flash and solid-state drives.
[0037] One or more computer processors 106 may include, for example, a processor, a microprocessor, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any other programmable logic device (PLD) that can be configured to run the operating system and applications of the electronic device 100 and to facilitate the capture, processing, display, and storage of the images described herein.
[0038] Therefore, any of the processes described herein may be stored as instructions on a non-temporary computer-readable storage device, thereby allowing a processor to utilize these instructions to perform various steps of the processes described herein. Similarly, the device described herein includes memory (e.g., memory 108) and one or more processors (e.g., processor 106) operably coupled to the memory. One or more processors are configured to receive instructions from memory and execute these instructions to perform various steps of the processes described herein. Any of the processes described herein may be performed using the device described herein as a method for capturing and displaying images.
[0039] In some implementations, the electronic device 100 includes a set of depth sensors 114 (which may include a single depth sensor or multiple depth sensors), each of which is configured to compute depth information about a portion of the environment in front of the electronic device 100. Specifically, each of the set of depth sensors 114 can compute depth information within a coverage area (e.g., the widest lateral range for which its depth sensor can provide depth information). In some cases, the coverage area of one of the set of depth sensors 114 may at least partially overlap with the field of view of at least one of the set of cameras 112 (e.g., the spatial extent of the scene that the camera can capture using its image sensors), thereby enabling the depth sensor to compute depth information associated with the field of view(s) of the set of cameras 112.
[0040] Information from the depth sensor can be used to calculate the distance between the depth sensor and various points in the environment around the electronic device 100. In some cases, depth information from a set of depth sensors 114 can be used to generate a depth map as described herein. The depth information can be calculated in any preferred way. In one non-limiting embodiment, the depth sensor can utilize stereo imaging, where two images are taken from different positions, and the depth information can be calculated using the distance (parallax) between corresponding pixels in the two images. In another embodiment, the depth sensor can utilize structured light imaging, where the depth sensor can image the scene while projecting a known illumination pattern (typically using infrared illumination) onto the scene, and then calculate the depth information by observing how the pattern is distorted by the scene. In yet another embodiment, the depth sensor can utilize time-of-flight sensing, where the depth is calculated based on the time it takes for light (typically infrared) emitted from the depth sensor to return from the scene. Time-of-flight depth sensors may utilize direct or indirect time of flight, illuminate the entire coverage area at once, or illuminate only a subset of the coverage area at a given time (e.g., by one or more spots, stripes, or other patterns that may be fixed or scann across the coverage area). In cases where the depth sensor utilizes infrared illumination, this infrared illumination may be available under certain ambient conditions without being perceived by the user.
[0041] In some implementations, the electronic device 100 includes an eye-tracker 116. The eye-tracker 116 may be configured to determine the position of the user's eyes relative to the electronic device 100 (or a particular component thereof). The eye-tracker 116 may include any suitable hardware for identifying and locating the user's eyes, such as one or more cameras, depth sensors, or a combination thereof. In some cases, the eye-tracker 116 may also be configured to detect the location / direction of the user's gaze. It should be understood that the eye-tracker 116 may include a single module configured to determine the position of both of the user's eyes, or it may include multiple units, each configured to determine the position of the user's corresponding eye.
[0042] As an addition or alternative, the electronic device 100 may include a set of sensors 118 capable of determining the movement and / or orientation of the electronic device 100 (or a particular component thereof). For example, the set of sensors 118 shown in Figure 1 includes an accelerometer 120 and a gyroscope 122. Information from the set of sensors 118 may be used to calculate attitude information of the electronic device 100 (or a particular component thereof). This may allow the electronic device 100 to track its position and orientation as it is moved. As an addition or alternative, information captured by a set of cameras 112 and / or a set of depth sensors 114 may be used to calculate attitude information of the electronic device 100 (or a particular component thereof), for example, by using simultaneous localization and mapping (SLAM) technology.
[0043] In addition, the electronic device 100 may include a set of input mechanisms 124 that the user can operate to interact with and provide input to an augmented reality system (e.g., a touchscreen, soft keys, a keyboard, a virtual keyboard, buttons, knobs, joysticks, switches, dials, or combinations thereof). The electronic device 100 may include a communication unit 126 that receives application and operating system data using Wi-Fi, Bluetooth, near-field communication (NFC), cellular, and / or other wireless communication technologies. The electronic device 100 is not limited to the components and configurations shown in Figure 1 and may include other or additional components (e.g., a GPS sensor, a compass or other direction sensor, a physiological sensor, a microphone, a speaker, a haptic engine, etc.) in multiple configurations as desired.
[0044] Figure 2 shows an operating environment 200 including an electronic device 202 which may be part of an augmented reality system for performing the processes described herein. The electronic device 202 may be configured in any manner as described above with respect to the electronic device 100 of Figure 1. For example, the electronic device 202 may include a set of cameras including a first camera 112a and a second camera 112b, and a set of displays including a first display 110a and a second display 110b. In addition, the electronic device 202 is shown as including a depth sensor 114 and an eye-tracker 116 as described herein.
[0045] The first camera 112a has a first field of view 230a and is positioned to capture an image of a first portion of the physical environment surrounding the electronic device 202. Similarly, the second camera 112b has a second field of view 230b and is positioned to capture an image of a second portion of the physical environment surrounding the electronic device 202. The first field of view 230a and the second field of view 230b may overlap at least partially so that the first and second cameras 112a, 112b can simultaneously capture a common portion of the physical environment surrounding the electronic device 202. For example, the operating environment 200 shown in Figure 2 includes real-world objects 206 that are located in both the first field of view 230a and the second field of view 230b. Therefore, when the electronic device 202 is positioned as shown in Figure 2, the images captured by both the first and second cameras 112a, 112b will include real-world objects 206.
[0046] As will be described in more detail herein, images captured by the first and second cameras 112a, 112b may be displayed on a set of displays as part of an augmented reality environment (for example, as part of a pass-through mode) to provide the user with a representation of their physical environment. For example, images captured by the first camera 112a may be displayed on the first display 110a (for example, as part of an augmented reality environment) and may be seen by the user's first eye 204a, such as when the electronic device 202 is worn on the user's head or otherwise held close to the user's head. Similarly, images captured by the second camera 112b may be displayed on the second display 110b (for example, as part of an augmented reality environment) and may be seen by the user's second eye 204b.
[0047] Depending on the relative arrangement of the components of the electronic device 202, and the relative arrangement between the electronic device 202 and the user's eyes 204a, 204b, the images captured by the first camera 112a and the second camera 112b may provide a viewpoint of the operating environment 200 that differs from the view the user would have of their physical environment directly without the electronic device 202. Specifically, the set of cameras may be spatially offset from the set of displays in one or more directions. For example, the imaging plane of the first camera 112a may be separated from the first display 110a by a distance of 220a along a first (e.g., horizontal) direction, and the imaging plane of the second camera 112b may be separated from the second display 110b by a distance of 220b along the first direction. The imaging planes of the first and second cameras 112a, 112b may also, in some cases, be rotated relative to the first and second displays 110a, 110b, respectively.
[0048] In addition, the first and second eyes 204a, 204b may be separated from the first and second displays 110a, 110b by corresponding distances (for example, the first eye-to-display distance 222a between the first display 110a and the first eye 204a, and the second eye-to-display distance 222b between the second display 110b and the second eye 204b). The first camera 112a is separated by a first reference point (represented by object 206 in Figure 2) by a first scene distance 232a, and the second camera 112b is separated by a second reference point (which may be the same point as the first reference point or a different point) by a second scene distance 232b. The offset (spatial and / or rotational) between the first camera 112a and the first display 110a, the first eye-to-display distance 222a, and the first scene distance 232a can collectively characterize the difference between the viewpoint of the first eye 204a and the viewpoint of the first camera 112a with respect to the first reference point. Similarly, the offset between the second camera 112b and the second display 110b, the second eye-to-display distance 222b, and the second scene distance 232b can collectively characterize the difference between the viewpoint of the second eye 204b and the viewpoint of the second camera 112b with respect to the second reference point.
[0049] Therefore, it may be desirable to adjust the images captured by the camera set to account for these viewpoint differences (e.g., to generate perspective-corrected images) before they are displayed, thereby reducing user discomfort when users view these images. However, these viewpoint adjustments may be undesirable if the user wishes to use the camera set to capture images or videos (e.g., to save for later viewing or to share with others). These captured images may later be viewed by different devices (which may have different configurations than device 202), viewed from different viewpoints (e.g., by different users or users positioned at different distances from the device), and / or viewed in specific ways within an augmented reality environment (e.g., viewed within a viewing window or frame in an augmented reality environment). In these cases, presenting perspective-corrected images may lead to an undesirable user experience, and further adjusting these images to account for new viewing environments may undesirably introduce or amplify artifacts.
[0050] The process described herein includes the steps of acquiring a set of images from one or more cameras (for example, one or more of the camera sets of electronic devices 100 and 202 in Figures 1 and 2) and processing these images using a plurality of separate image processing pipelines. Specifically, a first image processing pipeline generates images that can be displayed by a set of displays, and a second image processing pipeline generates a set of output images that can be stored for later viewing or shared with others. The first and second image processing pipelines may receive the same images as input, but utilize different image processing techniques tailored to their respective outputs.
[0051] At least some of the set of images are captured during a media capture event. Specifically, a set of cameras may be used to capture images in one or more shooting modes (e.g., a photo mode that can capture still images, a video mode that can capture video, a panoramic mode that can capture panoramic photos, a portrait mode that can capture still photos with artificial bokeh applied, etc.). Generally, during these modes, the device may display a camera user interface that shows a “live preview” (e.g., via a set of displays of electronic devices 100, 202 in Figures 1 and 2). The live preview is a stream of images captured by the set of cameras and presented in real time, and may represent the field of view (which may be a subset of the field of view of one of the cameras) that is being captured when the camera initiates a media capture event. In other words, the live preview allows the user to see what part of the scene is currently being captured and to decide when to capture a photo or video.
[0052] The live preview may be presented to the user as part of an augmented reality environment, which operates in pass-through mode. In these cases, at least a portion of the augmented reality environment includes a recreation of the user's physical environment, thereby enabling the user to perceive their physical environment and understand which parts of the scene will be captured during a media capture event. The augmented reality environment may include one or more graphical user elements or other features to help guide the user when capturing images, such as transitions between a recreation of the physical environment and virtual content or a modified representation of the physical environment, indicating the boundaries of the image being captured by the device.
[0053] When the augmented reality system initiates a media capture event, the set of cameras captures images and generates media according to the current photo mode (e.g., captures and generates still images when in photo mode, or captures and generates video when in video mode), which may then be stored locally on device 100 (e.g., in a non-temporary computer-readable storage device as part of memory 108), sent to a remote server for storage, or sent to another system (e.g., as part of a messaging or streaming application). In some cases, a frame buffer of images captured by the set of cameras may exist, and it should be understood that some images collected before the initiation of the media capture event may be used in generating the captured media to be stored or sent.
[0054] A media capture event may be initiated when an electronic device described herein receives a capture request. A capture request may be received under certain predetermined conditions (for example, a software application running on the device may automatically request, using appropriate user permissions, that a set of cameras initiate a media capture event when certain criteria are met), or when a user gives a command to initiate a media capture event command, such as by interacting with the shutter control unit on the user interface, pressing a designated button on the electronic device, or giving a voice command.
[0055] The processes described herein may be performed as part of a media capture event, may utilize a first image processing pipeline to generate a live preview image displayed to the user, and may utilize a second image processing pipeline to generate captured media to be stored or transmitted. For example, Figure 3 shows an exemplary image stream 300 containing a sequence of images of a scene captured by a set of cameras in a device (e.g., a set of cameras in electronic devices 100 and 202 in Figures 1 and 2). In embodiments where the set of cameras includes two cameras, the image stream 300 includes a first camera image stream 300a of image 302 captured by a first camera (e.g., a first camera 112a in electronic device 202 in Figure 2) and a second camera image stream 300b of image 304 captured by a second camera (e.g., a second camera 112b in electronic device 202 in Figure 2).
[0056] A set of cameras may continuously capture an image stream 300 while an electronic device is operating, or may initiate the capture of an image stream 300 in response to certain conditions (e.g., the device entering pass-through mode). When the image stream 300 includes multiple camera image streams (e.g., first and second camera image streams 300a, 300b), images from each of the camera streams (e.g., images captured simultaneously from different cameras) may be grouped and processed together. For example, the image stream 300 may include a series of image pairs, each image pair including a first image captured by a first camera (e.g., an image from the first camera image stream 300a) and a second image captured by a second camera (e.g., an image from the second camera image stream 300b).
[0057] For example, Figure 3 shows the image stream 300 as containing a total of 16 images (eight images 302 labeled L1-L8 in the first camera image stream 300a, and eight images 304 labeled R1-R8 in the second camera image stream 300b), and thus having eight image pairs (a first image pair containing images L1 and R1, a second image pair containing images L2 and R2, etc.). The set of cameras can continue capturing the image stream 300 for as long as necessary to perform any desired media capture event. A predetermined number of images (or image pairs) from the image stream 300 may be temporarily stored in a buffer at a given time (e.g., a new image replaces the oldest image in the buffer), thereby making a certain number of recently captured images available to the device at any given time during the capture of the image stream 300.
[0058] The process described herein may select a first set 306 (or pair of images) from the image stream 300, and a set of live preview images is generated using the first set 306, as will be described in more detail below. It should be understood that when the process described herein is used to select a set of images for additional image processing (e.g., to generate images from them), the process does not need to wait for all images to be selected before starting this image processing. As a non-limiting example, using the first set 306, the device may begin generating live preview images from images L1 and R1 before other images from the first set 306 (such as images L7, L8, R7, and R8) are captured. In other words, the first image processing pipeline may continuously generate live preview images.
[0059] As shown in Figure 3, a capture request 308 may be received during the capture of an image stream (for example, as shown there, it may be received during the capture of images L4 and R4). In response to receiving the capture request 308, a second set 310 of images or pairs of images may be selected as part of the media capture event. The second set 310 of images may be used (for example, by a second image processing pipeline) to generate a set of output images that form the captured media to be stored or transmitted as described herein. In some cases, as described above, the second set 310 of images may include one or more images or pairs of images (for example, images L3 and R3) that were captured before the capture request 308. The number of images 310 selected to be in the second set 310 may be predetermined (for example, a particular photo-taking mode may select a specific number of images associated with each capture request 308) or may be determined dynamically (for example, in video mode, the number of images in the second set 310 may depend on when a termination request is received to stop the media capture event, thereby allowing the user to capture videos of varying lengths). In some variants, receiving capture request 308 may optionally prompt the camera set to temporarily change one or more operating parameters while capturing images. For example, in response to receiving capture request 308, the camera set may change the frame rate at which images are captured, or intentionally change the exposure time of one or more images (e.g., an image pair containing images L5 and R5 is shown to have been captured with a longer exposure time than the surrounding image pair).
[0060] Some or all of the images used to generate output images as part of a media capture event (e.g., a second set of images 310) may also be used to generate live preview images displayed to the user (e.g., via a first set of images 306). Thus, different image processing pipelines may receive the same images as input but generate multiple different images using different image processing techniques. For example, Figure 4 shows a variation of process 400 that uses multiple image processing pipelines to generate a first set of transformed images for display to the user (e.g., as part of a live preview in an augmented reality environment) and a second set of transformed images used to generate one or more output images as part of a media capture event. Blocks executed as part of process 400 may be executed as methods or stored as instructions on a non-temporary computer-readable storage device, and as a result, a processor may utilize these instructions to execute various blocks of the process described herein. Similarly, the device described herein includes memory (e.g., memory 108) and one or more processors (e.g., processor 106) operably coupled to the memory, the one or more processors being configured to execute instructions causing one or more processors to execute blocks of process 400.
[0061] Specifically, process 400 captures an image stream 402 using a set of cameras 404. In some cases, the set of cameras 404 includes a first camera 404a and a second camera 404b, and the image stream 402 includes an image 402a captured by the first camera 404a and an image 402b captured by the second camera 404b. In these cases, the image stream 402 may form a series of image pairs, each image pair including a first image captured by the first camera 404a and a second image captured by the second camera 404b.
[0062] Process 400 includes a display image processing pipeline 406 configured to generate a first set of transformed images 408 from images in image stream 300 (for example, a first set 306 of images from image stream 402 in Figure 3A). As part of generating the first set of transformed images 408, the display image processing pipeline 406 may apply a first transformation operation to the images in image stream 402. In some cases, this first transformation operation includes a perspective correction operation in which the images are transformed based on the difference between the viewpoint of one of the set of cameras 404 and the user's viewpoint. In these cases, the display image processing pipeline 406 may receive and utilize viewpoint data 410 representing information about the set of cameras 404 and the user's viewpoint. The perspective correction operation is described in more detail herein with reference to Figure 5.
[0063] For each image in the image stream 402 input to the display image processing pipeline 406, the display image processing pipeline 406 may generate a corresponding transformed image of a first set of transformed images 408. For example, an image 402a captured by a first camera 404a may be used to generate a first subset 408a of these transformed images 408, and an image 402b captured by a second camera 404b may be used to generate a second subset 408a of the transformed images 408. Thus, the first set of transformed images 408 may include a first set of transformed image pairs, and each image pair in the image stream 402 is used to generate a corresponding transformed image pair of the first set of transformed image pairs.
[0064] In step 412, the first set of transformed images is displayed on a set of displays (e.g., a set of 110 displays of the electronic device 100 in Figure 1). The first set of transformed images may also be displayed as part of an augmented reality environment (e.g., in pass-through mode), which may provide a live preview as described herein. If the set of displays includes a first display and a second display (e.g., the first and second displays 110a, 110b of the electronic device 202 in Figure 2), the first subset 408a of the transformed image 408 may be displayed on the first display, and the second subset 408b of the transformed image 408 may be displayed on the second display. In this way, the user's first eye can see the image captured by the first camera (via the first display), and the user's second eye can see the image captured by the second camera (via the second display). If the set of displays includes a single display, each pair of transformed images may be displayed simultaneously on the single display. For example, the first transformed image of a pair may be displayed on a first portion of the display (which can be seen by the user's first eye), and the second transformed image of a pair may be displayed on a second portion of the display (which can be seen by the user's second eye). In some cases, the display image processing pipeline 406 may add virtual content to the transformed image so that the user can see this virtual content when the transformed image is displayed to the user in block 412.
[0065] In addition, a capture request may be received while the image stream 402 is being captured. In block 414, in response to receiving the capture request, a set of images 416 is selected from the image stream 300 (for example, a second set of images 310 from the image stream 402 in Figure 3). The selected set of images 416 includes one or more images captured after the capture request was received, and in some cases may also include one or more images captured before the capture request was received, as described above with respect to Figure 3. The selected set of images 416 may include a first subset 416a of images captured by the first camera 404a (for example, selected from image 402a) and a second subset 416b of images captured by the second camera 404b (for example, selected from image 402b). Therefore, the selected set of images 416 may include a selected set of image pairs, where each image pair includes a first image captured by a first camera 404a and a second image captured by a second camera 404b.
[0066] Process 400 may utilize an output image processing pipeline 418 configured to generate a set of output images 420 using a selected set of images 416. In some cases, the set of output images 420 includes a first subset 420a of output images generated from a first selected subset 416a of images, and a second subset 420b of output images generated from a second subset 416b of images. Thus, the set of output images 420 may include a set of output image pairs. These output image pairs may include one or more stereo images and / or stereo videos. In stereoscopic imaging, two images taken from different perspectives, when presented to different eyes of the user, can allow the user to perceive depth. Thus, a stereo image represents a single pair of images that can later be presented to the user as a three-dimensional still image, while a stereo video represents a series of image pairs that can later be presented to the user sequentially as a three-dimensional video. Depending on the photo-taking mode, the set of output images 420 may include a single stereo image, multiple stereo images (e.g., generated as part of a burst photo-taking mode), stereo video only, or a combination of stereo images and stereo video. If the set of output images 420 includes both stereo images and stereo video, the stereo image may be presented to the user as a three-dimensional image, but the stereo video clip associated with the stereo image may also be played.
[0067] As part of generating a set of output images 420, the output image processing pipeline 418 may generate a second set of transformed images (not shown) from a selected set of images 416. The second set of transformed images may be used to generate the output images 420. In some cases, the output image processing pipeline 418 applies a second transformation operation to the images in the image stream 402 that is different from the first transformation operation applied by the display image processing pipeline 406. In this way, the display image processing pipeline and the output image processing pipeline may sometimes receive the same images for processing, but they process these input images separately (e.g., in parallel) to provide different outputs. In some variations, the second transformation operation includes an image modification operation, as will be described in more detail with respect to Figure 6.
[0068] In some cases, the output image processing pipeline 418 is also configured to generate metadata 422 associated with the set of output images 420. This metadata (an example of which is described in more detail below with respect to Figure 7) may be stored along with the set of output images 420. The output image processing pipeline 418 may also receive context information 424 in addition to the selected set of images 416. The context information 424 may include information about the device (or its components, such as a set of cameras) and / or the scene while the image stream 404 is being captured.
[0069] For example, in some variant forms, context information 424 includes a depth map 426. The depth map 426 includes a matrix of pixels, each pixel corresponding to a distinct location in the scene captured by a set of cameras and having a distinct depth value representing the distance between the device (or a part thereof, such as an individual camera in the set of cameras) and the distinct location in the scene. In some cases, the depth map 426 may be derived from depth information generated by a depth sensor, in cases where the device includes a depth sensor (e.g., depth sensor 114 of electronic devices 100 and 202 in Figures 1 and 2). As an addition or alternative, the depth map 426 may be derived from an image analysis of images (or multiple images) captured by one or more cameras of the device, such as one or more images in an image stream 402. In these cases, the images (one or more) may be analyzed using monocular or stereoscopic depth estimation techniques to estimate depth information about different parts of the scene. It should also be understood that the output image processing pipeline 418 may receive multiple depth maps, such as a first depth map for the first camera 404a (representing the scene distance to the first camera 404a) and a second depth map for the second camera 404b (representing the scene distance to the second camera 404b).
[0070] As an addition or alternative, context information 424 may include information about the content of the scene captured by the selected set of images 416. For example, context information 424 may include a scene information map 428. Specifically, the scene information map 428 includes a matrix of pixels, each pixel corresponding to a distinct location in the scene and having a distinct identifier (and optionally a confidence value associated with the identifier) indicating the type of object present at that location. For example, the scene information map 428 may be configured to identify locations in the scene that contain human faces, such that each pixel of the scene information map indicates whether a face is present at a particular location in the scene, or includes a value (such as a normalized value between 0 and 1) representing the confidence that a face is present at a particular location in the scene. In some variations, context information 424 may include multiple scene information maps (each associated with a different type of object) or a single scene information map 428 that may indicate the presence of multiple types of objects in the scene (e.g., people, animals, trees, etc.).
[0071] Specifically, the scene information map 428 (or more maps) may provide an understanding of the scene's content, which can be used when generating a set of output images 420 and / or metadata 422. For example, the output image processing pipeline may be configured to apply different image processing techniques to people in the scene compared to other objects present in the scene. Additionally or alternatively, the scene information map 428 (or more maps) may be used to select a default parallax value, which is included in the metadata 422 and can be used later (for example, when displaying the output images) to control the parallax at which the output images are displayed (for example, in a display area in an augmented reality environment).
[0072] The scene information map 428 can be generated in any preferred manner. For example, as will be readily understood by those skilled in the art, some or all of the selected set of images 416 and / or the depth map 426 can be analyzed using image segmentation and / or image classification techniques. These techniques can be used to find a given set of objects in the scene, the selection of which may be set based on user preferences, the photography mode that is active when the selected set of images 416 is captured, and / or a desired effect selected by the user that is applied to generate the set of output images 420. In some variations, the scene information map 428 may be based at least in part on user input. For example, the user can identify a region of interest in the scene (e.g., by determining that the user's gaze is related to the scene, by determining that the user's gestures are related to the scene, other user inputs, a combination thereof, etc.), and identify one or more objects within the region of interest to generate the scene information map 428.
[0073] Additionally or alternatively, context information 424 may include pose information 430 associated with a device (or one or more of its components). The pose information 430 may include the position and / or orientation of the device during the capture of a selected set of images 416. This may be used to generate similar pose information associated with a set of output images 420. The pose information associated with a set of output images 420 may be stored as metadata 422 and used later (for example, when displaying the output images) to control the position and / or orientation in which the output images are displayed (for example, in an augmented reality environment).
[0074] In block 432, the set of output images 420 (including metadata 422, if the set of output images 420 is associated with metadata 422) may be stored in a non-temporary computer-readable storage device (for example, as part of memory 108 of electronic device 100 in Figure 1). Additionally or alternatively, the set of output images 420 may be transmitted to a different device (for example, as part of an email, message, live stream, etc., or for remote storage). As part of storing and / or transmitting the set of output images 420 (including metadata 422, if the set of output images 420 is associated with metadata 422) may be encoded or written in a specified format. As an unrestricted example, stereo images may be written as a High Efficiency Image File format, such as a High Efficiency Image Container (HEIC) file. As another unrestricted example, stereo video may be written as a High Efficiency Video Coding (HEVC) file, such as a Multiview High Efficiency Video Coding (MV-HEVC) file.
[0075] By separating the processing of the image stream 402 between the display image processing pipeline 406 and the output image processing pipeline 418, these image processing pipelines can be coordinated to produce different types of images. Thus, in block 412, the process can generate a first set of transformed images 408 in a manner coordinated for real-time display of these images (e.g., as part of a live preview and / or pass-through mode in an augmented reality environment). These images may be discarded after being presented to the user, or they may be optionally stored as part of block 432 in a non-temporary computer-readable storage device (e.g., as part of memory 108 of electronic device 100 in Figure 1). In addition, process 400 can generate a set of output images 420 that can be coordinated to different intended viewing experiences. For example, a set of output images 420 may be intended to be viewed on different devices (having the same or different configuration as the device used to run process 400) and / or as part of a particular user experience with different viewing requirements than those of block 412.
[0076] It should be understood that various blocks of process 400 can occur simultaneously and continuously. For example, a particular image of the first set of transformed images 408 may be displayed in block 412 at the same time that the display image processing pipeline 406 is generating other images of the first set of transformed images 408. This may also occur at the same time that the first and second cameras 404a, 404b are capturing images to be added to the image stream 402, and while the output image processing pipeline 418 is generating the output image 420 as described herein.
[0077] Figure 5 shows a variant of the display image processing pipeline operation 406 of process 400. As shown therein, the display image processing pipeline 406 includes a perspective correction operation 500 that generates a set of perspective-corrected images (e.g., images from image stream 402) from a set of images received by the display image processing pipeline 406. Generally, a perspective correction operation takes an image captured from one viewpoint ("source viewpoint") and transforms it so that it appears as if it were captured from a different viewpoint ("target viewpoint"). Specifically, the perspective correction operation can use the user's viewpoint as the target viewpoint and thus transforms the image so that it appears as if it were captured from the user's viewpoint. In this way, when the perspective-corrected images are displayed to the user in block 412 of process 400, objects presented in these images may appear at the same height and distance as real-world objects in the user's physical environment. This can improve the accuracy of the reproduction of the user's physical environment in the augmented reality environment as part of passthrough and / or live preview mode.
[0078] The perspective correction operation 500 may use the viewpoint data 410 to determine the source viewpoint (e.g., the viewpoint of a camera used to capture an image) and the target field of view (e.g., the user's viewpoint or a specific eye of the user). For example, the viewpoint data 410 may include position information including one or more distances and / or angles (e.g., the first and second distances 220a, 220b described with respect to Figure 2) that represent the spatial offset and / or rotational offset between the set of cameras and the set of displays, respectively. In some cases, the offset between the set of cameras and the set of displays may be known in advance (e.g., if the set of cameras and the set of displays have a fixed spatial relationship), (e.g., if one or more of the set of cameras and / or the set of displays are movable within the device), or may be determined dynamically by the device.
[0079] The viewpoint data 410 may include information about the spatial relationship between the user and the device (or its components). This may include a single set of values representing the relative position of the user's eyes to the device, or it may include different sets of values for each eye (e.g., the first and second eye-to-display distances 222a, 222b as described with respect to Figure 2). The relative position between the user and the device may be determined, for example, using an eye-tracker (e.g., the eye-tracker 116 in electronic devices 100 and 202 in Figures 1 and 2). Similarly, the viewpoint data 410 may include information about the scene being captured by a set of cameras (e.g., the distance from the device to a specific object in the scene). This may include distance information determined by a depth sensor (e.g., the depth sensor 114 in electronic devices 100 and 202 in Figures 1 and 2) by one or more of the set of cameras, a combination thereof, etc.
[0080] Collectively, viewpoint data 410 can be used to determine the difference between a source viewpoint and a target viewpoint for a particular image. Perspective correction operation 500 may use the difference between the source viewpoint and the target viewpoint to apply a transformation (e.g., a homography transformation or a warping function) to the input image (e.g., the image in the image stream 402) to generate a perspective-corrected image. Specifically, for each pixel of the input image at its pixel location in the untransformed space, a new pixel location in the perspective-corrected image is determined in the transformed space of the transformed image. Examples of perspective correction operations are discussed in more detail in U.S. Patent No. 1,0832427, entitled "Scene camera retargeting," and U.S. Patent No. 1,1373271, entitled "Adaptive image warping based on object and distance information," the contents of which are incorporated herein by reference in their entirety and attached herein by reference.
[0081] If the first set of images includes a set of image pairs, the perspective correction operation 500 includes a first perspective correction operation 500a and a second perspective correction operation 500b. For each image pair, the first perspective correction operation 500a generates a first perspective-corrected image from the first image of the image pair. Similarly, the second perspective correction operation 500b generates a second perspective-corrected image from the second image of the image pair. Specifically, the first perspective correction operation 500a may be applied to an image 402a captured by a first camera 404a and may generate a perspective-corrected image based on the difference between the viewpoint of the first camera 404a and the viewpoint of the user (e.g., the viewpoint of the user's first eye). Similarly, a second perspective correction operation 500b may be applied to an image 402b captured by a second camera 404b, and may generate a perspective-corrected image based on the difference between the viewpoint of the second camera 404b and the user's viewpoint (e.g., the viewpoint of the user's second eye).
[0082] The perspective-corrected images generated by the perspective correction operation 500 may form a first set of transformed images 408. In some cases, additional processing steps may be applied to these images. For example, in block 502, the display image processing pipeline 406 may add virtual content, such as one or more graphical elements of a graphical user interface or virtual objects, to the first set of transformed images 408. This may enable the process to generate an augmented reality environment with virtual content, and enable the user to experience the virtual content when the first set of transformed images 408 is displayed to the user in block 412 of process 500. It should be understood that the display image processing pipeline 406 may perform additional processing steps, such as cropping a portion of an image, as part of generating the first set of transformed images 408.
[0083] Figure 6 shows a modified form of the output image processing pipeline 418 of process 400. As shown therein, the output image processing pipeline 418 includes a transformation unit 600 that performs transformation operations on images input to the output image processing pipeline 418 to generate transformed images. Thus, the output image processing pipeline may be configured to generate a set of transformed images (for example, a second set of transformed images different from a first set of transformed images 408 generated by the display image processing pipeline 406) from a selected set of images 416. The image generation unit 602 generates a set of output images 420 from this set of transformed images. The transformation operations performed by the transformation unit 600 may differ from the perspective correction operations 500 used in the display image processing pipeline operations 406. For example, in some modified forms, the transformation unit 600 applies image correction operations that transform the images and project these images onto a common target image plane. In the embodiment of Figure 6, the transformation unit 600 includes a dewarp operation 604 and an alignment operation 606 that collectively perform image correction operations. In some cases, a set of cameras may have a fisheye lens or other similar optical element designed to increase the field of view of these cameras. Images captured by these cameras may be useful for recreating a three-dimensional environment (for example, as part of the augmented reality environment described herein), but these images typically contain distortions when viewed as two-dimensional images (for example, the proportions of objects may be altered, and straight lines in the physical environment may appear as curves in the captured image).
[0084] Therefore, as will be readily understood by those skilled in the art, the dewarp operation 604 may generate a dewarped image to account for lens distortion. In other words, the image is dewarped from its physical lens projection to a linear projection such that straight lines in three-dimensional space are mapped to straight lines in a two-dimensional image. For each image in the selected set of images 416 received by the output image processing pipeline 418, a first dewarp operation 604a may generate a first dewarped image from a first image of the image pair (e.g., one of a first subset 416a of images captured by the first camera 404a), and a second dewarp operation 604a may generate a second dewarped image from a second image of the image pair (e.g., one of a second subset 416b of images captured by the second camera 404b). As long as the first and second cameras 404a and 404b have different configurations (e.g., different fields of view and / or different lens designs), the first and second dewarp operations 604a and 604b may be adjusted to accommodate the differences between these cameras.
[0085] For each pair of images in the selected set of images 416, the alignment operation 606 may align the first and second dewarped images generated by the dewarped operation 604 so that the epipolar lines of the dewarped images are aligned horizontally. In this way, the transformation unit 600 may generate a transformed image pair for each pair of images it receives that emulates having a common image plane (even if they can be captured by cameras having different image planes). These transformed image pairs may be used to generate the output image 420 via the image generation unit 602.
[0086] Figure 7 shows a modified image generation unit 602 configured to generate a set of output images 420 as part of process 400. Optionally, the image generation unit 602 may be configured to generate metadata 422 associated with the set of output images 420. As shown in Figure 7, the image generation unit 602 includes a still image processing unit 700 configured to generate a set of still images, a video processing unit 702 configured for video, and a metadata processing unit 704 that generates metadata associated with the output images. For a given media capture event, the output image processing pipeline 418 uses some or all of these when generating a set of output images 420, depending on the photo-taking mode.
[0087] The still image processing unit 700 is configured to receive a first group of transformed images 701 (which may be some or all of the transformed images generated by the transformation unit 600) and can apply an image fusion operation 706 to generate one or more fused images 708 from two or more of the first group of transformed images 701. If the transformed images 701 includes a set of transformed image pairs, the image fusion operation 706 can generate one or more fused image pairs to form one or more fused stereo output images. In these cases, the first fused image 708a of the fused image pair can be generated by applying the first image fusion operation 706a to a first subset of the transformed images 701 (e.g., all generated from image 416a captured by the first camera 404a). The second fused image 708b of the fused image pair can be generated by applying the second image fusion operation 706b to a second subset of the transformed images 701 (e.g., all generated from image 416b captured by the second camera 404b). Generally, an image fusion operation includes the steps of selecting a reference image from a group of selected images and modifying some or all of its pixel values using information from other selected images. The image fusion operation 706 may produce a fused output image having reduced noise and / or a higher dynamic range, and may utilize any suitable image fusion technique, as will be readily understood by those skilled in the art.
[0088] The video processing unit 702 is configured to receive a second group of converted images 703 (which may be some or all of the converted images generated by the conversion unit 600) and may perform a video stabilization operation 710 on the second group of converted images 703 to generate a stabilized video 712. If the second group of converted images 703 includes a set of image pairs, the stabilized video 712 may be a stereo video including a first set of images 712a (e.g., generated from image 416a captured by the first camera 404a) and a second set of images 712b (e.g., generated from image 416b captured by the second camera 404b). Thus, the video processing unit 702 may output a stereo video. If the set output image 420 includes both a fused output image and a stereo video, the video processing unit 702 may optionally utilize one or more fused images 708 as part of the video stabilization operation 710 (e.g., stabilizing frames of the stereo video against a fused stereo image). The video stabilization operation 710 may utilize any suitable electronic image stabilization technique, as will be readily understood by those skilled in the art. In some cases, the video stabilization operation 710 may receive contextual information 424 (e.g., scene information, orientation information, etc.) as described above, which may assist in performing video stabilization.
[0089] In some variations, the image generation unit 602 may optionally include an additional content unit 730 configured to add additional content to still images generated by the still image processing unit 700 and / or videos generated by the video processing unit 702. For example, in some cases, the additional content unit 730 may apply virtual content 732 (e.g., one or more virtual objects) and / or other effects 734 (e.g., filters, style transfer operations, etc.) to the output images generated by the image generation unit 602.
[0090] The metadata processing unit 704 is configured to generate metadata 422 associated with the set of output images 420. Depending on the type of metadata 422 being generated, the metadata processing unit 704 may receive a third group of images 705 (which may include some or all of the transformed images generated by the transformation unit 600, the output images generated by the transformation unit 600, and / or images captured or generated as part of process 400). In addition, or alternatively, the metadata processing unit 704 may receive contextual information 424, such as those described herein, which may be used to generate the metadata 422.
[0091] In some variations, the metadata processing unit 704 is configured to generate field of view information 714 relating to a set of cameras used to capture a selected set of images 416 used to generate a set of output images 420. This field of view information 714 may be fixed (e.g., if the field of view of the set of cameras is fixed), (e.g., if one or more of the set of cameras have an optical zoom function that provides a variable field of view), or dynamically determined. Additionally or alternatively, the metadata processing unit 704 is configured to generate baseline information 716 representing one or more separation distances between the cameras of the set of cameras (e.g., the separation distance between a first camera 404a and a second camera 404b). Again, this information may be fixed (e.g., if the set of cameras are in fixed positions relative to each other), or dynamically determined (e.g., if one or more of the set of cameras are movable with the device housing the set of cameras).
[0092] In some cases, the metadata processing unit 718 is configured to generate pose information 718 associated with a set of output images. Specifically, the pose information 718 may describe the pose of the output images, thereby describing the relative position and / or orientation of the scene captured by the set of output images 420. This information may include, or be derived from, pose information 430 associated with the underlying images (e.g., a selected set of images 416) used to generate the output images 420.
[0093] In some cases, the metadata processing unit 718 is configured to set a default parallax value 720 for a set of stereo output images. The default parallax value for a given stereo image represents the recommended parallax at which the images of the stereo image are presented when the stereo image is viewed (e.g., the distance at which these images are separated). The step of changing the parallax at which the stereo image is presented will affect how the user perceives the depth of the scene and can therefore be used to highlight different parts of the scene captured by the set of output images 420.
[0094] In some cases, it may be desirable to select a default set of disparity values based on the scene captured by the set of output images 420. In these cases, context information 424, such as a scene information map 428, a depth map 426, and / or pose information 430, may be used to select the default set of disparity values. This context information 424 may be associated with the set of output images 420 (e.g., generated from images or information captured simultaneously with the selected images 416 used to generate the set of output images 420). This context information 424 may enable the metadata processing unit 704 to determine one or more aspects of the scene captured by the set of output images 420 (e.g., the presence and / or location of one or more objects in the scene) and select default disparity values based on these determinations.
[0095] The metadata 422 generated by the metadata processing unit 704 may include information applicable to all images in the set of output images, and / or information specific to a particular image or pair of images. For example, the metadata 422 may include pose information 718 for one or more images in the output images 420, but may also include a single set of baseline information 716 applicable to all of the output images 420. Similarly, the set of default disparity values 720 may include a single value applicable to all pairs of images in the set of output images 420, or it may include different values for different pairs of images (for example, a first stereo image may have a first default disparity value, and a second stereo image may have a second default disparity value).
[0096] In the preceding description, certain technical terms have been used for illustrative purposes to provide a complete understanding of the described embodiments. However, it will be apparent to those skilled in the art that specific details are not required to practice the described embodiments. Accordingly, the preceding descriptions of the specific embodiments described herein are presented for illustrative and explanatory purposes only. They are not intended to be exhaustive or to limit embodiments to the exact form disclosed. It will be apparent to those skilled in the art that many modifications and variations are possible in light of the above teachings.
[0097] [Item 1] It is a device, Camera set, A set of displays, Memory and The memory comprises one or more processors operably coupled to the memory, and the one or more processors are configured to provide the one or more processors The aforementioned set of cameras is used to capture an image stream. By applying a first transformation operation to the images in the image stream, a first set of transformed images is generated from the image stream. The set of displays is used to display the first set of the converted images. Receive the capture request, The system is configured to execute an instruction to generate a set of stereo output images in response to receiving the capture request, and the step of generating the set of stereo output images is: The steps include selecting a set of images from the aforementioned image stream, A step of generating a second set of converted images from the set of images by applying a second conversion operation, which is different from the first conversion operation, A device comprising the step of generating a set of stereo output images using a second set of the converted images. [Item 2] The step of generating the set of stereo output images is: The device according to item 1, comprising the step of generating a fused stereo output image from two or more pairs of images from the set of images. [Item 3] The device described in item 1, wherein the set of stereo output images includes stereo video. [Item 4] The step of generating the set of stereo output images is: The device according to item 3, comprising the step of performing a video stabilization operation on a second set of the converted images. [Item 5] A device according to any one of items 1 to 4, wherein the first conversion operation is a perspective correction operation. [Item 6] A device according to any one of items 1 to 5, wherein the second conversion operation is an image correction operation. [Item 7] The device according to any one of items 1 to 6, wherein the step of generating the set of stereo output images includes the step of generating metadata associated with the set of stereo output images. [Item 8] The device according to item 7, wherein the metadata includes field of view information of the set of cameras. [Item 9] The device according to item 7 or 8, wherein the metadata includes pose information for the set of stereo output images. [Item 10] A device according to any one of items 7 to 9, wherein the metadata includes a set of default disparity values for the set of stereo output images. [Item 11] The device according to item 10, wherein the step of generating metadata associated with the set of stereo output images includes the step of selecting the set of default disparity values based on the scene captured by the set of stereo output images. [Item 12] It is a method, The steps include capturing an image stream using the device's camera set, The steps include: generating a first set of transformed images from the image stream by applying a first transformation operation to the images in the image stream; The steps include displaying a first set of the converted images in a set of device displays, The steps include receiving a capture request and The step of generating a set of stereo output images in response to receiving the capture request, wherein the step of generating a set of stereo output images is The steps include selecting a set of images from the aforementioned image stream, A step of generating a second set of converted images from the set of images by applying a second conversion operation, which is different from the first conversion operation, A method comprising the step of generating a set of stereo output images using a second set of the converted images. [Item 13] The step of generating the set of stereo output images is: The method according to item 12, comprising the step of generating a fused stereo output image from two or more pairs of images from the set of images. [Item 14] The method according to item 12, wherein the set of stereo output images includes stereo video. [Item 15] The step of generating the set of stereo output images is: The method according to item 14, further comprising the step of performing a video stabilization operation on a second set of the converted images. [Item 16] The method according to any one of items 12 to 15, wherein the first conversion operation is a perspective correction operation. [Item 17] The method according to any one of items 12 to 16, wherein the second conversion operation is an image correction operation. [Item 18] The method according to any one of items 12 to 17, wherein the step of generating the set of stereo output images includes the step of generating metadata associated with the set of stereo output images. [Item 19] The method according to item 18, wherein the metadata includes field of view information of the set of cameras. [Item 20] The method according to item 18 or 19, wherein the metadata includes pose information of the set of stereo output images. [Item 21] The metadata is as described in any one of items 18 to 20, wherein the metadata includes a set of default disparity values for the set of stereo output images. [Item 22] The method according to item 21, wherein the step of generating metadata associated with the set of stereo output images includes the step of selecting the set of default disparity values based on the scene captured by the set of stereo output images. [Item 23] The method according to any one of items 12 to 22, comprising the step of adding virtual content to the first set of converted images. [Item 24] A non-temporary computer-readable medium containing instructions, wherein, when the instructions are executed by at least one computing device, the at least one computing device causes the computing device to perform an operation including the steps described in any one of items 12 to 23. [Item 25] It is a device, First camera A second camera, A set of displays, Memory and The memory comprises one or more processors operably coupled to the memory, and the one or more processors are configured to provide the one or more processors A step of capturing a first set of image pairs of a scene using the first camera and the second camera, wherein each image pair in the first set of image pairs is The first image captured by the first camera, A capturing step including a second image captured by the second camera, A step of generating a first set of transformed image pairs from the first set of image pairs, wherein the step of generating the first set of transformed image pairs includes, for each image pair in the first set of image pairs, A step of generating a first perspective-corrected image from the first image based on the difference between the viewpoint of the first camera and the viewpoint of the user, A generating step includes the step of generating a second perspective-corrected image from the second image based on the difference between the viewpoint of the second camera and the viewpoint of the user, The steps include displaying the first set of the converted image pairs on the set of displays, The steps include receiving a capture request and In response to receiving the capture request, the steps include selecting a second set of image pairs from a first set of image pairs, A step of generating a second set of transformed image pairs from the second set of the aforementioned image pairs, wherein the step of generating the second set of transformed image pairs is performed for each image pair of the first set of the aforementioned image pairs, The steps include generating a first dewarped image from the first image, The steps include generating a second dewarped image from a second first image, A generation step including the step of aligning the first and second dewarped images, A device that performs the steps of generating a set of output images using a second set of the aforementioned converted image pairs. [Item 26] The step of generating the set of output images is, The device according to item 25, comprising the step of generating a fused stereo output image from two or more image pairs from a second set of the aforementioned transformed image pairs. [Item 27] The device according to item 25, wherein the set of output images includes a stereo video formed from a second set of the converted image pairs. [Item 28] The step of generating the set of output images is, The device according to item 27, comprising the step of performing a video stabilization operation on a second set of the converted image pairs. [Item 29] The device according to any one of items 25 to 28, wherein the step of generating the set of output images includes the step of generating metadata associated with the set of output images. [Item 30] The device according to item 29, wherein the metadata includes field of view information from at least one of the first camera or the second camera. [Item 31] The device according to item 29 or item 30, wherein the metadata includes pose information for the set of output images. [Item 32] The set of output images includes a set of stereo images, The device according to any one of items 29 to 31, wherein the metadata includes a set of default disparity values for the set of stereo output images. [Item 33] The device according to item 32, wherein the step of generating metadata associated with the set of output images includes the step of selecting the set of default disparity values based on the scene captured by the set of output images. [Item 34] The device according to any one of items 25 to 33, wherein the processor is configured to add virtual content to the first set of the converted image pairs. [Item 35] It is a method, A step of capturing a first set of image pairs of a scene using the first camera and the second camera of the device, wherein each image pair in the first set of image pairs is The first image captured by the first camera, A capturing step including a second image captured by the second camera, A step of generating a first set of transformed image pairs from the first set of image pairs, wherein the step of generating the first set of transformed image pairs includes, for each image pair in the first set of image pairs, A step of generating a first perspective-corrected image from the first image based on the difference between the viewpoint of the first camera and the viewpoint of the user, A generating step includes the step of generating a second perspective-corrected image from the first image based on the difference between the viewpoint of the second camera and the viewpoint of the user, The steps include displaying the first set of the converted image pairs on a set of displays of the device, The steps include receiving a capture request and In response to receiving the capture request, the steps include selecting a second set of image pairs from a first set of image pairs, A step of generating a second set of transformed image pairs from the second set of image pairs, wherein for each image pair in the first set of image pairs, The steps include generating a first dewarped image from the first image, The steps include generating a second dewarped image from a second first image, A generation step including the step of aligning the first and second dewarped images, A method comprising the step of generating a set of output images using a second set of the aforementioned transformed image pairs. [Item 36] The step of generating the set of output images is, The method according to item 35, comprising the step of generating a fused stereo output image from two or more image pairs from a second set of the converted image pairs. [Item 37] The method according to item 35, wherein the set of output images includes a stereo video formed from a second set of the converted image pairs. [Item 38] The step of generating the set of output images is, The method according to item 37, further comprising the step of performing a video stabilization operation on a second set of the converted image pairs. [Item 39] The method according to any one of items 35 to 38, wherein the step of generating the set of output images includes the step of generating metadata associated with the set of output images. [Item 40] The method according to item 39, wherein the metadata includes field of view information from at least one of the first camera or the second camera. [Item 41] The method according to item 39 or 40, wherein the metadata includes pose information for the set of output images. [Item 42] The set of output images includes a set of stereo images, The method according to any one of items 39 to 41, wherein the metadata includes a set of default disparity values for the set of stereo output images. [Item 43] The method according to item 42, wherein the step of generating metadata associated with the set of output images includes the step of selecting the set of default disparity values based on the scene captured by the set of output images. [Item 44] The method according to any one of items 35 to 43, comprising the step of adding virtual content to a first set of the converted image pairs. [Item 45] A non-temporary computer-readable medium containing instructions, wherein, when the instructions are executed by at least one computing device, the at least one computing device causes the computing device to perform an operation including the steps described in any one of items 35 to 44.
Claims
1. It is a device, Camera set, A set of displays, Memory and The memory comprises one or more processors operably coupled to the memory, and the one or more processors are configured to provide the one or more processors The aforementioned set of cameras is used to capture an image stream. A first set of transformed images is generated from the image stream by applying a perspective correction operation to the images in the image stream, which transforms the viewpoint of the images in the image stream. The set of displays is used to display the first set of the converted images. Receive the capture request, The system is configured to execute an instruction to generate a set of stereo output images in response to receiving the capture request, and the step of generating the set of stereo output images is: The steps include selecting a set of images from the aforementioned image stream, A step of generating a second set of transformed images from the set of images by applying an image correction operation, different from the perspective correction operation, which projects the images of the image stream onto a target image plane, to the set of images; A device comprising the step of generating a set of stereo output images using a second set of the converted images.
2. The step of generating the set of stereo output images is: The device according to claim 1, comprising the step of generating a fused stereo output image from two or more pairs of images from the set of images.
3. The device according to claim 1, wherein the set of stereo output images includes stereo video.
4. The step of generating the set of stereo output images is: The device according to claim 3, further comprising the step of performing a video stabilization operation on a second set of the converted images.
5. The device according to any one of claims 1 to 4, wherein the step of generating the set of stereo output images includes the step of generating metadata associated with the set of stereo output images.
6. The device according to claim 5, wherein the metadata includes field of view information of the set of cameras.
7. The device according to claim 5, wherein the metadata includes posture information for the set of stereo output images.
8. The device according to claim 5, wherein the metadata includes a set of default disparity values for the set of stereo output images.
9. The device according to claim 8, wherein the step of generating metadata associated with the set of stereo output images includes the step of selecting the set of default disparity values based on the scene captured by the set of stereo output images.
10. It is a method, The steps include capturing an image stream using the device's camera set, The steps include generating a first set of transformed images from the image stream by applying a perspective correction operation to the images in the image stream to transform the viewpoint of the images in the image stream, The steps include displaying a first set of the converted images in a set of device displays, The steps include receiving a capture request and The step of generating a set of stereo output images in response to receiving the capture request, wherein the step of generating a set of stereo output images is The steps include selecting a set of images from the aforementioned image stream, A step of generating a second set of transformed images from the set of images by applying an image correction operation, different from the perspective correction operation, which projects the images of the image stream onto a target image plane, to the set of images; A method comprising the step of generating a set of stereo output images using a second set of the converted images.
11. The step of generating the set of stereo output images is: The method according to claim 10, further comprising the step of generating a fused stereo output image from two or more pairs of images from the set of images.
12. The method according to claim 10, wherein the set of stereo output images includes stereo video.
13. The step of generating the set of stereo output images is: The method according to claim 12, further comprising the step of performing a video stabilization operation on a second set of the converted images.
Citation Information
Patent Citations
Defect image generation method based on generative adversarial network
CN114022586A
Imaging apparatus
JP2008141518A
Camera body, imaging apparatus, camera body control method, program, and recording medium with program recorded thereon
JP2012075078A
Imaging device and method for controlling the same device
JP2014228727A
Camera rig and stereo image capture
JP2018518078A