Multi-Camera Video Stabilization

The multi-camera video stabilization system addresses inconsistencies in image stabilization by using EIS and digital zoom, transforming data into a canonical space for seamless transitions, achieving smooth and high-quality video capture.

JP7799608B2Active Publication Date: 2026-01-15GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022536617
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-29
Filing Date
2021-07-02
Publication Date
2026-01-15
Estimated Expiration
2041-07-02

AI Technical Summary

Technical Problem

Existing devices with multiple camera modules face challenges in maintaining image stabilization consistency across different zoom levels and camera transitions, leading to undesirable effects such as camera shake, monocular parallax, and visible stutter in recorded videos.

Method used

A multi-camera video stabilization system that uses electronic image stabilization (EIS) and multi-camera digital zoom, transforming image data into a canonical space to align and stabilize images across different cameras, ensuring seamless transitions and consistent stabilization parameters.

Benefits of technology

The system provides smooth and high-quality video capture by maintaining temporal and spatial continuity, reducing camera shake and transitions between cameras, while efficiently managing computational resources and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007799608000034
    Figure 0007799608000034
  • Figure 0007799608000035
    Figure 0007799608000035
  • Figure 0007799608000036
    Figure 0007799608000036
Patent Text Reader

Abstract

A method, system, and apparatus for multi-camera video stabilization, including a computer program encoded on a computer storage medium, is provided. In some implementations, a video capture device includes a first camera and a second camera. The video capture device provides a digital zoom function that allows user-specified magnification changes within a digital zoom range during video recording. The video capture device is configured to use video data from different cameras across different portions of the digital zoom range. The video capture device can process image data captured using the second camera by applying a set of transforms including: (i) a first transform to a canonical reference space for the second camera; (ii) a second transform to the canonical reference space for the first camera; and (iii) a third transform for applying electronic image stabilization to the image data in the canonical reference space for the first camera.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background Some devices, such as smartphones, include multiple camera modules. These cameras can be used to record still images or video. In many situations, camera shake and other movements of the device can degrade the quality of the captured images and video. As a result, some devices include image stabilization features to improve the quality of the recorded image data. Summary of the Invention [Problem to be solved by the invention]

[0002] overview In some implementations, a device includes a multi-view camera system, e.g., a device with multiple camera modules with different fields of view. The device provides video stabilization for video captured using one or more of the device's camera modules. This can include providing various levels of image stabilization at different zoom (e.g., magnification or magnification) levels and managing image stabilization behavior consistently across different zoom levels and different camera modules. Even if the camera modules have a fixed field of view, the device can provide zoom functionality by implementing or approximating a continuous or smooth zoom along the range of field of view using, for example, digital zoom techniques. The device can provide features for capturing consistently stabilized video across a range of zoom settings, including when transitioning between video captures of different camera modules during video recording. In some implementations, the device detects when to transition between cameras for image capture during video recording. The device performs the transition and processes the captured data to generate a seamless output over the duration of the transition, e.g., an output video that maintains or smoothly adjusts parameters such as field of view, image stabilization, exposure, focus, and noise. This enables the system to generate video that uses different portions of video from different cameras, with transitions between cameras that are unobtrusive to the user, e.g., without monocular parallax, visible stutter, visible zoom pauses, or other glitches.

[0003] The primary goal of this system is to improve the smoothness of the video scene presented to the user. In other words, the system attempts to achieve both temporal continuity of the scene (e.g., reducing undesirable camera shake over time) and spatial continuity of the scene (e.g., reducing differences between videos captured from different cameras). This involves two key technologies: (1) electronic image stabilization (EIS), which provides scene continuity over time on a single camera, effectively providing temporal smoothing of the scene shown in the video feed, and (2) enhanced multi-camera digital zoom (e.g., gradual, incremental, or substantially continuous zoom using the outputs of multiple cameras), which effectively provides a spatial smoothing method to provide scene continuity across different cameras while avoiding clutter or interruptions near transitions between cameras.

[0004] EIS and multi-camera digital zoom can be efficiently combined using various techniques and transformations, discussed further below, including the use of a "canonical" camera space to represent image data. A canonical space can represent a conceptual camera view with fixed intrinsic properties that do not change over time. For example, a canonical space can be one that is unaffected by factors such as optical image stabilization (OIS) or voice coil motor (VCM) position and rolling shutter effects.

[0005] As an example, a device may include a first camera and a second camera, each providing a different view of a scene. The device may enable a zoom function that allows a user to smoothly change the zoom or magnification represented by the output video, even if one or more of the camera modules have a fixed view. In some implementations, two or more camera modules with fixed focal length lenses may be used to simulate continuous zoom across a zoom range by (1) using digital zoom (e.g., cropping and / or magnification) based on images from the first camera for a first portion of the range, and (2) using digital zoom based on images from the second camera for a second portion of the range. To improve the overall quality of the video, image stabilization processing may be dynamically adjusted relative to the current level of applied digital zoom. For example, each change in zoom along a simulated continuous zoom range may have a corresponding change in image stabilization parameters.

[0006] Image processing by the device can manage the transition between image captures from different cameras to provide a substantially seamless transition while maintaining consistency of EIS application and other image capture aspects such as focal length, exposure, etc. Zoom functionality can include, for example, digital zoom, which uses the output of a first camera as the device increasingly crops on the image captured by the first camera. Then, when a threshold level of zoom is reached and the zoomed-in area is within the field of view of the second camera, the device switches to recording the video captured using the second camera.

[0007] To ensure that the recorded video provides a smooth transition between the outputs of the different cameras, the device can use a series of transformations to relate the output of the second camera to the output of the first camera. These transformations may be implemented using a homography matrix or in other forms. In some implementations, the transformations involve mapping the image from the second camera to a canonical camera space by removing camera-specific time-dependent contributions from the second camera, such as rolling shutter effects and OIS lens movement. The second transformation can project the image data in the canonical image space of the second camera into the canonical image space of the first camera. This can align the field of view of the second camera with the field of view of the first camera and account for spatial differences (e.g., offsets) between the cameras in the device. Electronic image stabilization (EIS) processing can then be applied to the image data in the canonical image space of the first camera. This series of transformations provides a much more efficient processing technique than, for example, attempting to relate and align EIS-processed second camera image data with EIS-processed first camera image data. The EIS processing may be performed in a single camera space or frame of reference, even though the images are captured using different cameras with different fields of view, different inherent characteristics during image capture, etc. The output of the EIS processing in the canonical image space of the first camera may then be provided for storage as a video file (locally and / or remotely) and / or streamed for display (locally and / or remotely).

[0008] These techniques can apply a level of image stabilization tailored to the current zoom level to more effectively control camera shake and other unintentional camera movements. Additionally, the ability to smoothly transition between cameras during video capture can increase the resolution of the video capture without disruptive transitions. For example, as digital zoom is increasingly applied to the output of a camera with a wider field of view, resolution tends to decrease. As digital zoom increases, the resulting output image represents a smaller portion of the image sensor and therefore uses fewer pixels of the image sensor to generate the output image. The second camera can have a lens with a narrower field of view, allowing the narrower field of view to be captured across the entire image sensor. Once the video is zoomed in to the point where the output frame falls within the field of view of the second camera, the camera can transition video capture to using image data captured by the second camera. Thus, with EIS continuously and consistently applied while continuing to capture and record video files, the camera can switch between using different cameras for video capture in a virtually seamless manner. The switching can be transparent to the user, so that switching between cameras is not noticeable in the captured video footage or, optionally, in the user interface for the user.

[0009] Generally, the process of performing video capture and associated image processing can be computationally expensive, especially for high-resolution video capture. The techniques discussed herein provide, among other techniques, computationally efficient techniques for managing image stabilization and transitions between camera modules by mapping image data from both cameras into a single, common reference space and then applying the same type of EIS processing to the image data in that reference space. The reference space may be one in which time-dependent effects are removed, which further reduces the computation required for aligning images from different cameras and applying EIS processing.

[0010] In some implementations, the techniques described herein are implemented on battery-powered devices with limited power budgets and limited computational resources, such as phones, tablet computers, and other mobile devices. The processes discussed can be performed efficiently on battery power by the device, performing stabilization processing substantially in real time, concurrently with ongoing video capture, e.g., with image-stabilized video output being saved or streamed as video capture continues. Also, as described below, the techniques can coordinate the capture of video using camera modules, e.g., to adjust camera module settings for focus, exposure, etc., and to switch which camera module is used at different times. These techniques also execute efficiently, such that they can be performed in real time as video is captured, processed, and recorded while additional video continues to be captured.

[0011] In one general aspect, a method includes a video capture device having a first camera and a second camera providing a digital zoom function that enables user-specified magnification changes within a digital zoom range during video recording, the video capture device being configured to (i) use video data captured by the first camera over a first portion of the digital zoom range and (ii) use video data captured by the second camera over a second portion of the digital zoom range, and the method further includes, while capturing video using the second camera of the video capture device to provide a zoom level in the second portion of the digital zoom range, processing image data captured using the second camera by applying a set of transforms including: (i) a first transform to a second canonical reference space for the second camera; (ii) a second transform to the first canonical reference space for the first camera; and (iii) a third transform to apply electronic image stabilization to the image data in the first canonical reference space for the first camera.

[0012] In some implementations, the method includes, while capturing video using a first camera of a video capture device, processing image data captured using a second camera by applying a set of transformations including (i) a transformation to a first canonical reference space for the first camera and (ii) a transformation for applying electronic image stabilization to the data in the first canonical reference space for the first camera to provide a zoom level in a first portion of a digital zoom range.

[0013] In some implementations, the first camera and the second camera have different fields of view, and (i) the field of view of the second camera is contained within the field of view of the first camera, or (ii) the field of view of the first camera is contained within the field of view of the second camera.

[0014] In some implementations, the first camera and the second camera each include a fixed focal length lens assembly.

[0015] In some implementations, the canonical reference space for the second camera and the canonical reference space for the first camera are conceptual camera spaces defined by a predetermined, fixed set of camera-specific properties, such that projecting the image data into the canonical reference space removes time-dependent effects during capture of the video frames.

[0016] In some implementations, the first camera includes an optical image stabilization (OIS) system, and the first canonical reference space for the first camera is one in which the image data is represented with a consistent, predetermined OIS position.

[0017] In some implementations, the second camera includes an optical image stabilization (OIS) system, and the second canonical reference space for the second camera is one in which the image data is represented with a consistent predetermined OIS position.

[0018] In some implementations, a first camera provides image data that progressively captures image scan lines of an image frame, and a first canonical reference space for the first camera is one in which the image data has been corrected to remove distortion due to the progressive capture of the image scan lines.

[0019] In some implementations, the second camera provides image data that progressively captures image scan lines of the image frame, and the second canonical reference space for the second camera is one in which the image data has been corrected to remove distortion due to the progressive capture of the image scan lines.

[0020] In some implementations, the second transformation aligns the field of view of the second camera with the field of view of the first camera and adjusts for a spatial offset between the first and second cameras.

[0021] In some implementations, the first transform, the second transform, and the third transform each have a corresponding homography matrix, and processing the image data includes applying the homography matrix.

[0022] In some implementations, the method includes, during capture of video data using a first camera and processing of the video data from the first camera to apply electronic image stabilization, receiving user input indicating a change in zoom level to a particular zoom level in a second portion of a digital zoom range. In response to receiving the user input, the method includes recording a sequence of video frames in which a magnification of the video frames captured using the first camera is incrementally increased until a predetermined zoom level is reached, initiating video capture using a second camera, and recording a second sequence of video frames captured using the second camera, the second sequence of video frames providing the predetermined zoom level and providing increasing magnification of the video frames captured using the second camera until the particular zoom level is reached.

[0023] In some implementations, the second transformation is determined based at least in part on the focal length of the second camera.

[0024] In some implementations, the first transformation includes multiple different adjustments to different scanlines of image data captured using the second camera.

[0025] In some implementations, the third transformation includes electronic image stabilization for each particular one of the video frames using one or more video frames before the particular video frame and one or more video frames after the particular video frame.

[0026] In some implementations, the second camera has a smaller field of view than the first camera. The method includes receiving, during image capture using the first camera, user input indicating a change in zoom level for video capture, and determining, in response to receiving the user input, whether the changed zoom level is greater than or equal to a predetermined transition zoom level, the predetermined transition zoom level representing a smaller field of view than the field of view of the second camera.

[0027] In some implementations, the method includes storing data indicating (i) a first transition zoom level for transitioning from video capture using a first camera to video capture using a second camera, and (ii) a second transition zoom level for transitioning from video capture using the second camera to video capture using the first camera, where the first transition zoom level is different from the second transition zoom level, and the method further includes determining whether to switch between cameras for video capture by (i) comparing the requested zoom level to the first transition zoom level when the requested zoom level corresponds to a decrease in field of view, and (ii) comparing the requested zoom level to the second transition zoom level when the requested zoom level corresponds to an increase in field of view.

[0028] In some implementations, the first transition zoom level corresponds to a smaller field of view than the second transition zoom level.

[0029] In some implementations, the method includes determining, during recording of a video file, to switch from capturing video using a particular one of the cameras to capturing video using another one of the cameras, and in response to determining to switch, determining values ​​of video capture parameters used for image capture using the particular camera, setting values ​​of the video capture parameters for the other camera based on the determined video capture parameters, and after setting the values ​​of the video capture parameters for the other camera, initiating video capture from a second camera and recording the captured video from the second camera to a video file. Setting the video capture parameters includes adjusting one or more of exposure, image sensor sensitivity, gain, image capture time, aperture size, lens focal length, OIS status, or OIS level for the second camera.

[0030] Other embodiments of this aspect, as well as other embodiments discussed herein, include corresponding systems, apparatus, and computer programs encoded on computer storage devices and configured to perform the operations of the methods. One or more computer or other device systems may be so configured by software, firmware, hardware, or a combination thereof installed on the system that, during operation, causes the system to perform actions. One or more computer programs may be so configured by having instructions that, when executed by a data processing device, cause the data processing device to perform actions.

[0031] In another general aspect, one or more machine-readable media store instructions that, when executed by one or more processors, cause performance of operations, the operations including: a video capture device having a first camera and a second camera providing a digital zoom function that enables user-specified magnification changes within a digital zoom range during video recording, the video capture device being configured to (i) use video data captured by the first camera over a first portion of the digital zoom range and (ii) use video data captured by the second camera over a second portion of the digital zoom range; and, while capturing video using the second camera of the video capture device, to provide a zoom level in the second portion of the digital zoom range, the operations further include processing image data captured using the second camera by applying a set of transforms including: (i) a first transform to a second canonical reference space for the second camera; (ii) a second transform to the first canonical reference space for the first camera; and (iii) a third transform for applying electronic image stabilization to image data in the first canonical reference space for the first camera.

[0032] In another general aspect, a video capture device includes a first camera having a first field of view, a second camera having a second field of view, one or more position or orientation sensors, one or more processors, and one or more data storage devices that store instructions that, when executed by the one or more processors, cause performance of operations, the operations including the video capture device having the first camera and the second camera providing a digital zoom function that allows user-specified magnification changes within a digital zoom range during video recording, the video capture device (i) using video data captured by the first camera over a first portion of the digital zoom range, and (ii) using the digital zoom function to perform a digital zoom function. and (iii) a third transformation for applying electronic image stabilization to the image data in the first canonical reference space for the first camera, while capturing the video using the second camera of the video capture device to provide a zoom level in the second portion of the digital zoom range.

[0033] In some implementations, the first camera and the second camera have different fields of view, and the field of view of the second camera is contained within the field of view of the first camera. The first camera and the second camera can each include a fixed focal length lens assembly.

[0034] The details of one or more embodiments of the invention are set forth in the accompanying drawings and the description below. Other features and advantages of the invention will become apparent from the description, drawings, and claims. [Brief explanation of the drawings]

[0035] [Figure 1A] FIG. 1 illustrates an example of a device that provides multi-camera video stabilization. [Figure 1B] FIG. 1 illustrates an example of a device that provides multi-camera video stabilization. [Figure 2] FIG. 2 is a block diagram illustrating example components of the device of FIGS. 1A-1B. [Figure 3] FIG. 1 illustrates an exemplary technique for video stabilization. [Figure 4] 1A-1B are block diagrams illustrating additional examples of processing by the devices of FIGS. 1A-1B. [Figure 5A] FIG. 1 illustrates an example of a technique for multi-camera video stabilization. [Figure 5B] FIG. 1 illustrates an example of a technique for multi-camera video stabilization. [Figure 5C] FIG. 1 illustrates an example of a technique for multi-camera video stabilization. [Figure 6] FIG. 1 illustrates an example transformation that can be used to efficiently provide multi-camera video stabilization. DETAILED DESCRIPTION OF THE INVENTION

[0036] Like reference numbers and designations in the various drawings indicate like elements. Detailed Description 1A-1B illustrate an example device 102 that provides multi-camera video stabilization. Video stabilization is often an important feature for camera systems on mobile devices. In traditional single-camera video stabilization, video frames can be converted from image data actually captured during a shaky real camera trajectory to stabilized output for a time-smoothed virtual camera trajectory. In a single camera, frames acquired during the real camera trajectory can be projected onto modified output frames that represent frames along the smoothed virtual camera trajectory.

[0037] Some devices include multiple cameras facing the same direction (e.g., on the same side of the device) positioned to capture different views of the same scene. In a multi-camera system, the device may switch cameras over the course of capturing video so that a single video includes segments captured using different cameras. When a switch between cameras occurs, a single real camera trajectory no longer exists because different real cameras are being used. Nevertheless, the device should maintain the same virtual camera trajectory to maintain continuity and smoothness of the captured footage, even during transitions between cameras. As discussed herein, the virtual camera trajectory can be maintained across multiple cameras and smoothed over time so that switching between cameras does not cause distracting or disruptive effects (e.g., stutter, abrupt changes in image offset or field of view, interruptions in stabilization or abrupt changes in the applied stabilization level, etc.).

[0038] Multi-camera video stabilization systems can be implemented to provide a variety of benefits. In addition to smoothing video over time, transitions between periods of video capture from different cameras can also be smoothed. Additionally, the technology can be efficient enough to operate in real time, for example, to apply smoothing to captured video simultaneously with continuous video capture, and to conserve power to enable long-term use by battery-operated devices.

[0039] 1A, device 102 is shown as a phone, but may be another type of device, such as a tablet computer, a camera, etc. Device 102 includes a multi-camera module 108 that includes a first camera 110a and a second camera 110b. The two cameras 110a, 110b are located on the same side of device 102 and are therefore both positioned to capture images of the same scene 105 facing device 102 (e.g., on the back of the phone as shown). Cameras 110a, 110b may be securely coupled together to device 102 such that cameras 110a, 110b move together to the same extent and with the same movement as device 102.

[0040] The cameras have different fields of view. For example, a first camera 110a has a first field of view 120, and a second camera 110b has a second field of view 122 that is narrower than the first field of view 120. Although the fields of view 120, 122 may be significantly different, the two cameras 110a, 110b may have similar image resolution. The two fields of view 120, 122 may overlap. In particular, the field of view 122 of the second camera 110b may be completely contained within the field of view 120 of the first camera 110a. As an example, the first field of view 120 may be 77 degrees, and the second field of view 122 may be 52 degrees, with the 52-degree field of view 122 being mostly or completely within the 77-degree field of view 120. In this example, the cameras 110a, 110b each use a fixed focal length lens, e.g., a lens assembly without optical zoom. In other words, the lens focal length of each camera 110a, 110b may be fixed, independent of focus-related effects such as focus breathing.

[0041] One or more of the cameras 110a, 110b may include an optical image stabilization (OIS) module to reduce the effects of camera shake and other undesirable movements of the device 102. Regardless of whether the cameras 110a, 110b include an OIS module, the device 102 can use electronic image stabilization (EIS) to smooth captured video over time. As described further below, EIS processing can apply transformations to captured frames to project captured frames taken along a shaky real camera trajectory into stabilized output frames that represent a smoothed or filtered trajectory of a virtual camera.

[0042] The device 102 uses cameras 110a and 110b to provide an effective zoom range using image capture at different magnification levels from the different cameras 110a and 110b. The device 102 can provide the zoom range using digital zoom techniques applied to the output of the different cameras 110a and 110b at different portions of the zoom range. As an example, the device may provide an overall zoom range of 1.0x to 3.0x. The camera 110a, which provides a wider field of view 120, may be used to capture images at a first portion of the zoom range, e.g., 1.0x to 1.8x. When the zoom reaches a predetermined transition point, e.g., a particular level such as 1.8x, the device 102 switches from capturing video using the first camera 110a to capturing video using the second camera 110b. The camera 110b, which provides a narrower field of view 122, may be used to capture images at a second portion of the zoom range, e.g., 1.8x to 3.0x.

[0043] The illustrated example shows that device 102 includes a display that allows a user to view captured video 130 as it is being captured and recorded by device 102. Device 102 can provide a user with a control 132 for dynamically adjusting the zoom level during video capture. In this example, control 132 is an on-screen slider control displayed on the touchscreen of device 102, allowing the user to set the zoom to a desired position along the zoom range. In some implementations, the zoom level is adjusted in fine-grained increments, e.g., 0.2x, 0.1x, or smaller steps, allowing a user to gradually move across the supported zoom range in a manner that substantially approximates continuous zoom across the supported zoom range.

[0044] FIG. 1B shows an example of how different zoom levels can be provided using the outputs from cameras 110a and 110b. Respective fields of view 120 and 122 are shown with cropped portions 140a through 140e that can be used to provide different zoom levels. Output frames 150a through 150e show example frame outputs that can be provided at different zoom levels. Note that because the two cameras 110a and 110b are physically offset from one another within device 102, there are differences in their views of the scene. Even if the two cameras are viewing the same scene, objects within the scene will have slightly different positions at the cameras' outputs due to monocular parallax. Image transformations, described further below, can correct for this parallax and other differences between the outputs of cameras 110a and 110b.

[0045] Each of the output frames 150a-150c is derived from image data from a first camera 110a, which provides a wider field of view 120. When the zoom level reaches a threshold, such as 1.8x, the device 102 switches to using image data captured by a second camera 110b, which provides a narrower field of view 122. The transition between cameras 110a, 110b can occur at a zoom level that is entirely within the field of view 122 of the second camera 110b, such that there is enough image data to fill the output frame at the desired zoom level. In some implementations, the device 102 is configured to perform the transition when the zoom level corresponds to an area smaller than the full output of the second camera 110b, in order to preserve a margin of captured data that can be used for EIS processing.

[0046] Generally, to maximize output quality, it is advantageous to switch to a narrower field of view 122 near a zoom level where the field of view 122 can fill the output frame. This is because the narrower field of view 122 provides higher resolution for that region of the scene. The two cameras 110a, 110b may have similar resolutions, but at the narrower field of view 122, camera 110b can use its full resolution to capture a view of the scene, whereas the same view would be captured at a fraction of the resolution of the wider camera 110a. In this example, there is a zoom level, e.g., 1.7x, where the zoomed-in field of view 140c of the first camera 110a matches the full field of view 122 of the second camera 110b. At this point, only a significantly smaller portion of the image sensor for the first camera 110a is used to provide the output frame 150c, and therefore, the output frame 150c may be of lower resolution or lower quality than would be provided at a wider zoom level. In contrast, the full resolution of the image sensor of the second camera 110b can be used to provide that level of effective zoom or magnification, resulting in a higher quality output. By switching to the second camera 110b at or shortly after the balance point (e.g., 1.7x), the device 102 can provide higher quality at that zoom level and further zoom levels.

[0047] In some implementations, the zoom level set for switching between cameras 110a-110b is set after the point at which the second camera 110b can fill the image frame. For example, the transition point may be set at a zoom level that provides a margin to account for monocular parallax resulting from the different physical locations of cameras 110a-110b. As a result, if the second camera 110b can fill the frame at a 1.7x zoom level, device 102 may nevertheless delay switching to capture using the second camera 110b until, for example, a 1.8x or 1.9x zoom so that the full sensor output from camera 110b provides a buffer area of ​​image data (e.g., at the edges surrounding the area representing the desired frame capture area), so that any offset or other adjustments needed to correct for monocular parallax can be made, for example, to align field of view 122 with the enlarged area of ​​field of view 120 and still fill the output frame.

[0048] The device 102 applies video stabilization, e.g., smoothing of camera motion visible in frames over time, to reduce or eliminate the effects of device 102 motion. The device 102 can perform video stabilization, e.g., EIS processing, in real time or near real time, e.g., simultaneously with ongoing video capture for the video being stabilized. Video stabilization smoothing over time (e.g., across multiple frames) can be coordinated with transitions between the cameras 110a and 110b as zoom settings change. Nevertheless, as discussed below, an image data transformation can be used that maps the output of the second camera 110b to the canonical space for the first camera 110a, allowing a single EIS processing scheme to be consistently used for the output of both cameras 110a, 110b.

[0049] The techniques discussed herein can be effectively used with two or more fixed focal length lenses. Optionally, the techniques discussed herein can also be used with one or more lenses that provide optical zoom. For example, the techniques can be used to provide seamless effective zoom across a range that includes image capture from (i) multiple cameras with different optical zoom ranges that may or may not intersect or overlap, or (ii) one or more cameras that provide optical zoom and one or more fixed focal length cameras.

[0050] 2 illustrates an example of a device 102 that provides video stabilization. As described above, the device 102 includes a first camera 110a and a second camera 110b. One or more of the cameras 110a-110b may optionally include an OIS module 215a-215b. If the device 102 includes an OIS module 215a-215b, the device 102 may capture video frames while using the OIS module 215a-215b to at least partially counteract movement of the device 102 during frame capture. The device 102 also includes one or more device position sensors 220, one or more data storage devices 230, and an EIS module 255.

[0051] The device 102 can be any of a variety of types, including a mobile phone, a tablet computer, a camera module such as a camera, etc. In some implementations, the device 102 can include a computing system for performing the operations of the EIS module 255, which may be implemented in software, hardware, or some combination thereof. For example, the device 102 may include various processing components, such as one or more processors, one or more data storage devices that store executable instructions, memory, input / output components, etc. The processor that performs the EIS processing may include a general-purpose processor (e.g., the main CPU of a mobile phone or other device), a graphics processor, a co-processor, an image processor, a fixed-function EIS processor, or any combination thereof.

[0052] The EIS module 255 uses position data from both the device position sensor 220 and the OIS modules 215a-215b to stabilize video captured by the recording device. For example, the position data from the OIS modules 215a-215b can be used to determine an offset representing the effect of OIS movement on the expected camera view that would be inferred from the device position data. This allows the EIS module 215 to estimate an effective camera position that reflects the image sensor's actual view, even if the OIS modules 215a-215b change the camera's view of the scene relative to the device position. These techniques, along with other features described herein, allow the device 102 to effectively use OIS and EIS processing simultaneously and realize the benefits of both technologies.

[0053] Generally, OIS can be very effective at reducing blur within individual frames due to camera shake, and OIS can be somewhat effective at reducing motion visible across a series of frames. However, OIS used alone often suffers from various limitations. OIS modules may be limited in the speed with which they respond to motion and the magnitude of motion they can compensate for. Furthermore, the operation of OIS modules often causes distortions, such as jittery video, and can erroneously counter desired motion, such as panning. The EIS module 255 can mitigate the effects of these limitations using position data describing the internal motion of the OIS module.

[0054] Because the OIS modules 215a-215b attempt to compensate for recording device motion, device motion alone may not represent the true camera view used during video capture. If the EIS processing attempts to compensate for motion based solely on device motion, it may attempt to correct motion already compensated for by the OIS system. Furthermore, OIS generally only partially removes the effects of device motion, and the amount of compensation may vary from frame to frame. To provide high-quality stabilization, the EIS module 255 uses OIS position data along with device-level position data to vary the amount of stabilization applied to each frame, and in some implementations, to individual scan lines of a frame. This processing can provide effective stabilization while reducing or eliminating distortion in the video footage. For example, changes in the OIS lens shift position during frame capture can introduce distortion, especially when combined with a rolling shutter, which is typical of many camera modules. Using information about the OIS lens shift at different times during frame capture, the EIS module 255 can estimate the lens position when different portions of the frame were captured and correct the image. The EIS module 255 can also compensate to reduce the effects of OIS lens shift that interferes with panning or is otherwise undesirable.

[0055] Another way the EIS module 255 can enhance video is by analyzing data about later-captured frames. To process a particular frame, the EIS processing module may evaluate a set of camera positions within a time window that includes the time when one or more future frames were captured. Information about the future frames and their corresponding positions can be used in several ways. First, the EIS module 255 can apply filtering to the set of camera positions to smooth out motion patterns used to define image transformations to modify the frames. Second, the EIS module 255 can use the set of camera positions to evaluate the likelihood that consistent motion (e.g., panning) is present or being attempted, and then, if so, adjust the frame consistently with this motion. Third, the EIS module 255 can evaluate the frame's camera position relative to future camera positions and make adjustments for large future motion. For example, if large, rapid motion is identified for a future frame, the EIS module 255 can begin adjusting the frame's content before the motion begins. Rather than allowing large visible motion over several frames, the EIS module 255 can spread the motion over larger frames, with incremental image shifts occurring during earlier frames, gradually spreading the motion over a larger number of frames.

[0056] The EIS module 255 performs regional compositing of the output frame, for example, modifying the transformations applied to each scan line of the image frame, which allows the system to compensate for rolling shutter distortion, movement of the OIS modules 215a-215b, and various device movements that occur within the capture duration of a single frame.

[0057] 2 , device 102 may be any suitable device having a camera to capture video data, such as a camera, a cellular phone, a smartphone, a tablet computer, a wearable computer, or other device. While the example of FIG. 2 shows a single device capturing and processing video, functionality may optionally be spread among multiple devices or systems. For example, a first device may capture video frames and record location data and other parameters as metadata. The first device may provide the video frames and metadata to a second device, such as a local computing system or a remote server, that can perform the EIS processing described herein.

[0058] The first camera 110a may include lens elements, an image sensor, sensor reading circuitry, and other components. The OIS modules 215a-215b may include a sensor, a movable element, a processor, and a drive mechanism for moving the movable element. The movable element is located in the optical path of the first camera 110a. For example, the movable element may be a reflective or refractive element, such as a lens, mirror, or prism. In some implementations, the movable element is the image sensor of the first camera 110a. The sensor may include one or more gyroscopes or other sensors for detecting movement. The processor determines the amount and direction of movement required of the movable element to compensate for the movement indicated by the sensor and then commands the drive mechanism to move the movable element.

[0059] The device 102 may include one or more sensors that measure changes in the orientation of the device 102. device In some implementations, the device 102 includes a position sensor 220. device Position sensor 220 is separate from the sensors used by OIS modules 215a-215b. device The position sensor 220 can detect rotation of the device 102 around one or more axes. , the device position sensor 220 may be a three-axis gyroscope or an inertial measurement unit (IMU). Additionally or alternatively, other sensors may be used to determine the device position. For example, one or more accelerometers, one-axis gyroscopes, two-axis gyroscopes, etc. may be used to determine the position of the device 102. In general, any suitable sensor or combination of sensors may be used that allows the rotational position of the device 102 to be determined.

[0060] In some examples, the position data from the gyroscope sensors of the OIS modules 215a-215b may be Device 102 Separate device In addition to, or instead of, using the position sensor 220, data may be captured and stored. Nevertheless, it may be beneficial for the device 102 to use a gyroscope sensor with different characteristics than the OIS sensor. For example, the gyroscope sensor of the device 102 may provide measurements at a rate of approximately 400 Hz with a sensitive rotation range of greater than 100 degrees / second. Compared to a device-level sensor, a typical gyroscope sensor of an OIS module may provide measurements at a different rate and range, e.g., a rate of 5000 measurements / second or more with a sensitive rotation range of approximately 10 degrees / second. In some implementations, having a larger sensitive rotation range of the device-level sensor (e.g., to describe large movements) is beneficial, as is the more frequent measurement of the OIS module sensor (e.g., to detect small changes or high-frequency patterns). Thus, both types of data may be used together to determine the location of the device 102.

[0061] The device 102 includes one or more data storage devices 230 that store information characterizing the first camera 110a and the frame capture process. For example, the stored data can include calibration data 232 that indicates the relationship between the position of the OIS modules 215a-215b and the resulting offsets in the image data. Similarly, the calibration data 232 can indicate the correspondence between camera module lens focus positions and effective focal lengths for those focus positions (e.g., with different mappings for each camera 110a-110b), allowing the system to account for focus breathing. Additionally, the calibration data 232 or other stored data can indicate the correspondence between camera lens focus positions and subject distances, allowing for conversion from the lens focus position selected by the autofocus system to subject distance, which indicates the distance of the focused subject from the camera's sensor plane. The calibration data 232 also indicates the 3D spatial position of one camera 110a relative to the other camera 110b. Typically, this includes calibration data specifying the 3D rotation and 3D translation of one camera relative to the other. Typically, calibration is performed for each device manufactured to ensure the best user experience. As a result, the calibration data 232 can be highly accurate to the characteristics of a particular camera module (e.g., cameras 110a, 110b, their mounting structures, etc.) and the state of the module after manufacture. The stored data can include scan pattern data 234, which can indicate the reading characteristics of the image sensor in the first camera 110a. For example, the scan pattern data 234 can indicate the direction of the scan (e.g., scan lines are read from top to bottom), whether scan lines are read individually or in groups, etc.

[0062] During video capture, cameras 110a-110b, OIS modules 215a-215b, and device position sensor 220 may each provide information about the video capture process. First camera 110a provides video frame data 242a, e.g., a sequence of video image frames. Second camera 110b similarly provides video frame data 242b. Cameras 110a-110b also provide frame exposure data 244a-244b, which may include, for each captured frame, an indication of the exposure duration and a reference time indicating when the exposure occurred (e.g., the start or end time of the exposure). Cameras 110a-110b also provide lens focus position data 246a-246b indicating the lens focus position for each captured frame.

[0063] The OIS modules 215a-215b provide OIS position data 248a-248b indicating the position of the moving elements of the OIS modules 215a-215b at various times during video capture. For example, if the moving elements are movable lenses that shift to compensate for motion, the OIS modules 215a-215b can provide lens shift readings that identify the current position of the movable lens in each. The device 102 can record the lens shift position and the time that position occurred. In some implementations, the OIS position data 248a-248b is captured at a high frequency, e.g., at a rate higher than the frame rate of the video capture, so that multiple measurements are taken over the duration of each video frame exposure.

[0064] The device position sensor 220 provides device position data 250 indicative of the rotation and / or other movement of the device 102 during video capture. The device position can be measured at high frequencies, e.g., 200 Hz or higher. Thus, in many instances, measurements can be taken for multiple different times during the capture of each video frame.

[0065] The lens focus position data 246, OIS position data 248, and device position data 250 may all be recorded with a timestamp indicating the time the identified position occurred. The timestamp may be accurate, for example, to the nearest millisecond, allowing the data obtained from the various position measurements to be aligned in time. Additionally, the position of the device, OIS system, or lens focus mechanism may be interpolated to determine values ​​at times between measurements.

[0066] An example of potential timing for data capture is shown in chart 252. As shown, device position data 250 (e.g., gyroscope data) and OIS position data 248 (e.g., lens shift position data) may be captured at a rate higher than the video capture frame rate (e.g., 30 frames / second, 60 frames / second, etc.), and multiple positions of the device and OIS system can be determined for each video frame. As a result, for each scan line, which may be horizontal or vertical depending on the shutter orientation, different device positions and OIS settings can be used to determine the transformation of that scan line. Lens focus position data 246 may be captured at least once per image frame. This position data may be captured asynchronously with respect to the frame exposure, for example, by the gyroscope sensor data and OIS position data being sampled at a rate that exceeds, but is not necessarily synchronized with, the start or end of the image frame exposure.

[0067] Data obtained from cameras 110a-110b and other components is provided to EIS module 255 for processing. This processing may occur while video capture is in progress. For example, EIS processing may be performed substantially in real time, whereby a video file made accessible to a user at the end of video capture has been stabilized by EIS module 255. In some implementations, EIS processing may be performed at a later time, for example, after video capture is complete or by a different device than the one that recorded the video. EIS module 255 may be implemented in hardware, firmware, software, or a combination or subcombination thereof.

[0068] FIG. 2 illustrates only some of the functionality of the EIS module 255. The example in FIG. 2 illustrates EIS processing for video frames, including frames captured from a single one of cameras 110a-110b, and does not describe features used to provide digital zoom or to adjust EIS processing for zoom level. However, as described below, the device 102 can provide digital zoom during video capture and recording using image data captured from both cameras. For example, when a user zooms in to a threshold zoom amount, the device 102 can switch from using images captured from the first camera 110a to using images captured from the second camera 110b. Furthermore, because zooming in can accentuate visible shake in the video, the zoom level can adjust the transformations and other operations of the EIS module 225, for example, to apply greater stabilization levels as the zoom level increases. Adjusting EIS processing to account for digital zoom and transitions between multiple cameras is discussed with respect to FIGS. 4-6. In some implementations, the transformations can be achieved by applying a different transformation for each line of the captured image. For each line of the image, the device 102 can calculate a specific timestamp for that line, which is associated with the corresponding OIS and device position data.

[0069] 2 , the EIS module 255 includes a device position data handler 256 that periodically or continuously obtains updated device position data 250 from the device position sensor 220. The motion data handler estimates the current camera pose from the device position data 250. For example, a gyroscope signal may be obtained and used to estimate the device position of the device 102 at a high frequency, such as 200 Hz. This device position at a given time t is hereinafter referred to as R(t). This device position may indicate the rotational position of the device 102 about one, two, or three axes, for example. The device position may be expressed as a rotation matrix, with respect to a coordinate system, or in other forms. Each calculated device position may be labeled with a time indicating the time at which that position of the device 102 occurred.

[0070] The EIS module 255 includes an OIS position data handler 258, which periodically or continuously obtains OIS position readings, denoted as OIS position data 248. The OIS position data handler 258 converts the OIS readings into offsets that can be used with device position. For example, an OIS lens position can be converted into a two-dimensional pixel offset. To generate the offsets, the OIS position data handler 258 can use stored calibration data 232, which can provide conversion coefficients or matrices to convert from OIS position to a corresponding offset. Generating the offsets from the OIS position can take into account changes in the camera's effective focal length over time due to changes in lens focus position and / or lens zoom position, for example, if the first camera 110a is capable of optical zoom. Device Location Like the data handler 256, the OIS position data handler 258 labels each measurement and offset with the time that the data represents.

[0071] The EIS module includes a motion model builder 260 that receives the device position calculated by the device position data handler 256 and the OIS offset calculated by the OIS position data handler 258. Using this data, along with the frame exposure data 244 and lens focus position data 246, the motion model builder 260 generates a first transform 262 for the frame. For example, the first transform 262 can be a projection matrix that maps the real-world scene in the camera's view to the captured frame. This process is repeated for each frame. When generating the first transform 262 for a frame, the position of the OIS modules 215a-215b can be modeled as an offset from the primary device position determined from the gyroscope data. As discussed further below, the offset can take into account the camera's effective focal length at the time of capture by examining the effective focal length relative to the lens focus position at that time. The first transform 262 can separately describe the relationship between different subsets or regions of a single image frame. For example, different portions or components of the first transform 262 may describe how different scanlines of a frame map to a real-world scene. The device position, OIS module position, subject distance (e.g., the distance of the focused subject from the camera), and lens focus position can all be aligned using measurement timestamps and interpolated as necessary to provide accurate positions at the time of exposure of individual scanlines of a frame. For lenses with autofocus, the focus position is set depending on how far the subject is from the camera. Thus, a map exists that shows the relationship between lens focus position and subject distance. The mapping can be generated and calibrated, and the subject distance is used in later calculations for spatial transitions.

[0072] The first transform 262 generated by the motion model builder 260 is provided to a nonlinear motion filtering engine 270, which determines a second transform 272. This second transform 272 may be a second projection matrix P'i,j that projects the image data of the frame onto an output frame representing a stabilized version of that frame. Specifically, rather than operating on the captured image data, the second transform 272 may map the image projection Pi,j created using the first transform 262 onto the output frame. In some implementations, the two transforms 262, 272 may then be combined into a single transform that operates on the initially captured image data of the frame and maps it directly to the stabilized output frame.

[0073] To effectively stabilize the motion, the nonlinear motion filtering engine 270 can generate a second transform 272 to take into account motion that will occur in the future after the capture of the frame being processed. For example, for the current frame under analysis, the position of the recording device may not have moved significantly from the previous frame. Nevertheless, Nonlinear Motion Filtering If the engine 270 determines that significant motion will occur in a future frame, a second transform 272 may be generated to shift or otherwise modify the current frame to introduce visible motion into the video, so that large future motion can be spread as a series of gradual changes rather than abrupt changes. Similarly, if stabilization of a future frame results in cropping or other changes, a second transform 272 may be generated to at least partially propagate those changes to earlier frames for more gradual and consistent changes over a series of frames.

[0074] non-linear MovementThe filtering engine 270 can generate a second transform 272 from the virtual camera position relative to the frame. Rather than representing the actual position of the camera when the exposure occurred, the virtual camera position can represent an adjusted or virtual pose of the device 102 that will stabilize the video being recorded. The virtual position can represent a desired position for placing the virtual camera, for example, a position that would simulate a particular view or perspective of a scene. Generally, any camera position can be represented by its rotation and translation relative to a global reference frame. The virtual camera position can be represented as a rotation matrix, for example, a matrix indicating rotational offsets relative to a reference position. This can be a 3x3 matrix indicating rotational offsets relative to three axes of rotation. In some implementations, the stabilization process of the EIS module defines the position only in terms of rotational components, as these generally have the greatest impact on handheld video stability.

[0075] The virtual camera position for a frame can reflect adjustments to the estimated camera position to enhance video stabilization, correct for distortion and motion, facilitate panning, and otherwise enhance the video. The virtual camera position for a frame can be determined by generating an initial camera position that is adjusted based on various factors. For example, adjustments to the virtual camera position can be made through filtering of the device position based on motion detected before and after the frame, based on the amount of blur in the frame, based on the likelihood that panning is occurring, through adjustments to prepare for motion in future frames, and / or to ensure that image data covers the entire output frame. The various factors can be considered by generating a series of virtual camera positions for the frame that are altered, blended, or otherwise used to determine the final virtual camera position for that frame.

[0076] Just as transforms 262, 272 can have different mappings for different scanlines, different virtual camera positions can be determined for different scanlines of a frame to adjust for changes in device position, OIS module 215a-215b position, and / or lens focus position during frame capture. Thus, different virtual camera positions can be used for different portions of a frame. For efficiency, the virtual camera positions and corresponding components of second transform 272 can be calculated for an appropriate subset of the image sensor's scanlines, and then appropriate data can be interpolated for the remaining scanlines. For simplicity, the various examples below discuss a single scanline, such as the center scanline of the image sensor. The techniques for fully calculating virtual camera positions and corresponding projection matrix components can be used for multiple scanlines of an image frame, and even for each scanline individually, if desired.

[0077] As used herein, device position refers to the position of the device 102, for example, as indicated by the device position data 250 (e.g., gyroscope data) and the output of the device position data handler 256. This device-level position refers to the attitude or orientation of the device 102 without considering internal movement of the lens of the first camera 110a or movement of the OIS modules 215a-215b. Also, as used herein, camera position refers to a position corresponding to the effective or estimated view of the camera. The camera position may differ from the device position by accounting for movement of the OIS modules 215a-215b, lens breathing, and shifts due to other factors. Additionally, the camera position may be a virtual position, e.g., an approximate or virtual position that reflects an enhanced or modified view of the camera rather than the actual view of the camera.

[0078] The EIS module 255 then uses the output of the nonlinear motion filtering engine 270 to map each captured image frame to an output frame using the image warping engine 280. The second projection 272 may include components corresponding to each scan line of the frame, such that each portion of the frame is mapped to the output space, defining each pixel of the output frame. The processing of the EIS module 255 may be performed for each frame of the video.

[0079] 3 is a diagram illustrating an example of data used for video stabilization. The diagram shows a series of frames 310 captured by device 102. Each frame is labeled with a corresponding set of metadata 312 indicating, for example, exposure duration, exposure reference time (e.g., exposure start time, stop time, or other reference point), lens focus position, etc. Although not shown, device position data and OIS module position data are captured and time-stamped at various times during each exposure.

[0080] To stabilize frame 311, a time range around the capture of frame 311 is defined. This time range, or window of frames, is analyzed to determine how to transform frame 311. As used herein, the time of a frame, "t," generally refers to the time of capture of the center line, which is used to represent the reference time of capture of the frame. When referring to the time of an individual line (e.g., to account for a time that may differ from the primary time of capture for the center line of the frame), that time is referred to as t Lwhere L is the index or identifier of a particular scan line. The exposure time t of the center scan line of the frame 311 under analysis can be used as the center of the time range used for the analysis. A predetermined time offset Q can be used to set the time range from the range, e.g., [tQ, t+Q]. In some implementations, this time offset Q is approximately 200 ms. As a result, this range will include approximately 7 frames before and 7 frames after frame 311. Larger and smaller time offsets Q may also be used. Because the EIS module 255 uses the context of future frames during processing, processing of a frame is delayed until the appropriate number of subsequent frames have been captured.

[0081] In FIG. 3 , frame 311 is shown as being captured by an image sensor (e.g., of one of cameras 110a-110b). As described above, EIS module 255 defines first transform 262 from data indicating the actual position of device 102 during the capture of frame 311, as well as the positions of camera elements, such as OIS module elements, and lens focus position. The result of applying first transform 262 is projected image 330, which is shown with respect to output frame target 335. In some implementations, first transform 262 is determined using only data corresponding to the capture of frame 311. Transform 262 corresponds to the actual lens position of the camera, and thus projected image 330 estimates a mapping between image data and the actual scene in the camera's view.

[0082] EIS module 255 further adjusts the image data for frame 311 using a second transform 272. This second transform 272 corresponds to a virtual lens position, e.g., a virtual position that would result in more stable video if used to capture frame 311. This second transform 272, when applied to frame 311, produces projected image 340 that fully defines the data for output frame 335.

[0083] The second transform 272 that generates the projected image 340 may be generated from data corresponding to each frame within the time range from [tQ, t+Q]. The positions R(t) of the device 102 over this time period can be filtered to smooth motion across the range, for example, using a Gaussian filter. The set of positions R(t) in this case is the set that includes the positions of the device 102 at each of the center scan line capture times that occur within the range [tQ, t+Q]. Consider an example where this range encompasses the center scan line capture time t0 of the current frame being processed and the center scan line capture times of seven frames before and after the current frame being processed. The set of positions to be filtered is the set {R(t -7 ),R(t -6 ),…,R(t -1 ),R(t0),R(t1),…,R(t6),R(t -7 )}. The result of filtering at time t, i.e., the exposure of the center scan line of frame 311, can be used as the initial virtual camera position V0(t). Even with filtering, there may be undesired motion in the device position, or other factors that result in undesired motion. As a result, the initial virtual camera position V0(t) may be updated through a series of further operations. In some implementations, the filtered position R(t) is a position that does not assume OIS motion and, therefore, may be based on device position data 250 without considering OIS position data 248. In other implementations, OIS motion and offset may be factored into the set of positions that are filtered to generate the initial virtual camera position V0(t).

[0084] For example, a second virtual camera position V1(t) may be calculated by multiplying the position V0(t) by the final camera position VF(t) relative to the previous frame based on the amount of motion that occurs over the future frame. pre ) can be generated by interpolating the final camera position VF(t pre) may be a virtual camera position (e.g., used to generate a recorded output frame) relative to the center scan line of a frame captured immediately before frame 311. The camera position of the previous frame may be the position corresponding to the final virtual camera position, e.g., the transformation used to generate a stabilized output version of the previous frame. Interpolation may align visible changes in motion between frame 311 and the previous frame with visible changes in motion between frame 311 and a future frame.

[0085] A third virtual camera position V2(t) may be generated by interpolating V1(t) with the real device position R(t) based on the amount of camera motion blur present in frame 311. This may reduce the amount of stabilization applied to reduce the viewer's perception of blur. Because motion blur generally cannot be removed, this may reduce video stabilization when appropriate to produce a more natural result.

[0086] A fourth virtual camera position V3(t) may be generated to simulate or represent a position that occurs during consistent movement of device 102 over the time range [tQ, t+Q]. This position may be determined by applying a stable filter, such as a domain transform filter, to the estimated actual device position R(t) over the time range. While the filter is applied to the same set of device positions used to generate V0(t), this step represents a different type of filtering. For example, V0(t) may be generated through filtering that smooths but generally follows changes in the estimated actual device position over time and does not impose a predetermined shape or pattern. In contrast, V3(t) is generated by filtering the device pattern to fit a predetermined consistent movement pattern, such as a substantially linear panning or other movement that may potentially be intended by a user of device 102.

[0087] A fifth virtual camera position V4(t) may be generated as an interpolation of V3(t) and V2(t). The EIS module 255 may evaluate whether the change in device position over time is likely to represent panning of the device 102 and weight or adjust the interpolation accordingly. If panning is determined to be likely, V4(t) will be closer to the estimated panning position V3(t). If panning is determined to be less likely, V4(t) will be closer to position V2(t).

[0088] At the fifth virtual camera position V4(t), the EIS module 255 can evaluate the coverage that the corresponding transformation would provide to the output frame 335. Because it is desirable to fill the entire output frame 335 and not leave any pixels undefined, the EIS module 255 can determine a transformation, such as a projection matrix, that represents a view of the scene from the virtual camera position V4(t) and verify that the projected image will cover the output frame 335. To account for motion in future frames, the transformation can be applied to the portion of the scene that will be captured by the future image frames. The transformation and the corresponding virtual camera position V4(t) can be adjusted so that the current frame and each of the set of future frames, when mapped using the transformation, will all completely define the output frame 335. The resulting transformation can be set as the transform 272 and used to generate the stabilized output frame 335 for frame 311.

[0089] In some implementations, generating the stabilized output frame 335 relative to frame 311 may involve adjusting the time t L3 includes performing the EIS processing techniques described in for the exposed scan lines L. For example, the processing may be performed for the scan lines at certain intervals (e.g., every 100 scan lines, every 500 scan lines, etc.) or at certain reference points (e.g., 1 / 4 and 3 / 4 across the frame, or the top and bottom of the frame). When the virtual camera position and second transform 272 are determined for only an appropriate subset of the scan lines of frame 311, the transformations for the scan lines (e.g., corresponding portions of the projection matrix) are interpolated between the calculated positions. In this manner, an appropriate transformation is determined for each scan line, and each scan line may result in a different transformation being applied. In some implementations, the entire process of generating the virtual camera position and second transform 272 may be performed for each scan line of each frame, without relying on interpolation between data from different scan lines.

[0090] Once frame 311 is mapped to output frame 335, the result is saved and EIS module 255 begins processing the next frame. The process continues until each of the frames of the video has been processed.

[0091] The various elements used to generate the virtual camera positions and resulting transformations can be used in combination or separately. For example, depending on the implementation, some of the interpolations and adjustments used to create the virtual camera positions V0(t) through V4(t) may be omitted. For example, a different implementation may use any of the filtered camera positions V0(t) through V3(t) to determine a transformation for projecting the data into an output frame, instead of using V4(t) for that purpose. Thus, using any of the filtered camera positions V0(t), V1(t), and V2(t) to generate a stabilization transform may still improve video stability. Similarly, V3(t) may be effective in stabilizing video in which panning is occurring. Many other variations are within the scope of this disclosure, even if they take into account subsets of the different elements discussed.

[0092] The techniques discussed may be applied in a variety of ways. For example, rather than sequentially applying two transforms 262, 272 to image data, the recording device may generate a single combined transform that reflects the combined effect of both. Thus, generating stabilized image data using transforms 262, 272 may involve generating additional transforms or relationships that are ultimately used to stabilize the image data, rather than directly applying transforms 262, 272. Various techniques for image stabilization are described, and other techniques, such as those discussed in U.S. Patent 10,462,370, issued October 29, 2019, which is incorporated herein by reference, may additionally or alternatively be used.

[0093] Figure 4 is a block diagram illustrating additional examples of processing by device 102 of Figures 1A-1B. In addition to the elements shown in Figure 2, some of which are again represented in Figure 4, device 102 may include hardware and / or software elements for providing additional functionality represented in Figure 4.

[0094] The device 102 may include a zoom input processing module 410 that processes user input to the device 102 indicating a requested change in zoom level. When the device 102 captures video, the device 102 may receive user input to change the zoom level, even if EIS is in use. This may include input moving an on-screen slider control, gestures on a touchscreen, or other input. A 1.0x zoom level may represent the widest field of view available using the camera 110a with EIS in use. This may be a cropped section of the native image sensor resolution to provide margin for EIS processing. The module 410 may determine the desired zoom level based on the user input and, for example, move from a current zoom level (e.g., 1.0x) to a modified or desired zoom level (e.g., 2.0x).

[0095] Data indicating the requested zoom level is provided to the camera selector module 420, which determines whether to switch cameras used for video recording. For example, the camera selector module 420 can receive and use stored camera transition thresholds 430 that indicate the zoom level at which a transition between cameras should occur. For example, the second camera 110b can have a field of view corresponding to a 1.7x zoom level, and the transition threshold can be set at a 1.8x zoom level. The camera selector 420 can determine that a desired zoom level of 2.0 meets (e.g., is equal to or greater than) the 1.8x threshold, and therefore a camera change is appropriate for captures representing zoom levels of 1.8x and above. Multiple thresholds may be defined, such as a first threshold for zooming in (e.g., narrowing the field of view) and a second threshold for zooming out (e.g., widening the field of view). These thresholds can be different. For example, the first threshold can be 1.8 and the second threshold can be 1.7, so the transition point is different depending on whether the user is zooming in or out.

[0096] In some implementations, switching between cameras 110a and 110b is controlled simply by exceeding a threshold. Consider a situation in which a device is recording video using camera 110b and the user begins zooming out to a level that triggers a switch to use the wider-angle camera 110a instead. The threshold for switching from camera 110a to camera 110b when zoomed in may be 1.8x, and the lowest zoom position available to camera 110b may be 1.7x (e.g., the zoom level representing the full maximum field of view of the current camera 110b). Consequently, when zooming out, the transition must occur at the 1.7x zoom level because the second camera 110b cannot provide a wider field of view. To address this situation, when the zoom level is reduced to 1.9x, camera 110a begins streaming image data, even though the output frame is still based on the current image data output by camera 110b. This provides a period during a zoom level transition, for example, from 1.9x to 1.8x, during which both cameras 110a, 110b are capturing and streaming image data of the scene. This provides an adjustment period for initializing capture on camera 110a so that camera 110a can "warm up" and autofocus, autoexposure, and other processes begin to converge. This adjustment period may provide margin for the settings of camera 110a so that its settings and capture can be initialized and aligned to match those currently being used by camera 110a before the switch.

[0097] In this way, device 102 can predict and prepare for the need to switch between cameras based on factors such as the current zoom level, the direction of the zoom change, and user interaction with device 102. For example, upon detecting that a camera switch is likely to be required based on user input commanding a decrease in zoom to 1.9x, device 102 can command settings to begin capture on the next camera 110a to be used before or when the user commands the zoom level targeted for the transition. For example, if device 102 begins video capture and settings adjustment when a 1.9x zoom level is commanded by the user, and camera 110a is ready and operating at the desired settings by the time the user commands a 1.8x zoom level, device 102 can switch cameras when the 1.8x zoom level is commanded. Nevertheless, if, for some reason, settings convergence or implementation has not been completed by the time a 1.7x zoom level (e.g., the maximum field of view of the current camera 110b) is commanded, device 102 will force a switch when the zoom level reaches 1.7x.

[0098] The camera selector 420 provides a camera selection signal or other control data to the camera control module 440. The camera control module 440 reads video capture parameters from the cameras 110a, 110b and also sends commands or settings to set the video capture parameters. The video capture parameters can include, for example, which camera 110a, 110b is capturing image data, the frame capture rate (e.g., 24 frames per second (fps), 30 fps, 60 fps, etc.), exposure setting, image sensor sensitivity, gain (e.g., applied before or after image capture), image capture time (e.g., the effective “shutter speed” or duration that each scan line captures light during a frame), lens aperture size, subject distance (e.g., the distance of the focused subject from the camera), lens focus position or focal length, OIS status (e.g., whether OIS is enabled, the mode of OIS used, etc.), OIS lens position (e.g., horizontal and vertical offset, rotational position, etc.), applied OIS strength or level, etc. The camera control module 440 can set these and other video capture parameters for general video capture, such as exposure, frame rate, etc. 440 may also receive or have access to calibration data 232. A calibration procedure may be performed for each device, e.g., for device 102, as part of device manufacturing and quality assurance. The calibration data may indicate, for example, data for fine-tuning the relationship of cameras 110a, 110b to each other, the relationship between lens focus position and subject distance (e.g., focal plane distance) for different lens positions, etc.

[0099] The camera control module 440 can also enable and disable the cameras 110a, 110b at appropriate times to effect a camera switch indicated by the camera selector 420. For example, if input indicates a change in zoom level from 1.0x to 2.0x and data from the camera selector 420 indicates a change to the second camera 110b starting at 1.8x, the camera control module 440 can generate control instructions to effect this transition. Using stored information regarding the requested zoom speed and any limitations on the speed at which the zoom can be consistently performed, the camera control module 440 determines the time to effect the switch, e.g., the time at which the incremental or gradual zoom reflected in the image output reaches the 1.8x camera transition point. The time for this transition can be based on the duration or number of frames to continue capturing data with the first camera 110a until the digital zoom can smoothly reach the 1.8x zoom level, or the time specified by user input to reach the 1.8x zoom level.

[0100] In anticipation of a camera transition, the camera control module 440 can read video capture parameters from the current camera (e.g., the first camera 110a) and set corresponding video capture parameters for the camera to be used after the transition (e.g., the second camera 110b). This can include setting the same frame rate, the same exposure level, the same lens aperture, the same focal length, the same OIS mode or status (e.g., enabled or not), etc. Generally, before a transition between cameras, there is a period during which both cameras 110a, 110b are actively capturing video data simultaneously. For example, when switching from the first camera 110a to the second camera 110b, camera 110b will open and begin video capture before the zoom level reaches the threshold for switching to use the output of camera 110b for recorded video. The initial values ​​for the settings are based on the values ​​of the settings currently used for camera 110a. Camera 110b begins adjusting its operation toward the instructed settings, for example, to converge on the desired operating mode (e.g., the appropriate aperture setting, the correct focal length setting, the correct OIS setting, etc.). This process may include camera 110b or device 102 determining the appropriate settings, such as using an autofocus process to determine the correct focal length. After the settings have converged or calculated for camera 110b, the video stream used for recording will switch to the video stream output by camera 110b. In some cases, the parameters of cameras 110a and 110b may not be the same, but the parameters of camera 110b may nevertheless be set based on the parameters used for capture by camera 110a and may be set to promote or maintain consistency between the outputs. For example, the cameras 110a, 110b may not have the same aperture range available, and therefore the camera control module 440 may set equal or nearly equal exposure levels for the two cameras 110a, 110b, but with different combinations of settings for sensitivity / gain, capture time (e.g., shutter speed), and aperture.The camera control module 440 can activate the camera to be transitioned with the appropriate settings applied prior to the transition to use that camera for the final output frame recorded, so that image capture and the incoming video feed are available at or before the transition.

[0101] Video data is processed using an image processing module 450. This module 450 can receive captured video frames streamed from the camera 110a, 110b currently selected for video capture. Module 450 also receives sensor data from device position sensors, such as a gyroscope, an inertial measurement unit (IMU), or an accelerometer. Module 450 also receives (e.g., from the camera control module or memory) video capture parameter values ​​indicating the parameters used to capture the video frames. This can include metadata indicating OIS element positions, camera focus positions, subject distance, frame capture time, shutter speed / capture duration, etc. for each frame, for different portions of the frame, and even for specific scan lines or points in time within the frame capture or readout process (see FIG. 2, chart 252). Module 450 also receives data indicating the digital zoom level (which may be expressed, for example, as a magnification level, equivalent lens focal length, resulting field of view, level of cropping, etc.). Image processing module 450 applies transforms to captured image frames to remove artifacts due to rolling shutter, OIS system movement, focus breathing, etc. For example, module 450 can take video frames and transform or project them into the canonical space of the camera that captured the frames, where time-varying aspects of the frame's capture (e.g., OIS element movement, rolling shutter, etc.) are removed.

[0102] Module 450 can transform image frames into the canonical camera space of the first camera 110a, regardless of whether the image frames were captured by the first camera 110a or the second camera 110b. For example, for images captured using the second camera 110b, module 450 transforms the data from the canonical space of the second camera 110b to the canonical space of the first camera 110a by correcting for spatial differences between the positions of cameras 110a and 110b on the device and other factors, such as the focal position of camera 110b. This can align image data captured using the second camera 110b to the portion of the field of view of the first camera 110a so that the view of the scene is consistent across videos recorded from both cameras 110a and 110b. This technique is discussed further below with respect to FIGS. 5A-5C and 6.

[0103] The device 102 can include an EIS processing module 460 that receives and processes image data transformed into the canonical space of a primary camera. The "primary camera" refers to a camera among multiple cameras that is pre-designated as the reference for the other cameras. For example, the primary camera can be the first camera 110a with the widest field of view, and the output of any other camera (e.g., the second camera 110b) can be transformed or mapped to the canonical space of the first camera 110a. The module 450 maps the image data from both cameras 110a, 110b into a common, standardized canonical space that compensates for time-dependent variations within a frame (e.g., differences in capture time, device position, OIS position, etc.) for different scan lines of a frame. This significantly simplifies EIS processing by eliminating the need for the EIS processing module 460 to take into account time-varying capture characteristics within a frame. It also allows a single EIS processing workflow to be used for video data captured using either camera 110a, 110b. The EIS processing module 460 also receives a desired zoom setting, e.g., zoom level or field of view, potentially on a frame-by-frame basis. This enables the EIS processing module to apply an appropriate amount of stabilization for each frame according to the level of zoom used for that frame. As the zoom level increases and the image is magnified, the effects of camera motion are also magnified. Thus, the EIS processing module 460 can apply stronger stabilization as the zoom level increases to maintain a fairly consistent level of stability in the output video. The EIS processing module can use any or all of the EIS processing techniques described above with respect to FIGS. 2 and 3. Stabilization can be thought of as projecting image data from the canonical image space of the primary camera into a virtual camera space, where the image data is transformed to simulate an output as if the camera had a smoother motion trajectory during video capture than the actual camera would have.

[0104] After the EIS processing module 460 stabilizes the image data for a frame, the image data for that frame is output and / or recorded to a data storage device (e.g., a non-volatile storage medium such as flash memory). The zoom level that modules 410, 440 determine for a frame can be used to crop, upscale, or otherwise apply the required digital zoom level to the frame. As a result of these techniques, the device 102 can seamlessly transition between the two cameras 110a, 110b during capture, with the transition managed automatically by the device 102 based on the zoom level set by the user. Thus, the resulting video file can include video segments captured using different cameras 110a, 110b interspersed throughout the video file, with the data aligned and transformed to exhibit a smooth zoom transition while maintaining consistent EIS processing across segments from both cameras 110a, 110b.

[0105] The device 102 may include a video output and / or recording module 470. The video output and / or recording module 470 may be configured to receive the output of the EIS processing module 460 and provide it as a video file for storage locally on the device and / or remotely. Additionally or alternatively, the video output and / or recording module 470 may stream the output for display locally on the device 102 and / or remotely, for example, over a network to another display device.

[0106] 5A-5C illustrate examples of techniques for multi-camera video stabilization. In the examples described below, one of cameras 110a, 110b is designated as a primary camera, and the other is designated as a secondary camera. Through processing and transformations discussed below, when recorded video is captured using the secondary camera, the output of the secondary camera is mapped to the canonical space of the primary camera. For clarity, the first camera 110a is used as the primary camera, and the second camera 110b is used as the secondary camera. This means that in these examples, the camera with the wider field of view is designated as the primary camera. While this is desirable in some implementations, it is not required. Alternatively, these techniques may be used with a camera with a narrower field of view as the primary camera.

[0107] In general, a homography transformation is a transformation used to change from one camera space to another. The notation AHB represents a homograph that transforms points from camera space B to camera space A. A virtual camera refers to a synthetic camera view, such as the virtual camera from which the final scene (presented to the user and / or recorded in a video file) is generated. This effective camera position will typically be stabilized (e.g., appear as stationary as possible in position and orientation) for the entire duration of the video to provide as much temporal and spatial continuity as possible. As used herein, a "primary" camera is a primary camera used to define a frame of reference for generating the output video. The primary camera may be defined to be co-located spatially with the virtual camera, but need not be co-located temporally. In most of the examples below, the first camera 110a is used as the primary camera. A secondary camera is paired with the primary camera. The secondary camera is defined to be spatially separated from the virtual camera, and the output of the secondary camera will be warped into the virtual camera space if the secondary camera is the leading camera. In most of the examples below, the second camera 110b is used as the secondary camera. A "leading camera" refers to a camera currently open for video capture, e.g., the camera whose current image data is being used to generate the stored output video. A "trailing camera" is paired with a leading camera. The trailing camera is not currently open or active for video capture. Nevertheless, there may be a startup period during which the trailing camera begins capturing video in anticipation of acquiring the status of a leading camera before its output is actually used or stored in the output camera. A canonical camera is a conceptual camera with fixed intrinsic parameters that do not change over time; e.g., the canonical camera is not affected by the operation of an OIS or a voice coil motor (VCM) lens shift (e.g., for focusing). For each camera, there is a canonical camera (and corresponding image space), e.g., a canonical primary camera space and a canonical secondary camera space.

[0108] These parameters lead to two different use cases, depending on which of the cameras 110a, 110b is being used for capture. If the primary camera is leading, no spatial translation is required to map image data between the physical camera's space. Synthetic zoom and EIS processing can be performed simply with the EIS processing on the primary camera. On the other hand, if the secondary camera is leading, the system applies spatial translation to map the output to the primary camera's viewpoint for consistency across switches between captures with the two cameras 110a, 110b. This takes into account changes in focus in the scene that change focal length or subject distance, again maintaining consistency between the outputs of the two cameras 110a, 110b.

[0109] One of the challenges of providing digital zoom during recording is efficiently incorporating EIS processing with the digital zoom function, especially for the types of incremental or continuous zoom using multiple cameras described above. During video capture, device 102 can record video with EIS processing enabled to obtain a series of temporally stabilized image sequences. Device 102 can concatenate homographies for the zoom and EIS functions to achieve this effect.

[0110] Homographies for EIS and homographies for digital zoom using multiple cameras serve different purposes. A homograph for EIS transforms image data from the image sensor's current (e.g., actual) output frame (denoted by the subscript "R" or "real") to a virtual frame (denoted by the subscript "V" or "virt"). In the various equations and expressions below, the variable t (e.g., lowercase t) is the frame timestamp, which is typically related to the time of capture of the center scan line of the frame. However, because different scan lines are captured at different times, the timestamps for other scan lines may vary. When referring to the time of an individual scan line (e.g., to account for a time different from the primary time of capture for the center scan line of the frame), the time is t L , which represents the time of the capture timestamp of scanline L. When using a camera with a rolling shutter, the timestamp for each scanline of the frame will be slightly different, and therefore the term t L may be slightly different for different scan lines, and the camera position may also be slightly different for different scan lines. E Note that n refers to the extrinsic translation between the primary camera 110a and the secondary camera 110b, and does not represent any timestamp. T is the plane norm discussed below, which is independent of the external translation and time terms.

[0111] The EIS homography is V H R or H eis A homography from EIS is constructed to transform data from the current or "real" frame to the virtual frame by unprojecting points from the current or "real" frame into 3D space and then projecting them back into virtual space. This homography can be expressed as:

[0112]

number

[0113] In Equation 1, R V represents the 3x3 rotation matrix of the virtual camera, and R C represents the 3 × 3 rotation matrix of the currently used real cameras 110a to 110b, and R C and R V Both can be obtained from camera position data (e.g., gyroscope data). K is the intrinsic matrix, and K V is the intrinsic matrix of the virtual camera, and K C is the intrinsic matrix of the current camera (e.g., whichever camera 110a-110b is being used). The camera intrinsic data (e.g., based on the camera geometry, calibration data, and camera characteristics) may be stored in and retrieved from one or more data stores 230. The intrinsic matrix may be expressed as Equation 2 below:

[0114]

number

[0115] In Equation 2, f is the focal length and o x and o y is the principal point. In the above formula, R C , R V and K. C -1 R is time-dependent and is shown as a function of time t. These values ​​can hold or be based on information from previous frames or previous output frame generation processes to ensure temporal continuity. V is a 3x3 rotation matrix for virtual space, calculated based on a filter of the trajectory of past virtual camera frames, and potentially for several future frames as well, if delay and "look ahead" strategies are used. Cis a 3x3 rotation matrix for the current frame, calculated based on data from the gyroscope or device position sensor 220 that transforms the current frame to the first frame. C -1 is calculated based on the optical principal center and the current OIS value.

[0116] The homography for zooming transforms from the view of one camera 110a to the view of another camera 110b. For example, this homography transforms from the current frame of the primary camera 110a to the frame of the secondary camera 110b, main H sec This homography may be computed using the four-point technique to compute this homography, but by simplifying this homography as a Euclidean homography, the matrix itself can also be decomposed as follows:

[0117]

number

[0118] Point P on the image of the secondary camera (for example, the telephoto camera 110b) sec is expressed as a corresponding point P on the image of the main camera (e.g., the wide-angle camera 110a). main Given a Euclidean homography transformation that brings S to any scalar in S, this homography matrix can be decomposed as follows:

[0119]

number

[0120] In this formula, Ext(t) is the external transformation, which is a matrix that depends on the depth of the plane:

[0121]

number

[0122] The resulting combination is shown below:

[0123]

number

[0124] In Equation 4, n T is the plane norm (i.e., the vector normal to the focal plane), and D is the depth or distance at which the focal plane is located. E and T E indicate the external rotation and translation between the primary camera 110a and the secondary camera 110b, respectively.

[0125] Variable K main and K. sec are the intrinsic matrices of the primary camera 110a and secondary camera 110b, which have the same format as the EIS homography described above. main , K. sec , D is time-dependent. However, unlike the corresponding version in the EIS formulation, K main , K. sec , and D does not retain past information in the zoom formulation. K main and K. secBoth are calculated based on the current VCM and OIS values ​​(e.g., using OIS position data 248a-248b), which change from frame to frame. Variable D is related to the focus distance value during autofocus and corresponds to the subject distance, which may change over time during video capture and recording. As described above, subject distance refers to the distance of the focal plane from device 102 relative to the current focus selected for the camera. Subject distance can be determined, for example, using the camera's focus position from focus position data 246a-246b and calibration data, such as a look-up table, that shows the correspondence of focus setting or focus element position to subject distance, which indicates how far a focused subject is from the camera sensor plane. Given the positional offset between cameras 110a, 110b, the relationship between captured images may vary somewhat depending on subject distance, and the transformation can take these effects into account using subject distance, lens focus position, and / or other data.

[0126] The above two homography decompositions can be summarized as a series of operations: (1) unproject from the source camera, (2) transform from the source camera to the world 3D reference frame, (3) transform from the world to the target 3D reference frame, and (4) reproject back to the target camera, which are summarized in Table 1 below and also discussed with respect to Figures 5A-5C and 6.

[0127] [Table 1]

[0128] The next section describes techniques for combining the zoom homography and the EIS homography in an effective and computationally efficient manner. One technique that can be used is to set one camera, typically the one with the widest field of view, as the primary camera and map the processed output with respect to the primary camera.

[0129] If the primary camera (e.g., wide-angle camera 110a) is used for video capture, main H sec is the identity. As a result, the combined homography is simply the EIS homography H, as shown in Equation 7. eis It could be.

[0130]

number

[0131] term R D indicates the device rotation calculated based on data from a gyroscope or other motion sensor, and has the same value for both the primary camera 110a and the secondary camera 110b.

[0132] Specifically, the three homography transformations are: ● sec_can H sec : Transformation from real secondary camera to canonical secondary camera space with (0,0) OIS motion, 0 rolling shutter time, fixed focal length, and rotation about the frame center.

[0133]

number

[0134] where R sec -1 (t)*K sec -1 (t) is taken from each scan line or frame center, depending on what is needed, and this will work with the given OIS / VCM values. sec_can (t) is the R sec -1 Contrary to (t), this is a rotation about the frame center. ● main_can H sec_can : Transformation from canonical secondary camera to canonical primary camera space.

[0135]

number

[0136] In this formula, Ext(t) is the external transformation, which is a matrix that depends on the object distance (e.g., the depth in space of the focal plane from the camera sensor):

[0137]

number

[0138] where n T is the plane norm, and D is the plane n T is the depth of R E and T E is the external rotation and translation between the secondary and primary cameras. ● virt H main_can : A transformation from the canonical secondary camera to the stabilized virtual camera space, and this rotation is filtered by the EIS algorithm.

[0139]

number

[0140] R main_can =R sec_can Note that ∇x is the current actual camera rotation at the frame center, since the primary and secondary cameras are rigidly attached to each other.

[0141] Concatenating the three homographies gives the final homography:

[0142]

number

[0143] Here, the middle term main_can H sec_can comes from the zoom process. If the primary camera is ahead, the final homography simplifies to just the homography for EIS processing. The above formula becomes:

[0144]

number

[0145] this is, main_can H sec_can Setting the identity = H, the above equation is eis becomes the formula:

[0146]

number

[0147] From the above section, we generally have two formulas: When the secondary camera is leading, the formula is:

[0148]

number

[0149] When the main camera is leading, the formula is:

[0150]

number

[0151] From an engineering point of view, during EIS realization there are no secondary cameras, therefore all virtual cameras are located on the current leading camera, which means the following for the original EIS pipeline: When the main camera is ahead,

[0152]

number

[0153] When the secondary camera is ahead,

[0154]

number

[0155] From the basic formula of EIS,

[0156]

number

[0157]

number

[0158] term K V is always defined to be the same as the current camera, and by definition, K v_main and K. v_sec The only difference between is the FOV between the two cameras (the virtual camera is placed in the center of the image, so the principal point is the same).

[0159] To accommodate the above case, the system can intentionally ensure that the fields of view in the primary and secondary cameras match each other at the switch point. This field of view alignment is done efficiently through hardware cropping that scales the field of view from the secondary camera to the primary camera prior to any operation, and the matrix S is used in the following equation:

[0160]

number

[0161] From the software side, the homography formula fits to the following:

[0162]

number

[0163] The final applied homography might look like this:

[0164]

number

[0165] The transformation maps from the canonical subspace to a normalized canonical subspace that has the same field of view (e.g., scale) as the canonical subspace, but in which all translations / rotations caused by camera extrinsic factors are neutralized.

[0166] When a camera considered a secondary camera (e.g., telephoto camera 110b) is used for video capture, a series of transformations are used to convert or project the view from the secondary camera onto the view from the primary camera, allowing consistent image characteristics to be maintained over the course of capture using different cameras 110a, 110b. These transformations can include (1) deprojecting from the source camera (e.g., second camera 110b) to remove camera-specific effects such as OIS motion and rolling shutter; (2) converting to a canonical frame of reference, such as the three-dimensional world; (3) converting from the canonical frame of reference to the primary camera reference (e.g., aligning to the frame that would be captured by the first camera 110a); and (4) reprojecting from the primary camera's 110a frame onto a virtual camera frame with electronic stabilization applied. Depending on the order of the transformations, this technique can be performed with rectification first or stabilization first, resulting in different performance results. 5A-5C show different techniques for converting the captured image from the second camera 110b into a stabilized view that is aligned with and consistent with the stabilized view generated for the first camera 110a.

[0167] 5A illustrates an exemplary technique for multi-camera video stabilization that performs rectification first. This technique first rectifies the view from the second camera 110b (e.g., the telephoto camera in this example) to the first camera 110a (e.g., the wide-angle camera in this example), and then applies stabilization. This technique offers the advantage of being represented by a simple concatenation of the EIS homography and the primary secondary-to-primary homography, which increases the efficiency of processing the video feeds. This is shown in the following equation:

[0168]

number

[0169] FIG. 5B illustrates an exemplary technique for multi-camera video stabilization that performs stabilization first. This option stabilizes the image from the secondary camera frame into a stabilized virtual primary camera frame, then applies rectification to project the virtual secondary camera frame onto the primary camera frame. However, this technique is not as simple as the approach of FIG. 5A because the transformation from the secondary camera to the primary camera is defined in the primary camera frame coordinate system. Therefore, to perform the same transformation, the stabilized virtual primary camera frame must be rotated back from the virtual secondary frame. This is expressed in the following equation, where "virt_main" refers to the primary camera's virtual frame (stabilized), "virt_sec" refers to the secondary camera's virtual frame (stabilized), "real_main" refers to the primary camera's real frame (unstabilized), and "real_sec" refers to the secondary camera's real frame:

[0170]

number

[0171] By concatenating the items, the result is the same overall transformation as in Figure 5A. Even though there are different geometric definitions, the overall effect on the image is the same.

[0172]

number

[0173] 5C illustrates an example technique for multi-camera video stabilization that first aligns to a canonical camera. Using the geometric definitions discussed above, the implementation can be made more efficient by aligning from the current real camera view to a canonical camera view, which is defined as a virtual camera of the current leading camera, but with a fixed OIS (OIS_X=0, OIS_Y=0) and a fixed VCM (VCM=300).

[0174] The result is similar to the formula in the first case of Figure 5A. For example, this gives:

[0175]

number

[0176] However, in this case, K main -1 (t)K main (t) Instead of inserting pairs, K main_can -1 (t)*K main_can (t) is inserted, where K main_can denotes the intrinsic properties of the canonical main camera. eis and main H' sec is K main_can It is calculated using

[0177]

number

[0178] This approach offers significant advantages as the system does not need to query for metadata from both the secondary and primary cameras simultaneously.

[0179] The combined homography, H combinedcan be used with the canonical position representation and scanline-by-scanline processing. First, we will explain the canonical position. eis and main H sec From the original definition of , the system can be considered as the K sec and K main The system also uses OIS and VCM information for both the term K used in the EIS homography. C is also used.

[0180] In the following equations, certain variables are time-dependent variables that depend on the current OIS and VCM values ​​for the primary and secondary cameras, respectively. These time-dependent variables are K main -1 (t), K main (t), and K sec -1 (t). The time dependency of these terms means that the system would need to stream both OIS / VCM metadata of both cameras 110a, 110b simultaneously. However, using a canonical representation can significantly reduce the data collection requirements. For example, the term K sec C and K main C An example representation is shown below, where H is defined as a canonical camera model with both the OIS and VCM positioned at a predetermined standard or canonical position, where the canonical position is a predetermined position with OISX=OISY=0 and VCM=300. combined = H eis * main H sec For , the expression can be decomposed as follows:

[0181]

number

[0182] In the above formula, the term C sec(t) is a correction matrix that transforms the current camera view to the canonical camera view (e.g., removes the effect of the OIS lens position). Because the digital zoom homography is only active when capturing video from the second camera 110b (e.g., the telephoto camera), there is one correction matrix term C that transforms the current secondary camera view to the canonical secondary camera view. sec (t) is needed. In this new formula, K main C and K sec C are both constant over time for the homography due to digital zoom:

[0183]

number

[0184] As a result, the only time-dependent component is D(t), which depends on the subject distance for focus, which can be the distance from the camera to the focal plane. Therefore, the only new metadata needed will be the subject distance.

[0185] Next, we will explain the correction for each scan line. The correction for each scan line is done by adding another matrix S(t L , L). This additional matrix can be used to make corrections that are scanline specific, with potentially different adjustments indicated for each scanline L. The per-scanline correction matrix S(t L , L) is the current time t of the scanline (e.g., to obtain gyroscope sensor data at the appropriate time when the scanline was captured). L , and the scan line number L (used to correct for distortion or shift, for example, to align each scan line with the central scan line). As a result, the matrix S(t L , L) is the correction for each scan line L and its corresponding capture time t L may include:

[0186] From the original decomposed equation, adding the correction matrix line by line gives:

[0187]

number

[0188] The final homography can be determined by adding the canonical positions and the scanline-wise terms:

[0189]

number

[0190] In this example, H C eis = K V *R V (t)* R D -1 (t) *K main C-1 represents the stabilization matrix for the canonical main camera only. In addition, the term main H C sec = K main C * (R E -1 - D(t)*T E *n T ) * K sec C-1 represents the homography from the canonical minor view to the canonical major view. sec (t) is the homography that transforms the current secondary camera view into the canonical secondary camera view. The term S(t L , L) is a line-by-line correction that brings the deformation of each line to a reference position on the center line.

[0191] 6 illustrates example transformations that can be used to efficiently provide multi-camera video stabilization. The combination of EIS and continuous zoom can be interpreted as a combination of three homographies or transformations.

[0192] In this example, image frame 611 represents an image frame captured by second camera 110b. Through multiple transformations depicted in the figure, the image data is processed to remove visual artifacts (e.g., rolling shutter, OIS movement, etc.), aligned with the field of view of first camera 110a, and stabilized using the EIS techniques described above. Separate transforms 610, 620, 630, as well as various images 611-614, are shown for illustrative purposes. Implementations can combine the discussed operations and functionality without the need to separately generate intermediate images.

[0193] A first transform 610 operates on a captured image 611 from the actual secondary camera 110b (e.g., a telephoto camera) to transform the image 611 into a canonical secondary camera image 612. The transform 610 adjusts the image data of the image 611 so that the canonical secondary camera image 612 does not include OIS movement, does not include a rolling shutter, includes a fixed focal length, and involves rotation about the center of frame. In some implementations, the first transform 610 may adjust each image scan line of the image 611 individually to remove or reduce the effects of movement of the secondary camera 110b over the course of capturing the image 611, changes in focal length over the course of capturing the image 611, movement of the OIS system during the capture of the image 611, etc. This transform 610 sec_can H sec , i.e., a transformation from the second camera view (“sec”) to the canonical second camera view (“sec_can”). While OIS stabilization may have been used during the capture of image 611, neither image 611 nor image 612 was stabilized using EIS processing.

[0194] The second transform 620 is a transformation from the canonical second camera image 612 to the canonical first camera image 613. This can maintain the image data within the same canonical space of the image 612, but align the image data to the scene captured by the primary camera, e.g., the first (e.g., wide-angle) camera 110a. This transform 620 can correct for spatial differences between the second camera 110b and the first camera 110a. The two cameras 110a, 110b are located on a phone or other device with a spatial offset between them and potentially other differences in position or orientation. The second transform 612 can correct for these differences to project the image 612 onto the corresponding portion of the field of view of the first camera 110a. Additionally, the difference between the views of the cameras 110a, 110b can vary depending on the current depth of focus. For example, one or both of the cameras 110a, 110b may experience focus breathing, which adjusts the effective field of view depending on the focal length. The second transform 620 can take these differences into account, allowing the device 102 to use focal length to fine-tune the alignment of the image 612 relative to the primary camera canonical field of view. Typically, the same canonical parameters, such as for OIS position, are used for both the primary and secondary camera canonical representations, but if there are differences, these can be corrected for using the second transform 620.

[0195] A third transform 620 applies EIS to image 613 to generate stabilized image 614. For example, this transform 620 can convert image data from a canonical primary camera view into a virtual camera view in which the camera position is smoothed or filtered. For example, this transform 630 can be the second projection discussed above with respect to FIG. 3, in which the virtual camera position is filtered over time, changes in position between frames are used, motion is potentially allowed for due to blur, the possibility of panning versus accidental motion is considered, adaptation of future motion or adjustments to fill the output frame is performed, etc. The EIS processing can be performed using the current frame being processed as well as a window of previous frames and a window of future frames. Of course, the EIS processing can lag behind the image capture timing that collects the “future frames” needed for processing.

[0196] As mentioned above, the transforms 610, 620, 630 can be combined or integrated for efficiency, without the need to generate intermediate images 612 and 613. Rather, the device 102 may determine the appropriate transforms 610, 620, 630 and directly generate a stabilized image 614 that is aligned with and consistent with the stabilized image data generated from the image captured using the first camera 110a.

[0197] The example in Figure 6 shows a transformation from a video frame from a second camera 110b to the stabilized space of the first camera 110a, which is used when the zoom level corresponds to a field of view equal to or smaller than that of the second camera 110b. When the first camera 110a is used, for example, when the field of view is larger than that of the second camera 110b, only two transformations are required. Similar to the first transformation 610, a transformation is applied to remove time-dependent effects from the captured image, such as conditions that change for different scan lines of the video frame. Similar to transformation 610, this compensates for rolling shutter, OIS position, etc. However, the transformation directly projects the image data into the canonical camera space of the primary camera (e.g., camera 110a). From the primary camera's canonical camera space, only the EIS transformation, e.g., transformation 630, is required. Thus, when image data is captured using the primary camera, there is no need to relate the spatial characteristics of the two cameras 110a, 110b because the entire image capture and EIS processing is performed consistently in the frame of reference for the primary camera, e.g., from a canonical, non-time-dependent frame of reference for the primary camera.

[0198] The process in FIG. 6 illustrates various processes that can be used to convert or map the output of the second camera 110b to the stabilized output of the first camera 110a. This provides consistency in view when transitioning between using video captured by different cameras 110a, 110b during video capture and recording. For example, the video from the second camera 110b is aligned with the video from the first camera 110a, not only positioning the field of view in the correct position but also aligning the EIS characteristics. As a result, switching between cameras during image capture (e.g., by digital zoom above or below a threshold zoom level) can be done with EIS active, aligning the video feeds in a way that minimizes or avoids jerky motion, sudden image or view shifts, and visible video smoothness shifts (e.g., changes in EIS application) at the transition points between cameras 110a, 110b. Another advantage of this technique is that the video from camera 110a can be used without any adjustments or processing for consistency with the second camera 110b. Only the output of the second camera 110b is adjusted and aligned consistently with the view and characteristics of the first camera 110a.

[0199] The device 102 can also adjust video capture settings during video capture to better match the characteristics of video captured using different cameras 110a, 110b. For example, when transitioning from capturing using a first camera 110a to capturing using a second camera 110b, the device 102 can determine the characteristics, such as focal length, OIS parameters (e.g., whether OIS is enabled, the intensity or level of OIS applied, etc.), exposure settings (e.g., ISO, sensor sensitivity, gain, frame capture time or shutter speed, etc.), etc., used for the first camera 110a immediately before the transition. The device 102 can then cause the second camera 110b to use those settings, or settings that provide equivalent results, for the capture for the second camera 110b.

[0200] This may involve making setting changes prior to the transition to include video from the second camera 110b in the recorded video. For example, to provide time to make adjustments during operation, the device 102 may detect that a camera switch is appropriate or necessary and, in response, determine the current settings of the first camera 110a and instruct the second camera 110b to start using settings equivalent to those used by the first camera 110a. This may provide sufficient time, for example, to power on the second camera 110b, activate and achieve stabilization of the OIS system for the second camera 110b, adjust the exposure of the second camera 110b to match that of the first camera 110a, set the focus position of the second camera 110b to match the focus position of the first camera 110a, etc. When the second camera 110b is operating in an appropriate mode, e.g., with video capture parameters consistent with those of the first camera 110a, the device 102 switches from using the first camera 110a to using the second camera 110b for video capture. Similarly, when transitioning from the second camera 110b to the first camera 110a, the device 102 can determine the video capture parameter values ​​used by the second camera 110b and can set the corresponding video capture parameter values ​​for the first camera 110a before making the transition.

[0201] Several implementations have been described. Nevertheless, it will be understood that various modifications may be made without departing from the spirit and scope of the present disclosure. For example, the flow illustrated above may be used in various forms with steps rearranged, added, or removed. As another exemplary modification, while the above description primarily describes processing of image data being performed while video is being captured by the first and / or second camera, it will be understood that in some implementations, the first transformation to the second canonical reference space for the second camera, the second transformation from the second canonical reference space to the first canonical reference space for the first camera, and the third transformation for applying electronic image stabilization to the image data in the first canonical reference space for the first camera may instead be applied later (by device 102 or another, e.g., remote, device), e.g., when video is no longer being captured.

[0202] All of the embodiments of the present invention and the functional operations described herein may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Embodiments of the present invention may be implemented as one or more computer program products, e.g., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or controlling the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter providing a machine-readable propagated signal, or one or more combinations thereof. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, by way of example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may include code that creates an execution environment for the computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. A propagated signal is an artificially generated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a suitable receiver apparatus.

[0203] A computer program (also known as a program, software, software application, script, or code) can be written in any type of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.

[0204] The processes and logic flows described herein may be implemented by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows, as well as the implementation of the apparatus, may also be performed by special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0205] Processors suitable for executing computer programs include, by way of example, both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic, magneto-optical, or optical disks, or be operatively coupled to receive or transfer data from or to the one or more mass storage devices, or both. However, a computer need not have such devices. Furthermore, a computer can be incorporated into another device, such as a tablet computer, a mobile phone, a personal digital assistant (PDA), a mobile audio player, or a global positioning system (GPS) receiver, to name just a few. Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0206] To provide for user interaction, embodiments of the present invention can be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices can also be used to provide for user interaction; for example, feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; input from the user can be received in any form, including acoustic input, speech input, or tactile input.

[0207] Embodiments of the present invention may be implemented in a computing system that includes a back-end component, e.g., a data server; a middleware component, e.g., an application server; a front-end component, e.g., a client computer having a graphical user interface or web browser through which a user can interact with an implementation of the present invention; or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, e.g., a communications network. Examples of communications networks include local area networks ("LANs") and wide area networks ("WANs"), e.g., the Internet.

[0208] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

[0209] While this specification contains many details, these should not be construed as limitations on the scope of the invention or what may be claimed, but rather as descriptions of features specific to particular embodiments of the invention. Certain features that are described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, while features may be described above as operative in a combination, and may even initially be claimed as such, one or more features from a claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a subcombination or a variation of the subcombination.

[0210] Similarly, while operations are shown in a particular order in the figures, it should not be understood that such operations need to be performed in the particular order shown, or sequential order, or that all of the shown operations need to be performed, to achieve desirable results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated into a single software product or packaged into multiple software products.

[0211] Although specific embodiments of the present invention have been described, other embodiments are within the scope of the following claims. For example, the steps recited in the claims can be performed in a different order and still achieve desirable results.

Claims

1. 1. A method comprising: A method includes a video capture device having a first camera and a second camera providing a digital zoom function that allows user-specified magnification changes within a digital zoom range during video recording, the video capture device being configured to (i) use video data captured by the first camera over a first portion of the digital zoom range, and (ii) use video data captured by the second camera over a second portion of the digital zoom range, the method further comprising: a first transformation to a second canonical reference space for the second camera of the video capture device, while capturing video using the second camera to provide a zoom level in the second portion of the digital zoom range, processing image data captured using the second camera by applying a set of transformations including: (i) a first transformation to a second canonical reference space for the second camera; (ii) a second transformation from the second canonical reference space for the second camera to a first canonical reference space for the first camera; and (iii) a third transformation for applying electronic image stabilization to image data in the first canonical reference space for the first camera.

2. The method further comprises:

2. The method of claim 1, comprising, while capturing video using the first camera of the video capture device to provide a zoom level in the first portion of the digital zoom range, processing image data captured using the second camera by applying a set of transformations including: (i) the first transformation to the first canonical reference space for the first camera; and (ii) the third transformation for applying electronic image stabilization to image data in the first canonical reference space for the first camera.

3. 3. The method of claim 1 or 2, wherein the first camera and the second camera have different fields of view, and (i) the field of view of the second camera is contained within the field of view of the first camera, or (ii) the field of view of the first camera is contained within the field of view of the second camera.

4. The method of any one of claims 1 to 3, wherein the first camera and the second camera each include a fixed focal length lens assembly.

5. 5. The method of claim 1, wherein the second canonical reference space for the second camera and the first canonical reference space for the first camera are conceptual camera spaces defined by a predetermined, fixed set of camera-specific properties, such that projecting image data of the second canonical reference space into the first canonical reference space removes time-dependent effects during capture of video frames.

6. the first camera includes an optical image stabilization (OIS) system, and the first canonical reference space for the first camera is represented with image data having a consistent predetermined OIS position; or 6. The method of claim 5, wherein the second camera includes an optical image stabilization (OIS) system, and the second canonical reference space for the second camera is one in which image data is represented with a consistent, predetermined OIS position.

7. the first camera provides image data that progressively captures image scan lines of an image frame, and the first canonical reference space for the first camera is corrected so that image data is free of distortion resulting from the progressive capture of the image scan lines; or 7. The method of claim 6, wherein the second camera provides image data that progressively captures image scan lines of an image frame, and the second canonical reference space for the second camera is corrected so that image data is free of distortion resulting from the progressive capture of the image scan lines.

8. 8. The method of claim 1, wherein the second transformation aligns a field of view of the second camera with a field of view of the first camera and adjusts for a spatial offset between the first and second cameras.

9. 9. The method of claim 1, wherein the first transform, the second transform, and the third transform each have a corresponding homography matrix, and wherein processing the image data comprises applying the homography matrix.

10. The method further comprises: receiving a user input indicating a change in zoom level to a particular zoom level in the second portion of the digital zoom range during capture of video data using the first camera and processing of the video data from the first camera to apply electronic image stabilization; In response to receiving the user input, recording a sequence of video frames in which the magnification of the video frames captured using the first camera is incrementally increased until a predetermined zoom level is reached; initiating video capture using the second camera; and recording a second sequence of video frames captured using the second camera, the second sequence of video frames providing the predetermined zoom level and providing increasing magnifications of the video frames captured using the second camera until the particular zoom level is reached.

11. the first transformation includes a plurality of different adjustments for different scanlines of image data captured using the second camera; the second transformation is determined based at least in part on a focal length of the second camera; 11. The method of claim 1, wherein the third transformation comprises, for each particular one of the video frames, electronic image stabilization using one or more video frames before the particular one of the video frames and one or more video frames after the particular one of the video frames.

12. the second camera has a smaller field of view than the first camera; The method further comprises: receiving a user input indicating a change in zoom level for video capture during image capture using the first camera; and determining, in response to receiving the user input, whether the changed zoom level is greater than or equal to a predetermined transition zoom level, the predetermined transition zoom level representing a field of view that is smaller than the field of view of the second camera.

13. The method further comprises: storing data indicating (i) a first transition zoom level for transitioning from video capture using the first camera to video capture using the second camera, and (ii) a second transition zoom level for transitioning from video capture using the second camera to video capture using the first camera, the first transition zoom level being different from the second transition zoom level, the method further comprising:

13. The method of claim 1, comprising determining whether to switch between cameras for video capture by: (i) comparing a requested zoom level to the first transition zoom level when the requested zoom level corresponds to a decrease in field of view; and (ii) comparing the requested zoom level to the second transition zoom level when the requested zoom level corresponds to an increase in field of view.

14. The method further comprises: determining, during recording of a video file, to switch from capturing video using a particular one of the first camera and the second camera to capturing video using another one of the first camera and the second camera; In response to determining to switch, determining values ​​of video capture parameters used for image capture using the particular camera; setting values ​​of video capture parameters for the other camera based on the determined video capture parameters; after setting the values ​​of the video capture parameters for the other camera, initiating video capture from the second camera and recording the captured video from the second camera to the video file; 14. The method of claim 1, wherein setting the video capture parameters includes adjusting one or more of exposure, image sensor sensitivity, gain, image capture time, aperture size, lens focal length, OIS status, or OIS level for the second camera.

15. 1. A video capture device, comprising: a first camera having a first field of view; a second camera having a second field of view; one or more position or orientation sensors; one or more processors; and one or more data stores storing instructions that, when executed by said one or more processors, cause said instructions to perform the method of any one of claims 1 to 14.

Citation Information

Patent Citations

  • Device for generating multi-viewpoint image, image processor, method and computer program

    JP2003179800A

  • Video camera with zoom lens

    JP2012083685A

  • Systems and methods for digital video stabilization via constraint-based rotation smoothing

    JP2018085775A

  • Systems and methods for implementing seamless zoom function using multiple cameras

    US20170230585A1