Multi-camera video stabilization

The method addresses inconsistent image stabilization across zoom levels and camera transitions in multi-camera devices by using EIS and digital zoom, ensuring seamless video transitions and improved quality.

JP2026090246APending Publication Date: 2026-06-02GOOGLE LLC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
GOOGLE LLC
Filing Date
2025-12-26
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Devices with multiple camera modules face challenges in maintaining image stabilization consistency across different zoom levels and camera transitions, leading to undesirable camera shake and disruptions in video quality.

Method used

A method involving electronic image stabilization (EIS) and multi-camera digital zoom techniques, utilizing a canonical camera space to align and stabilize image data from different cameras, ensuring seamless transitions and consistent stabilization parameters.

Benefits of technology

The method provides smooth and uninterrupted video transitions between camera modules, enhancing video quality by maintaining temporal and spatial continuity while efficiently managing computational resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026090246000001_ABST
    Figure 2026090246000001_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, and apparatus for multi-camera video stabilization, including a computer program encoded on a computer storage medium. [Solution] A video capture device has a first camera and a second camera. The video capture device provides a digital zoom function during video recording. The video capture device is configured to use video data from different cameras across different parts of the digital zoom range. The video capture device can process image data captured using the second camera by applying a set of transformations including (i) a first transformation to a canonical reference space for the second camera, (ii) a second transformation to a canonical reference space for the first camera, and (iii) a third transformation for applying electronic image stabilization to image data in the canonical reference space for the first camera.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background Some devices, such as smartphones, contain multiple camera modules. These cameras can be used to record still images or videos. In many situations, camera shake and other device movements can degrade the quality of captured images and videos. As a result, some devices include image stabilization features to improve the quality of recorded image data. [Overview of the Initiative] [Problems that the invention aims to solve]

[0002] overview In some implementations, a device includes a multi-view camera system, for example, a device with multiple camera modules having different fields of view. This device provides video stabilization for video captured using one or more of the device's camera modules. This may include providing varying levels of image stabilization at different zoom levels (e.g., magnification or resizing) and managing the image stabilization behavior to be consistent across different zoom levels and different camera modules. Even if a camera module has a fixed field of view, the device may provide zoom functionality by, for example, using digital zoom techniques to perform or approximate continuous or smooth zooming along the range of the field of view. The device may provide features for capturing consistently stabilized video across a range of zoom settings, including when transitioning between video captures of different camera modules during video recording. In some implementations, the device detects when to transition between cameras for image capture during video recording. The device performs the transition and processes the captured data to produce output video that maintains or smoothly adjusts parameters such as field of view, image stabilization, exposure, focus, and noise over the duration of the transition. This allows the system to generate video that uses video from different parts of different cameras, with user-unobtrusive camera transitions, without, for example, monocular parallax, visible stutter, visible zoom pauses, or other glitches.

[0003] The primary goal of this system is to improve the smoothness of video scenes presented to the user. In other words, the system attempts to achieve both temporal continuity of scenes (e.g., reducing undesirable camera shake over time) and spatial continuity of scenes (e.g., reducing differences between videos captured from different cameras). This involves two key techniques: (1) electronic image stabilization (EIS) which effectively provides temporal smoothing of scenes shown in the video feed by providing scene continuity over time on a single camera, and (2) enhanced multi-camera digital zoom (e.g., incremental, incremental, or substantially continuous zoom using the outputs of multiple cameras) which effectively provides a spatial smoothing method to provide scene continuity between different cameras while avoiding disruption or interruptions near camera transitions.

[0004] EIS and multi-camera digital zoom can be efficiently combined by using various techniques and transformations, further discussed below, including the use of a "canonical" camera space to represent image data. A canonical space can represent a conceptual camera view with fixed, intrinsic characteristics that do not change over time. For example, a canonical space may be unaffected by factors such as optical image stabilization (OIS) or voice coil motor (VCM) position and rolling shutter effect.

[0005] For example, a device may include a first camera and a second camera that provide different fields of view of a scene. The device can enable a zoom function that allows the user to smoothly change the zoom or magnification represented by the output video, even if one or more of the camera modules have a fixed field of view. In some implementations, using two or more camera modules with fixed focal length lenses, continuous zoom can be simulated over a certain zoom range by (1) using digital zoom (e.g., cropping and / or magnification) based on the image from the first camera for a first portion of the zoom range, and (2) using digital zoom based on the image from the second camera for a second portion of the same zoom range. To improve the overall quality of the video, image stabilization processing can be dynamically adjusted for the current level of digital zoom applied. For example, each change in zoom along the simulated continuous zoom range may have a corresponding change in the image stabilization parameters.

[0006] The device's image processing can manage transitions between image captures from different cameras, providing a virtually seamless transition while maintaining consistency in the EIS application and other image capture parameters such as focal length and exposure. The zoom function may include, for example, digital zoom using the output of the first camera as the device increasingly crops the image captured by the first camera. Then, when the zoom threshold level is reached and the area being zoomed in is within the field of view of the second camera, the device switches to recording video captured using the second camera.

[0007] To ensure that recorded video provides a smooth transition between the outputs of different cameras, the device may use a series of transformations to associate the output of a second camera with the output of a first camera. These transformations may be implemented using homography matrices or in other forms. In some implementations, the transformations involve mapping images from the second camera to canonical camera space by removing camera-specific time-dependent contributions such as rolling shutter effects and OIS lens motion from the second camera. The second transformation can project image data in the canonical image space of the second camera onto the canonical image space of the first camera. This can align the field of view of the second camera with the field of view of the first camera and account for spatial differences (e.g., offsets) between cameras in the device. Electronic image stabilization (EIS) processing can then be applied to the image data in the canonical image space of the first camera. This series of transformations provides a much more efficient processing technique than, for example, attempting to associate and align EIS-processed second camera image data with EIS-processed first camera image data. EIS processing can be performed in a single camera space or reference frame, even though the images are captured using different cameras with different fields of view, different intrinsic characteristics, etc., during image capture. The output of the EIS processing in the canonical image space of the first camera can then be provided for storage (locally and / or remotely) as a video file, and / or streamed (locally and / or remotely) for display.

[0008] These technologies can apply a level of image stabilization adjusted to the current zoom level to more effectively control camera shake and other unintended camera movements. In addition, the ability to smoothly transition between cameras during video capture can increase the resolution of the video capture without interruption. For example, as digital zoom is increasingly applied to the output of cameras with a wider field of view, the resolution tends to decrease. As digital zoom increases, the resulting output image represents a smaller portion of the image sensor, and therefore less of the image sensor is used to produce the output image. The first camera uses pixels. The second camera may have a lens with a narrower field of view, allowing the narrower field of view to be captured across the entire image sensor. When the video is zoomed in to the point where the output frame enters the field of view of the second camera, the camera can transition the video capture to use the image data captured by the second camera. Thus, while continuing to capture and record video files, and with EIS continuously and consistently applied, the camera can switch to using different cameras for video capture in a substantially seamless manner. The switch can be transparent to the user, and therefore the switching between cameras is inconspicuous in the captured video footage or, optionally, in the user interface for the user.

[0009] In general, the processes of performing video capture and associated image processing can be computationally expensive, especially in the case of high-resolution video capture. The techniques discussed herein, among other techniques, provide a computationally efficient technique for managing image stabilization and transitions between camera modules by mapping image data from both cameras to a single common reference space and then applying the same type of EIS processing to the image data within that reference space. The reference space may be time-delimited, which further reduces the computation required for aligning images from different cameras and applying EIS processing.

[0010] In some implementations, the techniques described herein are implemented on battery-powered devices with limited power budgets and computing resources, such as phones, tablet computers, and other mobile devices. The processing discussed can be efficiently performed on battery power by the device, and stabilization processing is performed virtually in real time, concurrently with ongoing video capture, for example, while the image-stabilized video output is saved or streamed as video capture continues. Furthermore, as described below, the techniques can adjust video capture using camera modules, for example, to adjust camera module settings for focus, exposure, etc., and to switch which camera module is used at different times. These techniques are also efficiently performed so that they can be performed in real time as video is captured, processed, and recorded while additional video continues to be captured.

[0011] In one general context, a method includes a video capture device having a first camera and a second camera providing a digital zoom function that enables user-specified magnification changes within a digital zoom range during video recording, wherein the video capture device is configured to (i) use video data captured by the first camera over a first portion of the digital zoom range, and (ii) use video data captured by the second camera over a second portion of the digital zoom range, and the method further includes processing image data captured using the second camera by applying a set of transformations, which include (i) a first transformation to a second canonical reference space for the second camera, (ii) a second transformation to a first canonical reference space for the first camera, and (iii) a third transformation for applying electronic image stabilization to image data in the first canonical reference space for the first camera, while capturing video using the second camera of the video capture device to provide a certain zoom level in the second portion of the digital zoom range.

[0012] In some implementations, the method involves (i) converting to a first canonical reference space for the first camera and (ii) applying electronic image stabilization to the data in the first canonical reference space for the first camera while capturing video using a first camera of a video capture device to provide a certain zoom level in a first portion of the digital zoom range. This involves processing image data captured using a second camera by applying a set of transformations, including a transformation for which the image data was obtained.

[0013] In some implementations, the first camera and the second camera have different fields of view, such that (i) the field of view of the second camera is included within the field of view of the first camera, or (ii) the field of view of the first camera is included within the field of view of the second camera.

[0014] In some implementations, the first camera and the second camera each include a fixed focal length lens assembly.

[0015] In some implementations, the canonical reference space for the second camera and the canonical reference space for the first camera are conceptual camera spaces defined by a predetermined fixed set of camera intrinsic properties such that projecting image data onto the canonical reference space eliminates time-dependent effects during video frame capture.

[0016] In some implementations, the first camera includes an optical image stabilization (OIS) system, and the first canonical reference space for the first camera is such that the image data is represented with a consistent predetermined OIS position.

[0017] In some implementations, the second camera includes an optical image stabilization (OIS) system, and the second canonical reference space for the second camera is such that the image data is represented with a consistent predetermined OIS position.

[0018] In some implementations, a first camera provides image data that progressively captures the image scan lines of an image frame, and a first canonical reference space for the first camera is corrected so that the image data removes distortion due to progressive capture of the image scan lines.

[0019] In some implementations, a second camera provides image data that progressively captures the image scan lines of an image frame, and a second canonical reference space for the second camera is corrected so that the image data removes distortion caused by progressive capture of the image scan lines.

[0020] In some implementations, the second transformation aligns the field of view of the second camera with the field of view of the first camera and adjusts for the spatial offset between the first and second cameras.

[0021] In some implementations, the first, second, and third transformations each have corresponding homography matrices, and processing image data involves applying these homography matrices.

[0022] In some implementations, the method includes receiving user input indicating a change in zoom level to a specific zoom level in a second portion of the digital zoom range while capturing video data using a first camera and processing the video data from the first camera for applying electronic image stabilization. The method includes recording a sequence of video frames in which the magnification of the video frames captured using the first camera is gradually increased until a predetermined zoom level is reached, initiating video capture using a second camera, and recording a second sequence of video frames captured using the second camera, wherein the second sequence of video frames provides a predetermined zoom level and provides an increasing magnification of the video frames captured using the second camera until a specific zoom level is reached. .

[0023] In some implementations, the second transformation is determined at least partially based on the focal length of the second camera.

[0024] In some implementations, the first transformation involves multiple different adjustments to different scan lines of the image data captured using the second camera.

[0025] In some implementations, the third transformation includes electronic image stabilization for each video frame, using one or more video frames prior to a particular video frame and one or more video frames after a particular video frame.

[0026] In some implementations, the second camera has a smaller field of view than the first camera. The method includes receiving user input indicating a change in the zoom level for video capture during image capture using the first camera, and determining whether the changed zoom level is greater than or equal to a predetermined transition zoom level in response to receiving the user input, where the predetermined transition zoom level represents a field of view smaller than that of the second camera.

[0027] In some implementations, the method includes storing data indicating (i) a first transition zoom level for transitioning from video capture using a first camera to video capture using a second camera, and (ii) a second transition zoom level for transitioning from video capture using a second camera to video capture using a first camera, wherein the first transition zoom level is different from the second transition zoom level, and the method further includes determining whether to switch between cameras for video capture by (i) comparing the requested zoom level to the first transition zoom level when the requested zoom level corresponds to a decrease in the field of view, and (ii) comparing the requested zoom level to the second transition zoom level when the requested zoom level corresponds to an increase in the field of view.

[0028] In some implementations, the first transition zoom level corresponds to a smaller field of view than the second transition zoom level.

[0029] In some implementations, the method includes deciding to switch from capturing video using a specific camera to capturing video using another camera during video file recording; determining the values ​​of video capture parameters used for image capture using the specific camera in response to the decision to switch; setting the values ​​of the video capture parameters for the other camera based on the determined video capture parameters; and, after setting the values ​​of the video capture parameters for the other camera, starting video capture from the second camera and recording the captured video from the second camera to a video file. Setting the video capture parameters includes adjusting one or more of the following for the second camera: exposure, image sensor sensitivity, gain, image capture time, aperture size, lens focal length, OIS status, or OIS level.

[0030] Other embodiments of this aspect and other embodiments discussed herein include corresponding systems, devices, and computer programs encoded on a computer storage device and configured to perform the operation of the method. A system of one or more computers or other devices causes the system to perform an action during operation. Installed software, firmware, hardware, or a combination thereof can be configured in this way. One or more computer programs can be configured in this way by having instructions that, when executed by a data processing device, cause the data processing device to perform an action.

[0031] In another general embodiment, one or more machine-readable media store instructions that, when executed by one or more processors, cause the execution of an operation, the operation comprising a video capture device having a first camera and a second camera providing a digital zoom function that enables user-specified magnification changes within a digital zoom range during video recording, the video capture device being configured to (i) use video data captured by the first camera over a first portion of the digital zoom range, and (ii) use video data captured by the second camera over a second portion of the digital zoom range, the operation further comprising processing image data captured using the second camera by applying a set of transformations, including (i) a first transformation to a second canonical reference space for the second camera, (ii) a second transformation to a first canonical reference space for the first camera, and (iii) a third transformation for applying electronic image stabilization to image data in the first canonical reference space for the first camera, while capturing video using the second camera of the video capture device to provide a certain zoom level in the second portion of the digital zoom range.

[0032] In another general context, the video capture device comprises a first camera having a first field of view, a second camera having a second field of view, one or more position or orientation sensors, one or more processors, and one or more data storage devices that store instructions that, when executed by one or more processors, cause the execution of an operation, the operation comprising the video capture device having the first and second cameras providing a digital zoom function that enables a user-specified magnification change within a digital zoom range during video recording, the video capture device (i) using video data captured by the first camera over a first portion of the digital zoom range, and (ii) digital The operation is configured to use video data captured by a second camera over a second portion of a digital zoom range, and the operation further includes processing image data captured using a second camera by applying a set of transformations, including (i) a first transformation to a second canonical reference space for the second camera, (ii) a second transformation to a first canonical reference space for the first camera, and (iii) a third transformation for applying electronic image stabilization to image data in the first canonical reference space for the first camera, while capturing video using the second camera of the video capture device to provide a certain zoom level in the second portion of the digital zoom range.

[0033] In some implementations, the first and second cameras have different fields of view, with the field of view of the second camera being within the field of view of the first camera. The first and second cameras may each include a fixed focal length lens assembly.

[0034] Details of one or more embodiments of the present invention are described in the accompanying drawings and the following description. Other features and advantages of the present invention will become apparent from the description, drawings, and claims. [Brief explanation of the drawing]

[0035] [Figure 1A] This figure shows an example of a device that provides multi-camera video stabilization. [Figure 1B] This figure shows an example of a device that provides multi-camera video stabilization. [Figure 2] Figures 1A and 1B are block diagrams showing examples of device components. [Figure 3] This figure shows exemplary techniques for video stabilization. [Figure 4] Figures 1A and 1B show examples of additional processing by the devices. [Figure 5A] This figure shows examples of techniques for stabilizing multi-camera video. [Figure 5B] This figure shows examples of techniques for stabilizing multi-camera video. [Figure 5C] This figure shows examples of techniques for stabilizing multi-camera video. [Figure 6] This figure shows an exemplary transformation that can be used to efficiently provide multi-camera video stabilization. [Modes for carrying out the invention]

[0036] Similar reference numbers and names in various drawings refer to the same elements. Detailed explanation Figures 1A and 1B show an example of a device 102 that provides multi-camera video stabilization. Video stabilization is often an important feature for camera systems in mobile devices. In conventional single-camera video stabilization, video frames can be converted from image data actually captured during a shaky real camera trajectory to a stabilized output for a time-smoothed virtual camera trajectory. In a single camera, frames acquired during the real camera trajectory can be projected onto a modified output frame representing frames along a smoothed virtual camera trajectory.

[0037] Some devices include multiple cameras facing the same direction (e.g., on the same side of the device) that are positioned to capture different fields of view of the same scene. In a multi-camera system, the device may switch cameras during the process of capturing video so that a single video includes segments captured using different cameras. When a switch between cameras occurs, a single real camera trajectory no longer exists because different real cameras are being used. Nevertheless, the device should maintain the same virtual camera trajectory to maintain continuity and smoothness of the captured image, even during the period of camera transitions. As discussed herein, the virtual camera trajectory can be maintained between multiple cameras and smoothed over time so that the switch between cameras does not cause distracting or interruption effects (e.g., stutter, image offset or abrupt changes in field of view, interruption of stabilization or abrupt changes in the level of stabilization applied).

[0038] Multi-camera video stabilization systems can be implemented to offer various benefits. In addition to smoothing video over time, they can also smooth transitions between periods of video capture from different cameras. Furthermore, this technology can be sufficiently efficient to operate in real time, for example, to apply smoothing to captured video simultaneously with continuous video capture, and to conserve power to enable long-term use by battery-operated devices.

[0039] In the example in Figure 1A, device 102 is shown as a telephone, but it may be another type of device, such as a tablet computer or a camera. Device 102 includes a multi-camera module 108 which includes a first camera 110a and a second camera 110b. The two cameras 110a and 110b are positioned on the same side of device 102 and are therefore both positioned to capture images of the same scene 105 facing device 102 (for example, on the back of the illustrated telephone). Cameras 110a and 110b may be tightly coupled together to device 102 so that cameras 110a and 110b move together to the same extent and with the same motion as device 102.

[0040] The cameras have different fields of view. For example, the first camera 110a has a first field of view 120, and the second camera 110b has a second field of view 122 which is narrower than the first field of view 120. The fields of view 120 and 122 may be significantly different, but the two cameras 110a and 110b are the same. They may have different image resolutions. The two fields of view 120 and 122 can overlap. In particular, the field of view 122 of the second camera 110b can be completely contained within the field of view 120 of the first camera 110a. For example, the first field of view 120 may be 77 degrees and the second field of view 122 may be 52 degrees, and the 52-degree field of view 122 is almost or completely contained within the 77-degree field of view 120. In this example, cameras 110a and 110b each use a fixed focal length lens, for example, a lens assembly without optical zoom. In other words, the lens focal length of each camera 110a and 110b may be fixed, apart from focus-related effects such as focus breathing.

[0041] One or more of the cameras 110a and 110b may include an optical image stabilization (OIS) module to reduce the effects of camera shake and other undesirable movements of the device 102. Whether or not the cameras 110a and 110b include an OIS module, the device 102 can use electronic image stabilization (EIS) to smooth the captured video over time. As will be further described below, the EIS process can apply transformations to the captured frames to project the captured frames, taken along a shaky real camera trajectory, into a stabilized output frame representing the smoothed or filtered trajectory of the virtual camera.

[0042] Device 102 uses cameras 110a and 110b to provide an effective zoom range that uses image capture at different magnification levels from different cameras 110a and 110b. Device 102 can provide a zoom range using digital zoom techniques applied to the outputs of different cameras 110a and 110b in different parts of the zoom range. For example, the device may provide an overall zoom range of 1.0x to 3.0x. Camera 110a, which provides a wider field of view 120, may be used to capture images in a first part of the zoom range, for example, 1.0x to 1.8x. When the zoom reaches a predetermined transition point, for example, a specific level such as 1.8x, device 102 switches from capturing video using the first camera 110a to capturing video using the second camera 110b. Camera 110b, which provides a narrower field of view 122, may be used to capture images in a second part of the zoom range, for example, 1.8x to 3.0x.

[0043] The illustrated example shows that device 102 includes a display that allows the user to view the captured video 130 while the video 130 is being captured and recorded by device 102. Device 102 may provide the user with a control 132 for dynamically adjusting the zoom level during video capture. In this example, the control 132 is an on-screen slider control shown on the touchscreen of device 102, allowing the user to set the zoom to a desired position along the zoom range. In some implementations, the zoom level is adjusted in fine-grained increments, for example, 0.2x, 0.1x, or smaller steps, allowing the user to gradually move across the supported zoom range in a manner that substantially approximates a continuous zoom across the supported zoom range.

[0044] Figure 1B shows an example of how different zoom levels can be provided using the outputs from cameras 110a and 110b. Each field of view 120 and 122 is shown along with cropped portions 140a–140e that can be used to provide different zoom levels. Output frames 150a–150e show exemplary frame outputs that can be provided at different zoom levels. Note that there are differences in the view of the scene because the two cameras 110a and 110b are physically offset from each other within device 102. Even if the two cameras are viewing the same scene, subjects in the scene will have slightly different positions in the camera outputs due to monocular parallax, which will be further explained below. Image conversion can correct this parallax and other differences between the outputs of cameras 110a and 110b.

[0045] Output frames 150a to 150c are each derived from image data from the first camera 110a, which provides a wider field of view 120. When the zoom level reaches a threshold such as 1.8x, device 102 switches to using image data captured by the second camera 110b, which provides a narrower field of view 122. The transition between cameras 110a and 110b can occur at a zoom level that completely covers the field of view 122 of the second camera 110b, resulting in sufficient image data to fill the output frames at the desired zoom level. In some implementations, device 102 is configured to perform the transition when the zoom level corresponds to an area smaller than the entire output of the second camera 110b, in order to maintain a margin of captured data that can be used for EIS processing.

[0046] Generally, to maximize output quality, it is advantageous to switch to a narrower field of view 122 near the zoom level at which the field of view 122 can fill the output frame. This is because the narrower field of view 122 provides a higher resolution for that area of ​​the scene. The two cameras 110a and 110b may have similar resolutions, but in the narrower field of view 122, camera 110b can use its full resolution to capture a view of the scene, whereas the same view will be captured with only a fraction of the resolution of the wider camera 110a. In this example, there is a zoom level, e.g., 1.7x, at which the zoomed-in field of view 140c of the first camera 110a matches the full field of view 122 of the second camera 110b. In this respect, only a fairly small portion of the image sensor for the first camera 110a is used to provide the output frame 150c, and therefore the output frame 150c may be of lower resolution or quality than what is provided at wider zoom levels. In contrast, the full resolution of the image sensor of the second camera 110b can be used to provide that level of effective zoom or magnification, resulting in higher quality output. By switching to the second camera 110b at or immediately after the equilibrium point (e.g., 1.7x), device 102 can provide higher quality at its zoom level and further zoom levels.

[0047] In some implementations, the zoom level set to switch between cameras 110a and 110b is set after the point at which the second camera 110b can fill the image frame. For example, the transition point may be set to a zoom level that provides a margin to account for monocular parallax resulting from the different physical positions of cameras 110a and 110b. As a result, even if the second camera 110b can fill the frame at a zoom level of 1.7x, the device 102 may nevertheless delay switching to capture using the second camera 110b to, for example, a zoom of 1.8x or 1.9x, so that all sensor outputs from camera 110b provide a buffer area for image data (e.g., at the edges surrounding the region representing the desired frame capture area), and so that any offset or other adjustments required to compensate for monocular parallax can be made, for example, so that the field of view 122 aligns with the enlarged area of ​​the field of view 120 and still fills the output frame.

[0048] Device 102 applies video stabilization, such as smoothing of the camera motion visible within a frame over time, to reduce or eliminate the effects of the movement of Device 102. Device 102 can perform video stabilization, such as EIS processing, in real time or near real time, for example, concurrently with an ongoing video capture of the stabilized video. The video stabilization smoothing over time (e.g., across multiple frames) is applied as the zoom setting changes between camera 110a and camera 110b. It can be coordinated with movement. Nevertheless, as discussed below, an image data transformation may be used to map the output of the second camera 110b to the canonical space for the first camera 110a, allowing a single EIS processing scheme to be used consistently for the outputs of both cameras 110a and 110b.

[0049] The techniques discussed herein can be effectively used with two or more fixed focal length lenses. Optionally, the techniques discussed herein can also be used with one or more lenses that provide optical zoom. For example, the techniques can be used to provide a seamless effective zoom over a range including (i) multiple cameras having different optical zoom ranges that may or may not intersect or overlap, or (ii) one or more cameras providing optical zoom and one or more fixed focal length cameras.

[0050] Figure 2 shows an example of a device 102 that provides video stabilization. As described above, device 102 includes a first camera 110a and a second camera 110b. One or more of the cameras 110a to 110b may optionally include OIS modules 215a to 215b. If device 102 includes OIS modules 215a to 215b, it may capture video frames while using the OIS modules 215a to 215b to at least partially cancel out the movement of device 102 during frame capture. Device 102 also includes one or more device position sensors 220, one or more data storage devices 230, and an EIS module 255.

[0051] Device 102 can be any of various types, including a mobile phone, a tablet computer, or a camera module such as a camera. In some implementations, device 102 may include a computing system for performing the operation of the EIS module 255, which may be run in software, hardware, or any combination thereof. For example, device 102 may include various processing components, such as one or more processors, one or more data storage devices for storing executable instructions, memory, input / output components, etc. The processor performing the EIS processing may include a general-purpose processor (e.g., the main CPU of a mobile phone or other device), a graphics processor, a coprocessor, an image processor, a fixed-function EIS processor, or any combination thereof.

[0052] The EIS module 255 uses position data from both the device position sensor 220 and the OIS modules 215a-215b to stabilize the video captured by the recording device. For example, position data from the OIS modules 215a-215b can be used to determine an offset representing the effect of OIS motion on the expected camera view that would be inferred from the device position data. This allows the EIS module 215 to estimate an effective camera position that reflects the actual view of the image sensor, even if the OIS modules 215a-215b change the camera's view of the scene relative to the device position. Along with other features described herein, these techniques enable the device 102 to effectively use OIS and EIS processing simultaneously and realize the benefits of both techniques.

[0053] Generally, OIS can be very effective in reducing blur within individual frames due to camera shake, and it can be somewhat effective in reducing motion that appears across a series of frames. However, OIS used alone is often subject to various limitations. OIS modules may be limited in the speed at which they respond to motion and the magnitude of motion they can compensate for. Furthermore, the operation of OIS modules often causes distortion such as shaky video and incorrectly counteracts desired motion such as panning. This can happen. The EIS module 255 can mitigate the effects of these limitations by using position data that describes the internal movement of the OIS module.

[0054] OIS modules 215a-215b attempt to compensate for the motion of the recording device, so device motion alone may not represent the true camera view used during video capture. If EIS processing attempts to compensate for motion based solely on device motion, it may attempt to correct motion already compensated for by the OIS system. Furthermore, OIS generally only partially removes the effects of device motion, and the amount of compensation may vary from frame to frame. To provide high-quality stabilization, EIS module 255 uses OIS position data along with device-level position data to vary the amount of stabilization applied to each frame, and in some implementations, also to individual scan lines within the frame. This processing provides effective stabilization and can reduce or eliminate distortion in the video image. For example, changes in OIS lens shift position during frame capture can introduce distortion, especially when combined with a rolling shutter, which is typical of many camera modules. Using information about the OIS lens shift at different times during frame capture, EIS module 255 can estimate the lens position when different parts of the frame were captured and correct the image. The EIS module 255 can also compensate to reduce the effects of OIS lens shift that interfere with panning or are otherwise undesirable.

[0055] Another way the EIS module 255 can enhance video is through the analysis of data about frames captured later. To process a particular frame, the EIS processing module may evaluate a set of camera positions within a time window that includes the time when one or more future frames were captured. Information about future frames and their corresponding positions can be used in several ways. Firstly, the EIS module 255 can apply filtering to the set of camera positions to smooth motion patterns used to define image transformations to modify frames. Secondly, the EIS module 255 can use the set of camera positions to evaluate the likelihood of consistent motion (e.g., panning) existing or being attempted, and then, if that is likely, adjust the frame to be consistent with this motion. Thirdly, the EIS module 255 can evaluate the camera position of a frame against future camera positions and make adjustments for large future movements. For example, if a large, rapid movement is identified for a future frame, the EIS module 255 can begin adjusting the contents of the frame before that movement begins. The EIS module 255 can extend motion over a larger number of frames, rather than allowing large, visible motions to occur over several frames. Gradual image shifts occur in earlier frames, gradually spreading the motion over a greater number of frames.

[0056] The EIS module 255 performs regional synthesis of output frames, for example, by modifying the transformation applied to each scan line of the image frame. This allows the system to compensate for rolling shutter distortion, motion of OIS modules 215a-215b, and various device movements that occur within the capture duration of a single frame.

[0057] Referring further to Figure 2, device 102 could be any suitable device having a camera to capture video data, such as a camera, cellular phone, smartphone, tablet computer, wearable computer, or other device. While the example in Figure 2 shows a single device capturing and processing video, the functionality may, optionally, be spread across multiple devices or systems. For example, the first device may capture video frames and also record location data and other parameters as metadata. The first device and The metadata may be provided to a second device capable of performing the EIS processing described herein, such as a local computing system or a remote server.

[0058] The first camera 110a may include a lens element, an image sensor, a sensor reading circuit, and other components. The OIS modules 215a-215b may include a sensor, a movable element, a processor, and a drive mechanism for moving the movable element. The movable element is located in the optical path of the first camera 110a. For example, the movable element may be a reflective or refractive element, such as a lens, mirror, or prism. In some implementations, the movable element is the image sensor of the first camera 110a. The sensor may include one or more gyroscopes or other sensors for detecting motion. The processor determines the amount and direction of movement required of the movable element to compensate for the motion indicated by the sensor, and then commands the drive mechanism to move the movable element.

[0059] Device 102 includes one or more position sensors 220 that measure changes in the orientation of device 102. In some implementations, the position sensors 220 for device 102 are separate from the sensors used by the OIS modules 215a-215b. The position sensors 220 can detect rotation of device 102 around one or more axes. For example, the device position sensor 220 may be a 3-axis gyroscope or an inertial measuring unit (IMU). In addition, or as an alternative, other sensors may be used to determine the device position. For example, one or more accelerometers, a 1-axis gyroscope, a 2-axis gyroscope, etc., may be used to determine the position of device 102. In general, any suitable sensor or combination of sensors that enables the rotational position of device 102 to be determined can be used.

[0060] In some examples, position data from the gyroscope sensors of the OIS modules 215a-215b may be captured and stored in addition to, or instead of, using a separate position sensor 220 of the recording device 220. Nevertheless, it may be beneficial for device 102 to use a gyroscope sensor with different characteristics from the OIS sensor. For example, the gyroscope sensor of device 102 may provide measurements at a rate of about 400 Hz with a detectable rotation range greater than 100 degrees / second. Compared to a device-level sensor, a typical gyroscope sensor of an OIS module may provide measurements at different rates and ranges, e.g., a rate of 5000 measurements / second or more with a detectable rotation range of about 10 degrees / second. In some implementations, having a larger detectable rotation range of the device-level sensor (for example, to describe large movements) is beneficial, as is the more frequent measurements of the OIS module sensor (for example, to detect small changes or high-frequency patterns). Therefore, both types of data may be used together to determine the position of device 102.

[0061] Device 102 includes one or more data storage devices 230 that store information characterizing the first camera 110a and the frame capture process. For example, the stored data may include calibration data 232 showing the relationship between the positions of the OIS modules 215a-215b and the resulting offsets occurring in the image data. Similarly, the calibration data 232 may show the correspondence between the camera module lens focal positions and the effective focal distances for those focal positions (e.g., with different mappings for each camera 110a, 110b), allowing the system to account for focus breathing. In addition, the calibration data 232 or other stored data may show the correspondence between the camera lens focal positions and the subject distance, allowing the conversion from the lens focal position selected by the autofocus system to the subject distance, which indicates the distance of the in-focus subject from the camera's sensor plane. The calibration data 232 also shows the relative position of one camera 110a to the other camera 110b. It also indicates the approximate 3D spatial position. Typically, this includes calibration data that specifies the 3D rotation and 3D translation of one camera relative to the other camera. Typically, calibration is performed for each device manufactured to ensure the best user experience. As a result, the calibration data 232 can be highly accurate to the characteristics of a particular camera module (e.g., cameras 110a, 110b, their mounting structures, etc.) and the state of the module after manufacturing. The stored data may include scan pattern data 234 that can indicate the reading characteristics of the image sensor in the first camera 110a. For example, the scan pattern data 234 may indicate the direction of scanning (e.g., scan lines are read from top to bottom), whether scan lines are read individually or in groups, etc.

[0062] During video capture, cameras 110a-110b, OIS modules 215a-215b, and device position sensor 220 may each provide information about the video capture process. The first camera 110a provides video frame data 242a, for example, a sequence of video image frames. The second camera 110b similarly provides video frame data 242b. Cameras 110a-110b also provide frame exposure data 244a-244b, which may include, for each captured frame, an indication of the exposure duration and a reference time (e.g., the start or end time of the exposure) indicating when the exposure occurred. Cameras 110a-110b also provide lens focus position data 246a-246b indicating the lens focus position for each captured frame.

[0063] The OIS modules 215a to 215b provide OIS position data 248a to 248b indicating the position of the movable elements of the OIS modules 215a to 215b at various times during video capture. For example, if the movable element is a movable lens that shifts to compensate for motion, the OIS modules 215a to 215b can provide lens shift readings that identify the current position of the movable lens at each point. Device 102 can record the lens shift position and the time at which that position occurred. In some implementations, the OIS position data 248a to 248b are captured at a high frequency, for example, a rate higher than the video capture frame rate, so that multiple measurements are taken over the duration of each video frame exposure.

[0064] The device position sensor 220 provides device position data 250 indicating the rotation and / or other movement of the device 102 during video capture. The device position can be measured at high frequencies, e.g., 200 Hz or higher. Therefore, in many cases, measurements can be obtained for multiple different times during the capture of each video frame.

[0065] Lens focus position data 246, OIS position data 248, and device position data 250 can all be recorded with a timestamp indicating the time when the identified position occurred. The timestamp can be created with high precision, for example, down to the nearest millisecond, allowing for temporal alignment of data obtained from various position measurements. Furthermore, the positions of the device, OIS system, or lens focusing mechanism can be interpolated to determine values ​​over time between measurements.

[0066] An example of potential timing for data capture is shown in Chart 252. As illustrated, device position data 250 (e.g., gyroscope data) and OIS position data 248 (e.g., lens shift position data) may be captured at a rate higher than the video capture frame rate (e.g., 30 frames / second, 60 frames / second, etc.), and multiple positions of the device and OIS system can be determined for each video frame. As a result, for each scan line, which may be horizontal or vertical depending on the shutter direction, different device positions and OIS settings can be used to determine that scan line The conversion can be determined. Lens focus position data 246 may be captured at least once per image frame. This position data may be captured asynchronously with respect to the frame exposure, for example, by sampling gyroscope sensor data and OIS position data at a rate that exceeds the start or end of the image frame exposure and is not necessarily synchronized with it.

[0067] Data obtained from cameras 110a-110b and other components is provided to the EIS module 255 for processing. This processing may be performed while video capture is in progress. For example, EIS processing can be performed in virtually real time, so that the video file, which is made accessible to the user at the end of video capture, is stabilized by the EIS module 255. In some implementations, EIS processing may be performed later, for example, after video capture is complete, or by a device different from the one that recorded the video. The EIS module 255 can be implemented in hardware, firmware, software, or a combination or partial combination thereof.

[0068] Figure 2 shows only a portion of the functionality of the EIS module 255. The example in Figure 2 shows EIS processing on a video frame containing a frame captured from a single camera among cameras 110a-110b and does not describe the features used to provide digital zoom or to adjust the EIS processing with respect to the zoom level. However, as described below, device 102 can provide digital zoom during video capture and recording using image data captured from both cameras. For example, when the user zooms in to a threshold zoom amount, device 102 can switch from using the image captured from the first camera 110a to using the image captured from the second camera 110b. Furthermore, since zooming in can accentuate visible shake in the video, the zoom level can be adjusted to apply a greater level of stabilization as the zoom level increases, for example, by adjusting the transformations and other operations of the EIS module 225. Adjusting the EIS processing to consider digital zoom and transitions between multiple cameras is discussed with respect to Figures 4 to 6. In some implementations, the transformations can be achieved by applying different transformations to each line of the captured image. For each line in the image, device 102 can calculate a specific timestamp for that line, and this line is associated with the corresponding OIS and device location data.

[0069] Referring further to Figure 2, the EIS module 255 includes a device position data handler 256 that periodically or continuously acquires updated device position data 250 from the device position sensor 220. The motion data handler estimates the current camera orientation from the device position data 250. For example, a gyroscope signal can be acquired and used to estimate the device position of device 102 at a high frequency, such as 200 Hz. This device position at a given time t is hereafter referred to as R(t). This device position may represent, for example, the rotational position of device 102 with respect to one, two, or three axes. The device position may be expressed as a rotation matrix, or with respect to a coordinate system, or in other forms. Each calculated device position can be labeled with a time indicating the time at which device 102 obtained that position.

[0070] The EIS module 255 includes an OIS position data handler 258, which periodically or continuously acquires OIS position readings, indicated as OIS position data 248. The OIS position data handler 258 converts the OIS readings into offsets that can be used with the device position. For example, the OIS lens position can be converted into a two-dimensional pixel offset. To generate the offset, the OIS position data handler 258 may provide a conversion coefficient or matrix to convert the OIS position to the corresponding offset, stored in memory. The calibration data 232 can be used. Generating an offset due to the OIS position can take into account the change in the effective focal length of the camera over time due to changes in the lens focal position and / or lens zoom position, for example, if the first camera 110a is capable of optical zoom. Similar to the motion data handler 256, the OIS position data handler 258 labels each measurement and offset with the time the data represents.

[0071] The EIS module includes a motion model builder 260 that receives the device position calculated by the device position data handler 256 and the OIS offset calculated by the OIS position data handler 258. Using this data, along with the frame exposure data 244 and the lens focus position data 246, the motion model builder 260 generates a first transformation 262 for the frame. For example, the first transformation 262 may be a projection matrix that maps the real-world scene in the camera's view to the captured frame. This process is repeated for each frame. When generating the first transformation 262 for a frame, the positions of the OIS modules 215a-215b can be modeled as offsets from the principal device position determined from the gyroscope data. As will be discussed further below, the offset can take into account the camera's effective focal length at the time of capture by examining the effective focal length relative to the lens focus position at that time. The first transformation 262 can separately describe the relationships of different subsets or regions of a single image frame. For example, different parts or components of the first transformation 262 may describe how different scan lines of a frame are mapped to a real-world scene. Device position, OIS module position, subject distance (e.g., the distance of the in-focus subject from the camera), and lens focal position can all be aligned using measured timestamps and interpolated as needed to provide precise positions at the exposure of individual scan lines of the frame. For lenses with autofocus, the focal position is set depending on how far the subject is from the camera. Thus, a map exists showing the relationship between the lens focal position and the subject distance. The mapping can be generated and calibrated, and the subject distance is used in subsequent calculations for spatial transitions.

[0072] The first transformation 262 generated by the motion model construction unit 260 is provided to the nonlinear motion filtering engine 270, which determines the second transformation 272. This second transformation 272 can be a second projection matrix P'i,j that projects the image data of a frame onto an output frame representing a stabilized version of that frame. Specifically, the second transformation 272 can map the image projection Pi,j created using the first transformation 262 to the output frame, rather than performing calculations on the captured image data. In some implementations, the two transformations 262,272 can then be combined into a single transformation that performs calculations on the image data initially captured in the frame and maps it directly to the stabilized output frame.

[0073] To effectively stabilize motion, the nonlinear motion filtering engine 270 can generate a second transformation 272 to take into account motion that will occur in the future after the capture of the frame being processed. For example, for the current frame being analyzed, the position of the recording device may not have moved significantly from the previous frame. Nevertheless, if the engine 270 determines that significant motion will occur in a future frame, the second transformation 272 may be generated to shift or otherwise modify the current frame to introduce visible motion into the video, so that large future motion can be diffused as a series of gradual changes rather than abrupt changes. Similarly, if stabilization of a future frame results in cropping or other changes, the second transformation 272 may be generated to propagate those changes at least partially to previous frames for a more gradual and consistent change across a series of frames.

[0074] The nonlinear filtering engine 270 can generate a second transformation 272 from the virtual camera position relative to the frame. Rather than representing the actual position of the camera when exposure occurs, the virtual camera position can represent the adjusted or virtual orientation of the device 102 that will stabilize the recorded video. The virtual position can represent a desired position for positioning the virtual camera, for example, a position that will simulate a particular view or perspective of the scene. In general, any camera position can be represented by its rotation and translation relative to a global reference frame. The virtual camera position can be represented as a rotation matrix, for example, a matrix showing rotation offsets relative to the reference position. This may be a 3x3 matrix showing rotation offsets for three rotation axes. In some implementations, the stabilization process of the EIS module defines the position only with respect to rotation components, because these generally have the greatest impact on the stability of handheld video.

[0075] The virtual camera position relative to a frame may reflect adjustments to the estimated camera position to enhance video stabilization, correct distortion and motion, facilitate panning, and otherwise enhance the video. The virtual camera position relative to a frame may be determined by generating an initial camera position that is adjusted based on various factors. For example, adjustments to the virtual camera position may be made through filtering of device positions based on motion detected before and after the frame, based on the amount of blur in the frame, based on the likelihood of panning occurring, through adjustments to prepare for future motion in the frame, and / or to ensure that the image data covers the entire output frame. Various factors may be considered by generating a set of virtual camera positions for that frame, which are modified, mixed, or otherwise used to determine the final virtual camera position relative to the frame.

[0076] Just as transformations 262 and 272 can have different mappings for different scan lines, different virtual camera positions can be determined for different scan lines of a frame to adjust for changes in the device position, the positions of OIS modules 215a-215b, and / or lens focal position during frame capture. Therefore, different virtual camera positions can be used for different parts of a frame. For efficiency, the virtual camera positions and corresponding components of the second transformation 272 may be calculated for a suitable subset of the image sensor's scan lines, and then the appropriate data may be interpolated for the remaining scan lines. In the various examples below, for brevity, we will discuss a single scan line, such as the central scan line of the image sensor. Techniques for fully calculating the virtual camera positions and corresponding projection matrix components may be used for multiple scan lines of an image frame, and even individually for each scan line as needed.

[0077] As used herein, device position refers to the position of device 102, as indicated, for example, by device position data 250 (e.g., gyroscope data) and the output of device position data handler 256. This device-level position indicates the attitude or orientation of device 102, without considering the internal movement of the lens of the first camera 110a or the movement of OIS modules 215a-215b. Also, as used herein, camera position refers to the position corresponding to the effective or estimated view of the camera. By considering the operation of OIS modules 215a-215b, lens breathing, and other factors causing shifts, the camera position may differ from the device position. Furthermore, the camera position may be a virtual position, for example, an approximate or virtual position that reflects an enhanced or modified view of the camera rather than the actual view of the camera.

[0078] Next, the EIS module 255 uses the image warping engine 280 to perform nonlinear The output of the motion filtering engine 270 is used to map each captured image frame to an output frame. The second projection 272 may include components corresponding to each scan line of the frame, such that each part of the frame is mapped to the output space and each pixel of the output frame is defined. Processing of the EIS module 255 may be performed for each frame of the video.

[0079] Figure 3 shows an example of data used for video stabilization. This figure shows a series of frames 310 captured by device 102. Each frame is labeled with a corresponding set of metadata 312, which indicates, for example, the exposure duration, the exposure reference time (e.g., the start time, stop time, or other reference point of the exposure), the lens focus position, etc. Although not shown, device position data and OIS module position data are captured and timestamped at various times during each exposure.

[0080] To perform stabilization processing on frame 311, a time range near the capture of frame 311 is defined. This time range, or frame window, is analyzed to determine how frame 311 should be transformed. As used herein, the frame time "t" generally refers to the time of the central scanline capture, which is used to represent the reference time of the frame capture. When referring to the time of individual scanlines (for example, to consider times that may differ from the principal time of the capture relative to the central scanline of the frame), that time is t LThis is shown as L being the index or identifier of a particular scan line. The exposure time t of the central scan line of frame 311 being analyzed can be used as the center of the time range used for analysis. A predetermined time offset Q can be used to set the time range from which that range is derived, for example, [tQ, t+Q]. In some implementations, this time offset Q is approximately 200 ms. As a result, this range will include approximately 7 frames before and approximately 7 frames after frame 311. Larger and smaller time offset Qs may also be used. Since the EIS module 255 uses the context of future frames during processing, processing of a frame is delayed until a suitable number of subsequent frames have been captured.

[0081] In Figure 3, frame 311 is shown to be captured by an image sensor (for example, of one of the cameras 110a-110b). As described above, the EIS module 255 defines a first transformation 262 from data indicating the actual position of device 102 during the capture of frame 311, as well as the positions of camera elements such as OIS module elements and lens focal positions. The result of applying the first transformation 262 is the projected image 330 shown with respect to the output frame target 335. In some implementations, the first transformation 262 is determined using only the data corresponding to the capture of frame 311. The transformation 262 corresponds to the actual lens position of the camera, and therefore the projected image 330 estimates the mapping between the image data and the actual scene in the camera's view.

[0082] The EIS module 255 further adjusts the image data of frame 311 using a second transformation 272. This second transformation 272 corresponds to a virtual lens position, for example, a virtual position that would result in a more stable video when used to capture frame 311. When this second transformation 272 is applied to frame 311, it produces a projected image 340 that fully defines the data for output frame 335.

[0083] A second transformation 272 that generates the projected image 340 may be generated from data corresponding to each frame within the time range from [tQ, t+Q]. The position R(t) of the device 102 over this period can be filtered, for example, using a Gaussian filter to smooth the motion over that range. The set of positions R(t) here is [ This is a set containing the position of device 102 at each central scanline capture time occurring within the range tQ, t+Q]. Consider an example where this range includes the central scanline capture time t0 of the currently processed frame and the central scanline capture times of the seven frames before and after the currently processed frame. The set of positions to be filtered is the set {R(t -7 ),R(t -6 ), ..., R(t -1 ),R(t0),R(t1),...,R(t6),R(t -7 )} will result. The result of filtering at time t, i.e., the exposure of the central scan line of frame 311, can be used as the initial virtual camera position V0(t). Even with filtering, there may be undesirable movements in the device position, or other factors that cause undesirable movements. As a result, the initial virtual camera position V0(t) may be updated through a series of further operations. In some implementations, the position R(t) to be filtered is a position that does not assume OIS movement and can therefore be based on device position data 250 that does not take OIS position data 248 into account. In other implementations, OIS movement and offset may be included as elements in the set of positions that are filtered to generate the initial virtual camera position V0(t).

[0084] For example, the second virtual camera position V1(t) is determined by the amount of movement that occurs over future frames, based on the position V0(t) and the final camera position VF(t) relative to the previous frame. pre It can be generated by interpolation with ). Final camera position VF(t pre) can be the virtual camera position relative to the central scanline of the frame captured immediately before frame 311 (for example, the position used to generate the recorded output frame). The camera position of the previous frame can be the position corresponding to the final virtual camera position, for example, the position used to generate the stabilized output version of the previous frame. Interpolation can align the visible change in motion between frame 311 and the previous frame with the visible change in motion between frame 311 and future frames.

[0085] A third virtual camera position V2(t) can be generated by interpolating V1(t) with the actual device position R(t) based on the amount of camera motion blur present in frame 311. This can reduce the amount of stabilization applied to reduce the viewer's perception of blur. Since motion blur is generally impossible to eliminate, this can reduce video stability at appropriate times to produce a more natural result.

[0086] A fourth virtual camera position V3(t) can be generated to simulate or represent a position that occurs during consistent motion of device 102 over a time range [tQ, t+Q]. This position may be determined by applying a stable filter, such as a domain transformation filter, to the estimated actual device position R(t) over a time range. The filter is applied to the same set of device positions used to generate V0(t), but this step represents a different type of filtering. For example, V0(t) may be generated through filtering that smooths but generally follows changes in the estimated actual device position over time without imposing a predetermined shape or pattern. In contrast, V3(t) is generated by filtering the device pattern to fit a predetermined consistent motion pattern, such as substantially linear panning or other motion, which may be potentially intended by the user of device 102.

[0087] A fifth virtual camera position V4(t) can be generated as an interpolation of V3(t) and V2(t). The EIS module 255 can evaluate whether the change in device position over time is likely to represent panning of device 102 and weight or adjust the interpolation accordingly. If panning is deemed likely, V4(t) will be close to the estimated panning position V3(t). If panning is deemed unlikely, V4(t) will be closer to position V2(t).

[0088] At the fifth virtual camera position V4(t), the EIS module 255 can evaluate the coverage that the corresponding transformation will provide to the output frame 335. Since it is desirable to fill the entire output frame 335 and not leave any pixels undefined, the EIS module 255 may determine a transformation, such as a projection matrix representing the view of the scene from the virtual camera position V4(t), and verify that the projected image will cover the output frame 335. To account for motion in future frames, the transformation may be applied to the portion of the scene that will be captured by future image frames. The transformation and the corresponding virtual camera position V4(t) may be adjusted so that each pair of current and future frames, when mapped using the transformation, will all completely define the output frame 335. The resulting transformation may be set as transformation 272 and may be used to generate a stabilized output frame 335 for frame 311.

[0089] In some implementations, generating a stabilized output frame 335 for frame 311 involves time t for one or more other scan lines of frame 311. LThis includes performing the EIS processing technique described for the exposed scan lines L. For example, the processing may be performed on the scan lines at certain intervals (e.g., every 100 scan lines, every 500 scan lines, etc.) or at certain reference points (e.g., 1 / 4 and 3 / 4 across the frame, or the top and bottom of the frame). When the virtual camera position and the second transformation 272 are determined for only a suitable subset of the scan lines in frame 311, the transformations on the scan lines (e.g., corresponding parts of the projection matrix) are interpolated between the calculated positions. In this way, a suitable transformation is determined for each scan line, and each scan line may have different transformations applied as a result. In some implementations, the entire process of generating the virtual camera position and the second transformation 272 may be performed for each scan line in each frame without relying on interpolation between data of different scan lines.

[0090] When frame 311 is mapped to output frame 335, the result is saved, and the EIS module 255 begins processing the next frame. The process continues until each frame of the video has been processed.

[0091] The various elements used to generate the virtual camera positions and the resulting transformations can be used in combination or separately. For example, depending on the implementation, some of the interpolation and adjustments used to create the virtual camera positions V0(t) to V4(t) may be omitted. For example, in a different implementation, one of the filtered camera positions V0(t) to V3(t) may be used to determine the transformation for projecting the data onto the output frame, instead of using V4(t) for that purpose. Thus, using one of the filtered camera positions V0(t), V1(t), and V2(t) to generate a stabilization transformation can still improve the stability of the video. Similarly, V3(t) may be effective in stabilizing video where panning is occurring. Many other modifications are within the scope of this disclosure, even if they take into account a subset of the different elements discussed.

[0092] The techniques discussed can be applied in various ways. For example, instead of sequentially applying the two transformations 262 and 272 to the image data, the recording device may generate a single combined transformation that reflects the combined effects of both. Thus, generating stabilized image data using transformations 262 and 272 may involve generating further transformations or relationships that are ultimately used to stabilize the image data, rather than directly applying transformations 262 and 272. Various techniques for image stabilization are described, and other techniques, such as the technique discussed in U.S. Patent No. 10,462,370 issued on October 29, 2019, which is incorporated herein by reference, may be used in addition or as an alternative.

[0093] Figure 4 is a block diagram showing an example of additional processing by device 102 as shown in Figures 1A and 1B. In addition to the elements shown in Figure 2, some of which are again shown in Figure 4, device 102 may include hardware and / or software elements to provide additional functionality as shown in Figure 4.

[0094] Device 102 may include a zoom input processing module 410 that processes user input to Device 102 indicating a requested change in the zoom level. When Device 102 captures video, even if EIS is in use, Device 102 can receive user input to change the zoom level. This can include input to move an on-screen slider control, a gesture on a touchscreen, or other input. A zoom level of 1.0x can represent the widest field of view available using camera 110a with EIS in use. This may be a cropped section of the native image sensor resolution to provide a margin for EIS processing. Module 410 can determine the desired zoom level based on the user input and, for example, change from the current zoom level (e.g., 1.0x) or move to the desired zoom level (e.g., 2.0x).

[0095] Data indicating the requested zoom level is provided to a camera selector module 420, which determines whether to switch the camera used for video recording. For example, the camera selector module 420 can receive and use a stored camera transition threshold 430 indicating the zoom level at which a transition between cameras should occur. For example, the second camera 110b may have a field of view corresponding to a zoom level of 1.7x, and the transition threshold may be set to a zoom level of 1.8x. The camera selector 420 determines that the desired zoom level 2.0 satisfies the 1.8x threshold (e.g., is greater than or equal to the threshold), and therefore, the camera change is appropriate from a capture representing a zoom level of 1.8x or higher. Multiple thresholds may be defined, such as a first threshold for zooming in (e.g., narrowing the field of view) and a second threshold for zooming out (e.g., widening the field of view). These thresholds may differ. For example, the first threshold may be 1.8 and the second threshold may be 1.7, and therefore the transition point will differ depending on whether the user is zooming in or zooming out.

[0096] In some implementations, switching between cameras 110a and 110b is controlled simply by exceeding a threshold. Consider a scenario where the device is recording video using camera 110b, and the user begins zooming out to a level that triggers a switch to use the wider-angle camera 110a instead. The threshold for switching from camera 110a to camera 110b when zooming in is 1.8x, and the lowest zoom position available to camera 110b is 1.7x (e.g., the zoom level representing the full maximum field of view of the current camera 110b). Consequently, when zooming out, the transition must occur at zoom level 1.7x because the second camera 110b cannot provide a wider field of view. To address this situation, when the zoom level is reduced to 1.9x, camera 110a begins streaming image data, even though the output frame is still based on the current image data output by camera 110b. This provides a period during which both cameras 110a and 110b are capturing and streaming image data of the scene, for example, during a zoom level transition from 1.9x to 1.8x. This provides a stabilization period for initializing the capture on camera 110a so that camera 110a can "warm up" and autofocus, auto exposure, and other processes will begin to converge. This stabilization period may also provide a margin for camera 110a's settings so that camera 110a's settings and capture are initialized and aligned to match what is currently being used by camera 110a before the switch.

[0097] In this way, device 102 can predict and prepare for the need to switch between cameras based on factors such as the current zoom level, the direction of zoom change, and user interaction with device 102. For example, if device 102 detects that a camera switch is likely to be needed based on user input commanding to reduce the zoom to 1.9x, it can command the settings to begin capturing with the next camera 110a to be used, before the user commands the target zoom level for the transition, or when it becomes necessary. For example, if device 102 starts video capture and setting adjustments when the user commands a zoom level of 1.9x, and camera 110a is ready and operating with the desired settings by the time the user commands a zoom level of 1.8x, device 102 can switch cameras when the zoom level of 1.8x is commanded. Nevertheless, if for some reason the settings have not converged or been implemented by the time the zoom level of 1.7x (e.g., the maximum field of view of the current camera 110b) is commanded, device 102 will force a switch when the zoom level reaches 1.7x.

[0098] The camera selector 420 provides camera selection signals or other control data to the camera control module 440. The camera control module 440 reads video capture parameters from cameras 110a and 110b and also sends commands or settings to set the video capture parameters. Video capture parameters may include, for example, which camera 110a or 110b is capturing image data, the frame capture rate (e.g., 24 frames per second (fps), 30 fps, 60 fps, etc.), exposure settings, image sensor sensitivity, gain (e.g., applied before or after image capture), image capture time (e.g., the effective "shutter speed" or duration for each scan line to capture light during the frame), lens aperture size, subject distance (e.g., the distance of the subject in focus from the camera), lens focus position or focal length, OIS status (e.g., whether OIS is enabled, the mode of OIS used, etc.), OIS lens position (e.g., horizontal and vertical offset, rotation position, etc.), and the intensity or level of OIS applied. The camera control module 440 can set these and other video capture parameters for general video capture, such as exposure, frame rate, etc. The camera control module 232 can also receive or access calibration data 232. Calibration procedures can be performed for each device, for example, for device 102, as part of the manufacturing and quality assurance of the device. Calibration data can provide data for fine-tuning, for example, the relationship between cameras 110a and 110b to each other, the relationship between the lens focal position and the subject distance (e.g., the distance of the focal plane) for different lens positions, etc.

[0099] The camera control module 440 can also enable and disable cameras 110a and 110b at appropriate times to cause camera switching indicated by the camera selector 420. For example, if the input indicates a change in zoom level from 1.0x to 2.0x, and data from the camera selector 420 indicates a change to a second camera 110b starting at 1.8x, the camera control module 440 can generate a control command to cause this transition. Using stored information regarding the requested zoom speed and any constraints on the speed at which the zoom can be performed consistently, the camera control module 440 determines the time to make the switch, for example, the time it takes for a gradual or incremental zoom reflected in the image output to reach the 1.8x camera transition point. The time for this transition can be based on the duration or number of frames for which data continues to be captured with the first camera 110a until the digital zoom can smoothly reach the 1.8x zoom level, or on the time specified by user input to reach the 1.8x zoom level.

[0100] Anticipating a camera transition, the camera control module 440 can read video capture parameters from the current camera (e.g., the first camera 110a) and set corresponding video capture parameters for the camera to be used after the transition (e.g., the second camera 110b). This may include setting the same frame rate, same exposure level, same lens aperture, same focal length, same OIS mode or status (e.g., whether enabled), etc. Generally, there is a period before the transition between cameras in which both cameras 110a and 110b are actively capturing video data simultaneously. For example, when switching from the first camera 110a to the second camera 110b, camera 110b will open and begin video capture before the zoom level reaches the threshold that would cause it to switch to using the output of camera 110b for the recorded video. The initial values ​​for the settings are based on the values ​​of the settings currently used for camera 110a. Camera 110b begins adjusting its operation toward the instructed settings so that it converges to a desired operating mode (e.g., appropriate aperture setting, correct focal length setting, correct OIS setting, etc.). This process may include camera 110b or device 102 determining the appropriate settings, such as using an autofocus process to determine the correct focal length. After the convergence and calculation of settings for camera 110b, the video stream used for recording will switch to the video stream output by camera 110b. In some cases, the parameters of cameras 110a and 110b do not have to be the same, but the parameters of camera 110b may nevertheless be set based on the capture parameters used by camera 110a, and may be set to promote or maintain consistency between outputs. For example, cameras 110a and 110b do not necessarily have the same aperture range available, and therefore the camera control module 440 may set equivalent or nearly equivalent exposure levels for the two cameras 110a and 110b, but with different combinations of settings for sensitivity / gain, capture time (e.g., shutter speed), and aperture.The camera control module 440 can start the camera to be transitioned with appropriate settings applied prior to the transition so that the camera is used for the final recorded output frame, and thus image capture and incoming video feed are available at or before the transition.

[0101] The video data is processed using the image processing module 450. This module 450 can receive captured video frames streamed from the cameras 110a and 110b currently selected for video capture. Module 450 also receives sensor data from device position sensors such as gyroscopes, inertial measurement units (IMUs), and accelerometers. Module 450 also receives video capture parameter values ​​(e.g., from the camera control module or memory) indicating the parameters used to capture the video frames. This may include metadata indicating OIS element position, camera focal position, subject distance, frame capture time, shutter speed / capture duration, etc., for each frame, for different parts of a frame, and even for a specific scan line or point in time during the process of capturing or reading the frame (see Figure 2, Chart 252). Module 450 also receives data indicating the digital zoom level (e.g., may be expressed as magnification level, equivalent lens focal length, resulting field of view, cropping level, etc.). The image processing module 450 applies transformations to the captured image frame to remove artifacts such as rolling shutter, OIS system motion, and focus breathing. For example, module 450 can acquire a video frame and transform or project it into the canonical space of the camera that captured the frame, in which time-varying aspects of the frame capture (e.g., motion of OIS elements, rolling shutter, etc.) are removed.

[0102] Module 450 is configured such that the image frame is captured by either the first camera 110a or the second camera 110 Regardless of whether the image frame was captured with camera b, it can be converted to the canonical camera space of the first camera 110a. For example, in the case of an image captured using the second camera 110b, module 450 converts the data from the canonical space of the second camera 110b to the canonical space of the first camera 110a by correcting for spatial differences between the positions of cameras 110a and 110b on the device and other factors such as the focal position of camera 110b. This allows the image data captured using the second camera 110b to be aligned with the field of view of the first camera so that the view of the scene is consistent across the video recorded from both cameras 110a and 110b. This technique will be discussed further below with respect to Figures 5A-5C and Figure 6.

[0103] Device 102 may include an EIS processing module 460 that receives and processes image data converted to the canonical space of the primary camera. The "primary camera" is a camera among several cameras that is pre-designated as the reference for the other cameras. For example, the primary camera could be a first camera 110a with the widest field of view, and the output of any other camera (e.g., a second camera 110b) can be converted or mapped to the canonical space of the first camera 110a. Module 450 maps the image data from both cameras 110a and 110b to a common, standardized canonical space that compensates for time-dependent variations within the frame (e.g., differences in capture times, device positions, OIS positions, etc., between different scan lines in the frame). This significantly simplifies EIS processing by eliminating the need for the EIS processing module 460 to consider time-varying capture characteristics within the frame. It also enables a single EIS processing workflow to be used for video data captured using either camera 110a or 110b. The EIS processing module 460 also receives the desired zoom settings, such as zoom level or field of view, potentially frame by frame. This allows the EIS processing module to apply an appropriate amount of stabilization frame by frame, according to the zoom level used for the frame. As the zoom level increases and the image is magnified, the effects of camera motion are also amplified. Therefore, the EIS processing module 460 can apply stronger stabilization as the zoom level increases to maintain a nearly consistent level of stability in the output video. The EIS processing module can use any or all of the EIS processing techniques described above with respect to Figures 2 and 3. Stabilization can be thought of as projecting the image data from the canonical image space of the main camera into a virtual camera space, and the image data is transformed to simulate the output as if the camera had a smoother motion trajectory than the actual camera has during video capture.

[0104] After the EIS processing module 460 stabilizes the image data of a frame, the image data of that frame is output and / or recorded to a data storage device (e.g., a non-volatile storage medium such as flash memory). The zoom level determined by modules 410, 440 for a frame can be used to apply the required digital zoom level to the frame by cropping, upscaling, or in other manner. As a result of these techniques, device 102 can seamlessly transition between the two cameras 110a, 110b during capture, and the transition is automatically managed by device 102 based on the zoom level set by the user. Thus, the resulting video file may contain video segments captured using different cameras 110a, 110b, scattered throughout the video file, and the data is aligned and transformed to show smooth zoom transitions while maintaining consistent EIS processing across segments from both cameras 110a, 110b.

[0105] Device 102 may include a video output and / or recording module 470. The video output and / or recording module 470 receives the output of the EIS processing module 460 and stores it locally and / or remotely as a video file to the device. It may be configured to provide for storage in a separate location. In addition, or alternatively, the video output and / or recording module 470 may stream the output for display locally on device 102 and / or remotely to another display device, for example, over a network.

[0106] Figures 5A to 5C illustrate examples of techniques for multi-camera video stabilization. In the examples described below, one of cameras 110a and 110b is designated as the primary camera, and the other as the secondary camera. Through the processing and transformations discussed below, when the recorded video is captured using the secondary camera, the output of the secondary camera is mapped to the canonical space of the primary camera. For clarity, the first camera 110a is used as the primary camera, and the second camera 110b is used as the secondary camera. This means that in these examples, the camera with the wider field of view is designated as the primary camera. This is desirable in some implementations, but not required. Alternatively, these techniques may be used with a camera with a narrower field of view as the primary camera.

[0107] Generally, a homography transformation is a transformation used to change from one camera space to another. The notation AHB represents a homograph that transforms a point from camera space B to camera space A. A virtual camera is (passed to the user and / or in a video file). This refers to a composite camera view, such as a virtual camera, from which the final scene (to be recorded) is generated. This effective camera position will typically be stabilized over the entire duration of the video (e.g., appearing as stationary as possible in terms of position and orientation) to provide as much temporal and spatial continuity as possible. As used herein, the “primary” camera is the primary camera used to define the reference frame for generating the output video. The primary camera may be defined to be spatially the same as the virtual camera, but not necessarily temporally. In most of the following examples, a first camera 110a is used as the primary camera. A secondary camera is paired with the primary camera. The secondary camera is defined to be spatially separate from the virtual camera, and the output of the secondary camera will be warped into the virtual camera space if the secondary camera is the preceding camera. In most of the following examples, a second camera 110b is used as the secondary camera. “Preceding camera” refers to the camera currently open for video capture, e.g., the camera from which the current image data is being used to generate the saved output video. A “following camera” is paired with the preceding camera. The trailing camera is not currently open or active for video capture. Nevertheless, there may be a startup period in which the trailing camera begins capturing video in anticipation of acquiring status as the trailing camera before the trailing camera's output is actually used by or saved to the output camera. A canonical camera is a conceptual camera with fixed, intrinsic parameters that do not change over time; for example, a canonical camera is unaffected by OIS operation or voice coil motor (VCM) lens shift (for focusing, etc.). For each camera, there is a canonical camera (and corresponding image space), e.g., a canonical primary camera space and a canonical secondary camera space.

[0108] These parameters lead to two different use cases depending on which of cameras 110a and 110b is used for capture. When the primary camera is leading, spatial movement is not required to map the image data between the physical camera's space and space. Composite zoom and EIS processing can simply be done with the EIS processing for the primary camera. On the other hand, when the secondary camera is leading, the system applies spatial movement to map the output to the primary camera's viewpoint for consistency across switching between captures by the two cameras 110a and 110b. This again takes into account changes in focus in scenes where the focal length or subject distance is changed to maintain consistency between the outputs of the two cameras 110a and 110b.

[0109] One of the challenges of providing digital zoom during recording is to efficiently integrate EIS processing with the digital zoom function, particularly with respect to incremental or continuous zoom types using multiple cameras as described above. During video capture, device 102 can record video with EIS processing enabled to acquire a series of temporally stabilized image sequences. Device 102 can achieve this effect by concatenating homography for zoom and EIS functions.

[0110] The homography for EIS and the homography for digital zoom using multiple cameras serve different purposes. The homography for EIS converts image data from the current (e.g., actual) output frame of the image sensor (denoted by the subscript "R" or "real") to a virtual frame (denoted by the subscript "V" or "virt"). In the following various equations and expressions, the variable t (e.g., lowercase t) is the time stamp of the frame, which is usually related to the time of capture of the central scan line of the frame. However, since different scan lines are captured at different times, the time stamps for other scan lines can vary. When referring to the time of an individual scan line (e.g., to account for a time different from the main time of capture for the central scan line of the frame), the time is denoted as t L and represents the time of the capture time stamp of scan line L. When using a camera with a rolling shutter, the time stamps for each scan line of the frame are slightly different, and thus, the term t L can be slightly different for different scan lines, and the camera position can also be slightly different for different scan lines. Different terms T E refer to the extrinsic translation between the main camera 110a and the secondary camera 110b and it should be noted that it does not represent any time stamp. Similarly, n T is the plane norm discussed below and is independent of the extrinsic translation and time terms.

[0111] The EIS homography is V denoted as H R or H eis The homography from EIS is configured to convert data from the current or "actual" frame to a virtual frame by unprojecting points from the current or "actual" frame into 3D space and then projecting them back into virtual space. This homography can be expressed as follows:

[0112]

number

[0113] In Equation 1, R V This represents the 3x3 rotation matrix of the virtual camera, R C This represents the 3x3 rotation matrix of the currently used actual cameras 110a~110b, and R C and R V Both can be obtained from camera position data (e.g., gyroscope data). K is an internal matrix, and K V is the internal matrix of the virtual camera, and K C This is the internal matrix of the current camera (for example, whether camera 110a or 110b is being used). Camera internal data (for example, based on camera geometry, calibration data, and camera characteristics) is stored in one or more data storage devices 230 and can be retrieved from there. The internal matrix can be expressed as shown in Equation 2 below.

[0114]

number

[0115] In equation 2, f is the focal length, and o x and o y is the principal point. In the above equation, R C , R V and K C -1 These values ​​are time-dependent and are expressed as a function of time t. These values ​​can retain or be based on information for previous frames or previous output frame generation processes to ensure temporal continuity. V This is a 3x3 rotation matrix for virtual space, calculated based on a filter of the trajectories of past virtual camera frames, and potentially also calculated for some future frames if delay and "look ahead" strategies are used. CThis is a 3x3 rotation matrix for the current frame, calculated based on data from the gyroscope or device position sensor 220 that converts the current frame to the first frame. C -1 This is calculated based on the optical principal center and the current OIS value.

[0116] Homography for zooming is performed by converting the view from one camera 110a to the view from another camera 110b. For example, this homography converts the current frame from the primary camera 110a to the frame from the secondary camera 110b. main H sec This is shown as follows. This homography may be calculated using a four-point method, but by simplifying this homography as a Euclidean homography, the matrix itself can also be decomposed as follows:

[0117]

number

[0118] Point P on the image from the secondary camera (e.g., telephoto camera 110b). sec The corresponding point P on the image from the main camera (e.g., wide-angle camera 110a). main To bring up to any scalar of S, Assuming a Clid homography transformation, this homography matrix can be decomposed as follows:

[0119]

number

[0120] In this equation, Ext(t) is an external transformation, which is a matrix that depends on the depth of the plane:

[0121]

number

[0122] As a result, the combinations are shown as follows:

[0123]

number

[0124] In equation 4, n T R is the plane norm (e.g., a vector perpendicular to the focal plane), and D is the depth or distance where the focal plane is located. E and T E These indicate external rotation and external translation between the main camera 110a and the sub-camera 110b, respectively.

[0125] Variable K main and K sec These are the internal matrices of the main camera 110a and the sub-camera 110b. These have the same format as the EIS homography described above. In Equation 4, K main , K sec D is time-dependent. However, the corresponding bar in the EIS formulation Unlike John, K main , K sec , and D retain past information in the zoom formulation. No. K main and K sec Both change from frame to frame (for example, OIS position). (Using data 248a~248b) The values ​​are calculated based on the current VCM and OIS values. Variable D is related to the focal length value during autofocus and corresponds to the subject distance, which may change over time during video capture and recording. As described above, the subject distance refers to the distance of the focal plane from device 102 to the current focus selected for the camera. The subject distance can be determined, for example, using the camera's focal position from focal position data 246a~246b and calibration data such as a lookup table showing the correspondence of the focus setting or focal element position to the subject distance, indicating how far the focused subject is from the camera sensor plane. Assuming a positional offset between cameras 110a and 110b, the relationship between captured images may vary somewhat depending on the subject distance, and the transformation can take these effects into account using subject distance, lens focal position, and / or other data.

[0126] The two homography decompositions described above can be summarized as a series of operations: (1) deprojection from the source camera, (2) conversion from the source camera to the world 3D reference frame, (3) conversion from the world to the target 3D reference frame, and (4) reprojection back to the target camera. These are summarized in Table 1 below, and will also be discussed in relation to Figures 5A to 5C and Figure 6.

[0127] [Table 1]

[0128] The following section describes techniques for combining zoom homography and EIS homography in an effective and computationally efficient manner. One usable technique is to set one camera, typically the camera with the widest field of view, as the primary camera and map the processed output with respect to the primary camera.

[0129] If the main camera (e.g., wide-angle camera 110a) is used for video capture, then... mainH sec The identity is constant. As a result, the combined homography is simply EIS homography H, as shown in Equation 7. eis It is possible.

[0130]

number

[0131] term R D This indicates the device rotation calculated based on data from a gyroscope or other motion sensor, and has the same value for both the main camera 110a and the sub-camera 110b.

[0132] Specifically, the three homography transformations are as follows: ● sec_can H sec :(0,0)OIS motion, 0 rolling shutter time, fixed focal length, A transformation from the actual subcamera to the canonical subcamera space, accompanied by rotation around the frame center.

[0133]

number

[0134] Here, R sec -1 (t)*K sec -1 (t) is each scan line or frame, depending on what is required. Obtained from the center, this will work in conjunction with the given OIS / VCM values. sec_can (t) is R in each scan line sec -1 In contrast to (t), this is a rotation around the center of the frame. ● main_can H sec_can : Conversion from canonical secondary camera space to canonical primary camera space.

[0135]

number

[0136] In this equation, Ext(t) is an external transformation, which is a matrix that depends on the subject distance (e.g., the depth in space of the focal plane from the camera sensor):

[0137]

number

[0138] Here, n T is the plane norm, and D is the plane n T This is the depth, R E and T E This refers to external rotation and translation between the secondary camera and the primary camera. ● virt H main_can This is a conversion from a canonical subcamera to a stabilized virtual camera space, and this The rotation is filtered by the EIS algorithm.

[0139]

number

[0140] R main_can =R sec_can Please note that this represents the current actual camera rotation at the center of the frame, because the main camera and sub-camera are firmly attached to each other.

[0141] By concatenating three homographs, the final homograph is obtained:

[0142]

number

[0143] Here, the intermediate term main_can H sec_can This is due to the zoom process. If the primary camera is leading, the final homography is simplified to only the homography required for EIS processing. The above equation is as follows:

[0144]

number

[0145] this is, main_can H sec_can = Setting the identity, the above equation is the original H eis It becomes an equation:

[0146]

number

[0147] From the above section, we generally end up with two formulas: When the secondary camera is leading, the equation is as follows:

[0148]

number

[0149] When the main camera is leading, the formula is as follows:

[0150]

number

[0151] From an engineering perspective, during the implementation of EIS, there are no secondary cameras; therefore, all virtual cameras are located on top of the current preceding camera, which means the following for the original EIS pipeline: When the main camera is leading,

[0152]

number

[0153] When the secondary camera is leading,

[0154]

number

[0155] From the basic formula of EIS,

[0156]

number

[0157]

number

[0158] term K V It is always defined the same as the current camera, and according to the definition, K v_main and K v_sec Between The only difference is the field of view between the two cameras (since the virtual camera is positioned at the center of the image, the principal point is the same).

[0159] To address the above case, the system can intentionally ensure that the fields of view in the primary and secondary cameras coincide at the switching point. This field of view matching is efficiently performed through hardware cropping that scales the field of view from the secondary camera to the primary camera prior to all operations, and the matrix S is used in the following equation.

[0160]

number

[0161] From the software side, the homography formula conforms to the following:

[0162]

number

[0163] The final homography applied may be as follows:

[0164]

number

[0165] The transformation maps the canonical subspace to a normalized canonical subspace that has the same field of view (e.g., scale) as the canonical subspace, but in which all translations / rotations caused by camera-external factors are neutralized.

[0166] A camera considered to be a secondary camera (e.g., telephoto camera 110b) is used for video capture. When used, there is a set of transformations used to convert or project the view from the secondary camera to the view from the primary camera in order to ensure that consistent image characteristics are maintained over the duration of the capture using different cameras 110a and 110b. These transformations may include (1) deprojecting from the source camera (e.g., the second camera 110b) to remove camera-specific effects such as OIS motion and rolling shutter, (2) converting to a canonical reference frame such as a three-dimensional world, (3) converting from the canonical reference frame to a primary camera reference (e.g., aligning to the frame that will be captured by the first camera 110a), and (4) reprojecting from the frame of the primary camera 110a to a virtual camera frame to which electronic stabilization has been applied. Based on the order of transformations, this technique may be performed in parallelization first or stabilization first, which will result in different performance outcomes. Figures 5A to 5C illustrate different techniques for converting an image captured from the second camera 110b to a stabilized view that is aligned with and consistent with the stabilized view produced for the first camera 110a.

[0167] Figure 5A shows an exemplary technique for multi-camera video stabilization that performs parallelization first. This technique first parallelizes the view from the second camera 110b (e.g., the telephoto camera in this example) to the first camera 110a (e.g., the wide-angle camera in this example), and then applies stabilization. This technique has the advantage of being represented by a simple concatenation of EIS homography and primary-to-subprimary homography, which increases the efficiency of processing the video feed. This is shown in the following equation:

[0168]

number

[0169] Figure 5B shows an exemplary technique for multi-camera video stabilization that performs stabilization first. This option stabilizes the image from the secondary camera frame into a stabilized virtual primary camera frame, and then applies parallelization to project the virtual secondary camera frame onto the primary camera frame. However, this technique is not as simple as the method in Figure 5A because the transformation from secondary camera to primary camera is defined in the primary camera frame coordinate system. Therefore, to perform the same transformation, the stabilized virtual primary camera frame must be rotated back from the virtual secondary frame. This is expressed by the following equation, "virt_main " refers to the virtual frame (stabilized) of the main camera, "virt_sec" refers to the virtual frame (stabilized) of the secondary camera, and "real_main" refers to the actual frame (stabilized) of the main camera. "real_sec" refers to the actual frame from the secondary camera (not present), while "real_sec" refers to the actual frame from the secondary camera.

[0170]

number

[0171] By concatenating the items, the result is the same overall transformation as in Figure 5A. Even if different geometric definitions exist, the overall effect on the image is the same.

[0172]

number

[0173] Figure 5C shows an exemplary technique for multi-camera video stabilization, first parallelizing to a canonical camera. Using the geometric definitions discussed above, the implementation can be made more efficient by parallelizing from the current real camera view to a canonical camera view, where the canonical camera view is defined as a virtual camera of the current preceding camera, but has a fixed OIS (OIS_X=0, OIS_Y=0) and a fixed VCM (VCM=300).

[0174] This result is similar to the equation in the first case in Figure 5A. For example, this gives:

[0175]

number

[0176] However, in this case, K main -1 (t)K main (t) Instead of inserting pairs, K main_can -1 (t)*K main_can (t) will be inserted here, K main_can This shows the inherent characteristics of the canonical primary camera. Term H' eis and main H' sec is, K main_can It is calculated using

[0177]

number

[0178] This approach offers significant advantages because the system does not need to query metadata for both the secondary and primary cameras simultaneously.

[0179] Combined homography, H combined can be used with the canonical position representation and the per-scanline processing. First, the canonical position will be described. For the discussed H eis and main H sec from its original definition, the system uses the OIS and VCM information for both the K sec and K main terms used in the homography for digital zoom processing. This system also uses the term K used in the EIS homography. C

[0180] In the following equations, certain variables are time-dependent variables that depend on the current OIS and VCM values for each of the main camera and the sub-camera. These time-dependent variables include K main -1 (t), K main (t), and K sec -1 (t). The time-dependence of these terms means that the system will need to stream both the OIS / VCM metadata of both cameras 110a, 110b simultaneously. However, using the canonical representation can significantly reduce the data collection requirements. For example, the terms K include that both OIS and VCM are defined as a canonical camera model where both are placed at a predetermined standard or canonical position, and the canonical position is a predetermined position with OISX = OISY = 0 and VCM = 300. An exemplary representation is shown below. For the original H sec C and K main C its representation can be decomposed as follows: For the original H combined = H eis * main H sec it can be decomposed as follows:

[0181]

Equation

[0182] In the above equation, term C sec (t) converts the current camera view to a canonical camera view. This is a correction matrix (for example, to remove the effect of the OIS lens position). Since digital zoom homography is only active when capturing video from the second camera 110b (e.g., the telephoto camera), one correction matrix term C transforms the current subcamera view into a canonical subcamera view. sec Only (t) is needed. In this new formula, K main C oh Call K sec C Both are constant over time for homography for digital zoom:

[0183]

number

[0184] As a result, the only time-dependent component is D(t), which depends on the subject distance for the focal point, which can be the distance from the camera to the focal plane. Therefore, the only new metadata required would be the subject distance.

[0185] Next, we will explain the correction for each scan line. The correction for each scan line is expressed in another matrix S(t) on the right side of the above equation. L The use of a scan line correction matrix S(t) can be included. This additional matrix can be used to perform scan line-specific corrections with different adjustments potentially shown for each scan line L. L , L) is the current time t of the scan line (for example, to obtain gyroscope sensor data at the appropriate time when the scan line was captured). L , and the scan line number L (used, for example, to correct for deformation or shift in order to align each scan line with the central scan line). As a result, the matrix S(t L, L) is the correction for each scan line L and its corresponding capture time t L can include.

[0186] From the original decomposed equation, the addition of the correction matrix for each scan line gives:

[0187]

Number

[0188] The final homography can be determined by adding the canonical position and terms for each scan line:

[0189]

Number

[0190] In this example, H C eis = K V *R V (t) * R D -1 (t) * K main C-1 represents the stabilization matrix in only the canonical main camera. Additionally, the term main H C sec = K main C * (R E -1 - D(t)*T E *n T ) * K sec C-1 represents the homography from the canonical secondary view to the canonical main view. The term C sec (t) is the homography that makes the current secondary camera view the canonical secondary camera view. The term S(t L , L) is the correction for each scan line that brings the deformation of each scan line to the reference position of the center line.

[0191] Figure 6 shows an exemplary transformation that may be used to efficiently provide multi-camera video stabilization. The combination of EIS and continuous zoom can be interpreted as a combination of three homographs or transformations.

[0192] In this example, image frame 611 represents an image frame captured by the second camera 110b. Through several transformations shown in the figure, the image data is processed to remove visual artifacts (e.g., rolling shutter, OIS motion, etc.), aligned with the field of view of the first camera 110a, and stabilized using the EIS technique described above. The separate transformations 610, 620, and 630, as well as the various images 611-614, are shown for illustrative purposes. The implementation can combine the operations and functions discussed without the need to separately generate intermediate images.

[0193] The first transformation 610 operates on the image 611 captured from the actual secondary camera 110b (e.g., a telephoto camera) and transforms the image 611 into a canonical second camera image 612. The transformation 610 adjusts the image data of image 611 so that the canonical second camera image 612 does not include OIS motion, does not include rolling shutter, includes a fixed focal length, and involves rotation at the center of the frame. In some implementations, the first transformation 610 may individually adjust each image scan line of image 611 to eliminate or reduce the effects of motion of the second camera 110b during the capture of image 611, changes in focal length during the capture of image 611, and motion of the OIS system during the capture of image 611. This transformation 610 is sec_can H sec This is represented as a conversion from the second camera view ("sec") to the canonical second camera view ("sec_can"). During the capture of image 611 OIS stabilization may be used, but neither image 611 nor image 612 has been stabilized using EIS processing.

[0194] The second transformation 620 is a transformation from the canonical second camera image 612 to the canonical first camera image 613. This allows the image data to be kept within the same canonical space as image 612, but the image data can be aligned with the scene captured by the primary camera, e.g., the first (e.g., wide-angle) camera 110a. This transformation 620 can compensate for the spatial difference between the second camera 110b and the first camera 110a. The two cameras 110a, 110b are located on a phone or other device with a spatial offset between them and potentially other differences in position or orientation. The second transformation 612 can compensate for these differences so that image 612 can be projected onto the corresponding portion of the field of view of the first camera 110a. In addition, the difference between the views of cameras 110a, 110b may vary depending on the current depth of field. For example, one or both of cameras 110a, 110b may experience focus breathing, which adjusts the effective field of view depending on the focal length. The second transformation 620 can take these differences into account and allows device 102 to fine-tune the alignment of image 612 with respect to the primary camera canonical field of view using the focal length. Typically, the same canonical parameters for OIS position, etc., are used for both the primary camera canonical representation and the secondary camera canonical representation, but if there are differences, these can be corrected with respect to using the second transformation 620.

[0195] A third transformation 620 applies EIS to image 613 to generate a stabilized image 614. For example, this transformation 620 can transform image data from a canonical primary camera view into a virtual camera view in which the camera position is smoothed or filtered. For example, this transformation 630 can be the second projection discussed above with respect to Figure 3, in which the virtual camera position is filtered over time, changes in position between frames are used, motion is potentially tolerated due to blurring, the possibility of panning versus accidental motion is considered, and adaptation of future motion or adjustments to fill the output frame is performed. EIS processing can be performed using the current frame being processed, as well as windows of previous frames and future frames. Naturally, EIS processing may be delayed from the image capture timing that collects the "future frames" required for processing.

[0196] As described above, the transformations 610, 620, and 630 can be combined or integrated for efficiency, and there is no need to generate intermediate images 612 and 613. Rather, device 102 may determine the appropriate transformations 610, 620, and 630 and directly generate a stabilized image 614 that is aligned with the stabilized image data generated from the image captured using the first camera 110a and consistent with it.

[0197] The example in Figure 6 illustrates the transformation from video frames from the second camera 110b to the stabilized space of the first camera 110a, which is used when the zoom level corresponds to the same or smaller field of view as the second camera 110b. When the first camera 110a is used, for example, when the field of view is larger than that of the second camera 110b, only two transformations are required. Similar to the first transformation 610, the transformation is applied to remove time-dependent effects from the captured image, such as conditions that change for different scan lines of the video frame. Similar to transformation 610, this compensates for rolling shutter, OIS position, etc. However, the transformation projects the image data directly into the canonical camera space of the primary camera (e.g., camera 110a). From the canonical camera space of the primary camera, only an EIS transformation, e.g., transformation 630, is required. Therefore, when image data is captured using the main camera, it is not necessary to relate the spatial characteristics of the two cameras 110a and 110b, because the overall image capture and EIS processing is performed consistently, for example, from a canonical, time-independent reference frame for the main camera to a reference frame for the main camera.

[0198] The processing in Figure 6 illustrates various processes that can be used to convert or map the output of the second camera 110b to the stabilized output of the first camera 110a. This provides consistency in the view when transitioning between using videos captured by different cameras 110a, 110b during video capture and recording. For example, the video from the second camera 110b is aligned with the video from the first camera 110a, but not only is the field of view positioned correctly, the EIS characteristics are also matched. As a result, switching between cameras during image capture (e.g., by digital zoom above or below a threshold zoom level) can be done with the video feed matched in a way that minimizes or avoids jerky movements, sudden image or view shifts, and visible video smoothness shifts (e.g., changes in EIS application) at the transition point between cameras 110a, 110b, while EIS is active. Another advantage of this technique is that the video from camera 110a can be used without any adjustment or processing for consistency with the second camera 110b. Only the output of the second camera 110b is adjusted and aligned to be consistent with the view and characteristics of the first camera 110a.

[0199] Device 102 can also adjust video capture settings during video capture to better match the characteristics of the video captured using different cameras 110a and 110b. For example, when transitioning from a capture using the first camera 110a to a capture using the second camera 110b, device 102 can adjust the focal length, OIS parameters (e.g., the same parameters used for the first camera 110a immediately before the transition) that were used for the first camera 110a. The device can then determine characteristics such as whether OIS is enabled, the intensity or level of OIS applied, and exposure settings (e.g., ISO, sensor sensitivity, gain, frame capture time, or shutter speed). Device 102 can then cause the second camera 110b to use those settings, or settings that provide equivalent results, for capture for the second camera 110b.

[0200] This may involve making configuration changes prior to the transition in order to include video from the second camera 110b in the recorded video. For example, to provide time to make adjustments during operation, device 102 may detect whether a camera switch is appropriate or necessary, and in response determine the current settings of the first camera 110a and instruct the second camera 110b to start using settings equivalent to those used by the first camera 110a. This can provide sufficient time to, for example, power on the second camera 110b, activate and achieve stabilization of the OIS system for the second camera 110b, adjust the exposure of the second camera 110b to match the exposure of the first camera 110a, and set the focal position of the second camera 110b to match the focal position of the first camera 110a. If the second camera 110b is operating in the appropriate mode, for example, using video capture parameters consistent with those of the first camera 110a, device 102 switches from using the first camera 110a for video capture to using the second camera 110b. Similarly, when transitioning from the second camera 110b to the first camera 110a, device 102 can determine the video capture parameter values ​​used by the second camera 110b and set the corresponding video capture parameter values ​​for the first camera 110a before making the transition.

[0201] Several implementations have been described. Nevertheless, it will be understood that various modifications can be made without departing from the spirit and scope of this disclosure. For example, the flows shown above may be used in various forms by rearranging, adding, or removing steps. As another exemplary modification, although the above description primarily describes that the processing of image data is performed while video is being captured by the first and / or second cameras, it will be understood that in some implementations, the first transformation to a second canonical reference space for the second camera, the second transformation from the second canonical reference space to the first canonical reference space for the first camera, and the third transformation for applying electronic image stabilization to the image data in the first canonical reference space for the first camera may instead be applied later (by device 102 or another, e.g., remote device), e.g., when video is no longer being captured.

[0202] All embodiments of the present invention and the functional operations described herein may be implemented in digital electronic circuits, or in computer software, firmware, or hardware including the structures disclosed herein and their structural equivalents, or in one or more combinations thereof. Embodiments of the present invention may be implemented as one or more modules of computer program instructions encoded on a computer-readable medium for execution by one or more computer program products, for example, a data processing device, or for controlling the operation of a data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of a material that provides a machine-readable propagating signal, or one or more combinations thereof. The term “data processing device” includes, for example, all devices and machines for processing data, including programmable processors, computers, or multiple processors or computers. In addition to hardware, a device may include code that creates the execution environment for the computer program, such as processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. It may include a code. A propagating signal is an artificially generated signal, for example, a mechanically generated electrical, optical, or electromagnetic signal produced to encode information for transmission to a suitable receiver device.

[0203] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as part of a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program, or multiple cooperating files (e.g., files that store one or more modules, subprograms, or portions of code). A computer program can be located on one computer or site, or distributed across multiple sites and executed on multiple computers interconnected by a communication network.

[0204] The processes and logical flows described herein can be executed by one or more programmable processors executing one or more computer programs to perform functions by operating input data and generating output. Execution of the processes and logical flows, as well as implementation of the apparatus, can also be done by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).

[0205] Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Generally, a processor will receive instructions and data from read-only memory, random-access memory, or both. Essential elements of a computer are a processor for executing instructions, and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or will be operationally coupled to receive data from or transfer data to such mass storage devices, or both. However, a computer is not required to have such devices. Furthermore, a computer can be incorporated into another device, for example, a tablet computer, a mobile phone, a personal digital assistant (PDA), a mobile audio player, or a Global Positioning System (GPS) receiver, to name just a few. Computer-readable media suitable for storing computer program instructions and data include, for example, all forms of non-volatile memory, media, and memory devices, including semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into special-purpose logic circuits.

[0206] To provide interaction with the user, embodiments of the present invention are implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, on which the user can provide input to the computer. This is possible. Interaction with the user can also be provided using other types of devices; for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or haptic feedback; input from the user may be received in any form, including acoustic input, voice input, or haptic input.

[0207] Embodiments of the present invention may be implemented, for example, in a computing system including a backend component as a data server; for example, in a computing system including a middleware component such as an application server; for example, in a computing system including a frontend component such as a client computer having a graphical user interface or a web browser that allows a user to interact with an embodiment of the present invention; or in a computing system of any combination of one or more such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication, for example, by a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), for example, the Internet.

[0208] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship between a client and a server arises from computer programs that run on each computer and have a client-server relationship with each other.

[0209] This specification contains many details, which should not be construed as limitations on the scope of the invention or the scope of the claims, but rather as descriptions of features specific to particular embodiments of the invention. Certain features described herein in the context of separate embodiments may also be realized in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be realized separately or in any preferred partial combination in multiple embodiments. Furthermore, features may be described above as acting in a certain combination, and may even be initially claimed as such, but in some cases one or more features from the claimed combination may be removed from that combination, and the claimed combination may be directed towards a partial combination or a variation of a partial combination.

[0210] Similarly, although operations are shown in a specific order in the diagrams, it should not be understood that such operations must be performed in that specific or sequential order shown to achieve the desired result, or that all shown operations must be performed. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated into a single software product or packaged into multiple software products.

[0211] Specific embodiments of the present invention have been described. Other embodiments are within the scope of the following claims. For example, the steps described in the claims may be performed in a different order and still achieve the desired results.

Claims

1. It is a method, A video capture device having a first camera and a second camera includes providing a digital zoom function that enables user-specified magnification changes within a digital zoom range during video recording, wherein the video capture device is configured to (i) use video data captured by the first camera over a first portion of the digital zoom range, and (ii) use video data captured by the second camera over a second portion of the digital zoom range, and the method further includes, A method comprising processing image data captured using the second camera by applying a set of transformations, including (i) a first transformation to a second canonical reference space for the second camera, (ii) a second transformation from the second canonical reference space for the second camera to the first canonical reference space for the first camera, and (iii) a third transformation for applying electronic image stabilization to image data in the first canonical reference space for the first camera, while capturing video using the second camera of the video capture device to provide a certain zoom level in the second portion of the digital zoom range.

2. The above method further, The method according to claim 1, wherein, in the first portion of the digital zoom range, while capturing video using the first camera of the video capture device to provide a certain zoom level, the method comprises processing image data captured using the second camera by applying a set of transformations including (i) the first transformation to the first canonical reference space for the first camera and (ii) the third transformation for applying electronic image stabilization to image data in the first canonical reference space for the first camera.

3. The method according to claim 1 or 2, wherein the first camera and the second camera have different fields of view, (i) the field of view of the second camera is included within the field of view of the first camera, or (ii) the field of view of the first camera is included within the field of view of the second camera.

4. The method according to any one of claims 1 to 3, wherein the first camera and the second camera each include a fixed focal length lens assembly.

5. The method according to any one of claims 1 to 4, wherein the second canonical reference space for the second camera and the first canonical reference space for the first camera are conceptual camera spaces defined by a predetermined fixed set of camera intrinsic properties such that projecting image data onto the canonical reference space eliminates time-dependent effects during video frame capture.

6. The first camera includes an optical image stabilization (OIS) system, and the first canonical reference space for the first camera is such that the image data has a consistent predetermined OIS position, or The method according to claim 5, wherein the second camera includes an optical image stabilization (OIS) system, and the second canonical reference space for the second camera is such that the image data has a consistent predetermined OIS position.

7. The first camera provides image data that progressively captures the image scan lines of an image frame, and the first canonical reference space for the first camera is such that the image data is the image It is corrected to remove distortion caused by the gradual capture of the image scan lines, or The method according to claim 6, wherein the second camera provides image data that progressively captures the image scan lines of an image frame, and the second canonical reference space for the second camera is corrected so that the image data removes distortion resulting from the progressive capture of the image scan lines.

8. The method according to any one of claims 1 to 7, wherein the second transformation aligns the field of view of the second camera with the field of view of the first camera and adjusts the spatial offset between the first camera and the second camera.

9. The method according to any one of claims 1 to 8, wherein the first transformation, the second transformation, and the third transformation each have a corresponding homography matrix, and processing the image data includes applying the homography matrix.

10. The above method further, During the capture of video data using the first camera and processing of the video data from the first camera for applying electronic image stabilization, the system receives user input indicating a change in the zoom level to a specific zoom level in the second portion of the digital zoom range. In response to receiving the user input, Recording a sequence of video frames in which the magnification of the video frames captured using the first camera is gradually increased until a predetermined zoom level is reached, To start video capture using the second camera mentioned above, The method according to any one of claims 1 to 9, comprising recording a second sequence of video frames captured using the second camera, wherein the second sequence of video frames provides a predetermined zoom level and provides an increasing magnification of the video frames captured using the second camera until the specific zoom level is reached.

11. The first transformation includes several different adjustments to different scan lines of the image data captured using the second camera, The second conversion is determined at least in part on the focal length of the second camera, The method according to any one of claims 1 to 10, wherein the third conversion includes, for each specific video frame among the video frames, electronic image stabilization using one or more video frames preceding the specific video frame and one or more video frames following the specific video frame.

12. The second camera has a smaller field of view than the first camera. The above method further, During image capture using the first camera, the system receives user input indicating a change in the zoom level for video capture, The method according to any one of claims 1 to 11, comprising determining whether the changed zoom level is equal to or greater than a predetermined transition zoom level in response to receiving the user input, wherein the predetermined transition zoom level represents a field of view smaller than the field of view of the second camera.

13. The above method further, (i) storing data indicating a first transition zoom level for transitioning from a video capture using the first camera to a video capture using the second camera, and (ii) storing data indicating a second transition zoom level for transitioning from a video capture using the second camera to a video capture using the first camera, wherein the first transition zoom level is different from the second transition zoom level, and the method further includes, The method according to any one of claims 1 to 12, comprising determining whether to switch between cameras for video capture by (i) comparing the requested zoom level to the first transition zoom level when the requested zoom level corresponds to a decrease in the field of view, and (ii) comparing the requested zoom level to the second transition zoom level when the requested zoom level corresponds to an increase in the field of view.

14. The above method further, During the recording of a video file, it is decided to switch from capturing video using a specific camera among the first and second cameras to capturing video using a different camera among the first and second cameras, In response to the decision to switch, Determining the values ​​of video capture parameters used for image capture using the aforementioned specific camera, Based on the determined video capture parameters, set the values ​​of the video capture parameters for the other camera, The process includes setting the values ​​of the video capture parameters for the other camera, starting video capture from the second camera, and recording the captured video from the second camera to the video file. The method according to any one of claims 1 to 13, wherein setting the video capture parameters includes adjusting one or more of the following for the second camera: exposure, image sensor sensitivity, gain, image capture time, aperture size, lens focal length, OIS status, or OIS level.

15. A video capture device, A first camera having a first field of view, A second camera having a second field of view, One or more position or orientation sensors, One or more processors, A video capture device comprising one or more data storage devices that store instructions, when executed by one or more processors, cause the method described in any one of claims 1 to 14 to be performed.