Smooth Continuous Zoom in Multi-Camera Systems via Image-Based Visual Features and Optimized Geometry Calibration

An image-based approach in multi-camera systems uses visual features to correct geometry-based warping transformations, addressing distortions caused by VCM/OIS adjustments and thermal effects, resulting in smoother and more stable zoom transitions.

JP2025536458APending Publication Date: 2025-11-06GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025518563
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-25
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

Multi-camera systems in computing devices experience perceptible image distortions, such as binocular disparity, during transitions between cameras due to changes in field of view, which are exacerbated by factors like VCM/OIS adjustments and thermal effects, leading to abrupt transitions and reduced image quality.

Method used

An image-based approach that utilizes image features to correct geometry-based warping transformations, combining image-based visual information with geometry-based calibration to improve the accuracy of warping transformations, reducing spatial differences between camera frames and mitigating the impact of inaccurate sensor metadata.

Benefits of technology

The solution effectively reduces visual artifacts during camera transitions, enhancing the smoothness and stability of zooming processes in multi-camera systems, improving the quality of captured images and videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536458000005
    Figure 2025536458000005
  • Figure 2025536458000006
    Figure 2025536458000006
  • Figure 2025536458000007
    Figure 2025536458000007
Patent Text Reader

Abstract

An exemplary method includes displaying an initial preview of a scene captured by a first camera operating within a first focal length range. The method includes detecting a zoom operation that is predicted to cause the first camera to reach a first range limit. The method includes activating a second camera operating within a second focal length range to capture a zoom preview of the scene. The method includes updating a geometry-based warping transform based on a comparison of respective image features from the initial preview and the zoom preview. The method includes aligning the zoom preview with the initial preview by applying the updated warping transform. The method includes displaying the aligned zoom preview of an image captured by the second camera while operating within the second range.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE / INCORPORATION BY REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Patent Application No. 63 / 377,581, filed September 29, 2022, which is incorporated herein by reference in its entirety. [Background technology]

[0002] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices. Some image capture devices are composed of multi-camera systems. The camera systems are configured to cooperatively meet different image capture requirements using their respective specifications. Smartphones can integrate multiple types of cameras with various focal lengths to address objects at different distances and scenes within different fields of view (FOV). Summary of the Invention

[0003]

[0001] The present disclosure generally relates to smooth transitions between multiple cameras. In one aspect, an image capture device may include multiple cameras. When transitioning between cameras, perceptible image distortions, such as binocular disparity, may occur due to changes in field of view. As described herein, a warping transformation is estimated from image-based data as well as available geometry metadata to warp an image from one camera to approximately align it with an image from another camera, thereby reducing perceptible image distortions during camera switching.

[0004] In a first aspect, a computer-implemented method is provided. The method includes displaying, on a display screen of the computing device, an initial preview of a scene captured by a first image capture device of the computing device, the first image capture device operating within a first focal length range. The method also includes detecting, by the computing device, a zoom operation predicted to cause the first image capture device to reach a limit of the first focal length range. The method further includes, in response to the detecting, activating a second image capture device of the computing device to capture a zoom preview of the scene, the second image capture device configured to operate within a second focal length range. Additionally, the method includes updating a geometry-based warping transform based on a comparison of respective image features from the initial preview and the zoom preview. The method further includes aligning the zoom preview with the initial preview by applying an updated warping transform, the updated warping transform reducing one or more visual artifacts caused by a change in field of view when transitioning from the initial preview to the zoom preview. The method also includes displaying, on the display screen of the computing device, the aligned zoom preview of an image captured by the second image capture device while operating within the second focal length range.

[0005] In a second aspect, a computing device is provided that includes a display screen, a first image capture device configured to operate within a first focal length range, a second image capture device configured to operate within a second focal length range, one or more processors, and data storage, the data storage storing computer-executable instructions that, when executed by the one or more processors, cause the mobile device to perform functions. The operations include displaying, by a display screen, an initial preview of a scene captured by a first image capture device, the operations further including detecting, by a computing device, a zoom operation that is likely to cause the first image capture device to reach a limit of a first focal length range, the operations further including activating a second image capture device to capture a zoom preview of the scene in response to the detecting, the operations further including updating a geometry-based warping transform based on a comparison of respective image features from the initial preview and the zoom preview, the operations further including aligning the zoom preview with the initial preview by applying the updated warping transform, the updated warping transform reducing one or more visual artifacts caused by changes in field of view when transitioning from the initial preview to the zoom preview, and the operations further including displaying, by a display screen of the computing device, the aligned zoom preview of an image captured by the second image capture device while operating within the second focal length range.

[0006] In a third aspect, an article of manufacture is provided. The article of manufacture may include a non-transitory computer-readable medium having stored thereon program instructions that, when executed by one or more processors of a computing device, cause the computing device to perform operations including displaying, by a display screen, an initial preview of a scene captured by a first image capture device, further including detecting, by the computing device, a zoom operation that is likely to cause the first image capture device to reach a limit of a first focal length range, activating a second image capture device in response to the detecting to capture a zoom preview of the scene, further including updating a geometry-based warping transform based on a comparison of respective image features from the initial preview and the zoom preview, further including aligning the zoom preview with the initial preview by applying the updated warping transform, wherein the updated warping transform reduces one or more visual artifacts caused by changes in field of view upon transitioning from the initial preview to the zoom preview, and further including displaying, by the display screen of the computing device, the aligned zoom preview of an image captured by the second image capture device while operating within a second focal length range.

[0007] In a fourth aspect, a system is provided, the system including: means for displaying, by a display screen, an initial preview of a scene captured by a first image capture device; means for detecting, by a computing device, a zoom operation that is likely to cause the first image capture device to reach a limit of a first focal length range; means for activating a second image capture device to capture a zoom preview of the scene in response to the detecting; means for updating a geometry-based warping transform based on a comparison of respective image features from the initial preview and the zoom preview; means for aligning the zoom preview with the initial preview by applying the updated warping transform, the updated warping transform reducing one or more visual artifacts caused by changes in field of view upon transitioning from the initial preview to the zoom preview; and means for displaying, by the display screen of the computing device, the aligned zoom preview of an image captured by the second image capture device while operating within the second focal length range.

[0008] Other aspects, embodiments, and implementations will become apparent to those skilled in the art from a reading of the following detailed description, with appropriate reference to the accompanying drawings. [Brief explanation of the drawings]

[0009] [Figure 1] 1 illustrates binocular disparity in a multi-camera system, according to an exemplary embodiment. [Figure 2] 1 is a flowchart of a workflow for image-based computation of a warping transformation, according to an example embodiment. [Figure 3] 1 is an exemplary sparse feature workflow for smooth continuous zoom in a multi-camera system, according to an exemplary embodiment. [Figure 4A]1 illustrates temporal feature matching and tracking, according to an example embodiment. [Figure 4B] 10 shows an example image of temporal feature matching and tracking, according to an example embodiment. [Figure 4C] 4 illustrates an example application 400 of temporal feature matching and tracking, according to an example embodiment. [Figure 5] 1 is an exemplary dense feature workflow for smooth continuous zoom in a multi-camera system, according to an exemplary embodiment. [Figure 6] 10 illustrates exemplary processing of delta data during a camera transition, according to an exemplary embodiment. [Figure 7] 10 is a table illustrating various cases for switching between a telephoto camera and a wide-angle camera, according to an example embodiment. [Figure 8A] 10 illustrates exemplary geometric relationships between each pair of matched pixels, according to an exemplary embodiment. [Figure 8B] 1 illustrates a workflow for determining geometric relationships between each pair of matched pixels, according to an example embodiment. [Figure 9] 1 illustrates an exemplary workflow for smooth continuous zoom in a multi-camera system, according to an exemplary embodiment. [Figure 10] 1 illustrates a distributed computing architecture in accordance with an exemplary embodiment. [Figure 11] FIG. 1 is a block diagram of a computing device in accordance with an exemplary embodiment. [Figure 12] 1 is a flowchart of a method according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] Exemplary methods, devices, and systems are described herein. Moreover, the words "exemplary" or "exemplary" are used herein to mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "exemplary" or "exemplary" should not necessarily be construed as preferred or advantageous over other embodiments or features. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.

[0011] Accordingly, the exemplary embodiments described herein are not intended to be limiting. The aspects of the present disclosure, as generally described and illustrated in the Figures herein, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are contemplated herein.

[0012] Furthermore, unless the context suggests otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should be viewed generally as component aspects of one or more overall embodiments, with the understanding that not all of the illustrated features are required for each embodiment.

[0013] overview A smartphone or other mobile device that supports image and / or video capture may include multiple cameras that use different specifications to cooperatively meet different image capture requirements. A smartphone may integrate multiple types of cameras with various focal lengths to view and / or capture objects at different distances and scenes within different fields of view (FOVs).

[0014] For example, a phone may be configured with a main camera with a medium focal length to meet normal photo / video capture requirements, a telephoto camera with a long focal length to capture distant objects, and an ultra-wide-angle camera with a short focal length to capture a larger FOV. During a photo / video capture session, a switch from the main to the telephoto camera may occur as the user continues to zoom in to focus on distant objects, and a switch from the main to the ultra-wide-angle camera may occur as the user continues to zoom out to capture a wider field of view. Multi-camera systems offer a much larger focal length range than a single camera. However, a sudden switch between cameras during zooming may result in a mismatch of views (known as "binocular disparity").

[0015] The warping transformation is estimated from available geometric metadata and image features to avoid binocular disparity and warp the image from one camera to approximately align with the image from the other camera, thereby reducing the perceptual change during camera switching. The warping transformation can include scaling, rotation, reflection, identity mapping, shearing, or various combinations thereof. For example, translation, similarity, affine mapping, and / or projective mapping may also be used as the warping transformation. Generally speaking, two planar images can be related by a warping transformation, such as a homography. For example, computer vision approaches for computing homographies can be used to warp image frames from one camera to the other. As described herein, homography calculations can be determined to reduce view disparity during camera switching while zooming. The homography calculations can use geometry information (not including image features) including metadata such as camera calibration data, focal length, etc. While a geometry-based solution may be used, the presence of electrical and / or mechanical components, such as a device's voice coil motor (VCM), optical image stabilization (OIS) adjustments, and / or thermal effects, can cause calibrations and changes in the focal length of a dynamic camera, which can result in errors in determining an accurate warping transformation for a smooth viewing experience, resulting in abrupt transitions.

[0016] Some existing approaches attempt to solve this problem. For example, physical camera views may be warped to the same coordinates, and the warping transformation may depend on the camera's calibration and focal length. However, because image-based features are not used, errors due to VCM / OIS adjustments and / or thermal effects may remain uncorrected. Another approach may be to blend multiple camera views and apply fade-style animations to obtain smooth transitions. However, this approach relies on the simultaneous display of images from different cameras. Theoretical models for using binocular disparity and motion parallax for depth estimation have been proposed, but they do not have any practical implementation to solve technical issues related to image capture devices.

[0017] This application relates to an "image-based" approach (sometimes referred to herein as ContiZoom) that better supports warping quality to overcome the adverse effects of VCM / OIS adjustments and / or thermal effects. In contrast to "geometry-based" approaches, the new "image-based" approach is designed to utilize image information and / or features as additional input to improve the warping transformation used in the geometry-based approach.

[0018] The approach described herein directly utilizes image features to provide a more accurate metric for computing the warping transformation, which reduces the spatial difference between image frames from the two cameras and mitigates the impact of many inaccurate sensor metadata from geometry-based approaches.

[0019] From a geometry perspective, thermal changes to the device can affect the principal point and therefore shift the entire FOV, causing the output from Camera Parameter Interpolation (CPI) to be unreliable. Because thermal changes affect focal length, with successive frames and continued use of the device, an already inaccurate focal length can become more unstable due to additional thermal effects.

[0020] These factors can be largely mitigated by utilizing image information (features) to adjust inaccurate geometry metadata. For example, image feature matching may be performed between two frames from two different physical cameras. Existing geometry metadata may be corrected based on the image features. The geometry-based warping transform may be recalculated based on the corrected geometry metadata based on the image features.

[0021] As described herein, dual images from a camera pair during switching are used for the image-based smooth zoom described herein. In the event that the quality of continuous zoom is adversely affected by thermal changes or inaccurate estimation of focal length, the resulting problems can be effectively solved by image-based visual information. Bundle adjustment can be applied to camera calibration and world points obtained from visual feature matching, so that optimized parameters produce more reliable homographies for image warping. Scene depth can be estimated from both image-based visual features and phase differences, which can result in improved smoothness of zoom during camera switching.

[0022] In some embodiments, image-based algorithmic processes may run at up to 30 frames per second (fps) and can be configured to work seamlessly with other camera features such as image distortion correction, video stabilization, etc. Computationally intensive steps such as visual feature extraction can be made less intensive by using multi-threading and DSP solutions.

[0023] Using image-based visual features has several advantages, such as the fact that images may be conveniently available from the device's camera system (e.g., in regular RGB format). Because image alignment during camera switches is a desired outcome of continuous zooming, warping estimated from the images themselves is more reliable and appropriate. Such warping effectively combines image-based visual features with geometry-based calibration and focal length, improving the smoothness and stability of zooming during camera switches.

[0024] Thus, the techniques described herein can improve image capture devices with multi-camera systems by reducing and / or eliminating visual discrepancies in images and / or videos during camera transitions, thereby improving their actual and / or perceived quality. Improving the actual and / or perceived quality of photos or videos can provide user experience benefits. These techniques are flexible and can be applied to a wide variety of videos in both indoor and outdoor settings.

[0025] In the following, the term "homography" is used to refer to an implementation of a warping transformation, and terms such as "warped" and "warping" may be used in the context of applying a warping transformation.

[0026] Smooth continuous zoom FIG. 1 illustrates binocular disparity in a multi-camera system according to an exemplary embodiment. For illustrative purposes, in FIG. 1, both cameras t1 and t2 face an object (one in focus and one out of focus). In some situations, camera t2 may be physically located adjacent to camera t1 (e.g., to the right, the other to the left, etc.). For example, the camera positions may be designed to mimic human left-eye / right-eye vision. In general, focused scene objects with the same depth (i.e., distance to the camera), such as focused object 110, can be warped almost perfectly from one camera t1 to another camera t2. For example, the focused object 110 of camera t1 warps to the focused object 110A of camera t2 without any mismatch. Planar objects with planes perpendicular to the camera's line of sight may exhibit such characteristics. Zooming in and / or out triggers a camera switch (e.g., between wide-angle and ultra-wide-angle, wide-angle and telephoto, etc.), leading to a change in FOV and a view discrepancy known as binocular disparity. For out-of-focus objects, a disparity (jump) between the images of the two cameras is perceptible. For example, a far object 105 and a near object 115 are out of focus in camera t1. Therefore, when the cameras are switched, the far object 105 in camera t1 is mapped to a far object 105A in camera t2, which is displaced from its actual position 105B. Similarly, the near object 115 in camera t1 is mapped to a near object 115A in camera t2, which is displaced from its actual position 115B.

[0027] As described herein, warping transformations can be applied to reduce binocular disparity by warping images from one camera to another. Focused scene objects (e.g., focused object 110) that have the same depth can be warped from one camera to another without perceptible parallax.

[0028] Warping inconsistencies may occur for out-of-focus objects (e.g., distant object 105, near object 115, etc.), or warping distortions may occur for planes across multiple depths. Inconsistencies such as those for out-of-focus objects depend on the depth and baseline of the camera. For example, an in-focus plane object with a plane tilted toward the camera's line of sight may have some "rotational" type inconsistencies.

[0029] FIG. 2 is a flowchart of a workflow 200 for image-based computation of a warping transformation, according to an example embodiment.

[0030] In block 210, the workflow includes acquiring frame-based data from the first and second cameras. A frame can be considered a unit of data processing. It contains input data required by image-based continuous zoom, including the image, pre-crop of the image, camera calibration with interior and exterior, auto focus distance, and other metadata.

[0031] At block 220, the workflow includes performing visual feature detection and matching to determine visual correspondence. Various visual feature detectors and descriptors may be used, such as SIFT (Scale-Invariant Feature Transform), SURF (Speeded Up Robust Features), FREAK (Fast REtinA Keypoint), or FAST (Features from Accelerated Segment Test). Additional and / or alternative visual features can be used in the pipeline, as long as the algorithm achieves sufficient quality of visual correspondence.

[0032] Image feature matching may be performed in two steps, such as feature extraction and feature matching. Various feature extraction and matching approaches may be used. For purposes of this specification, existing feature matching methods may be used, such as ArCore features or ILK features. The term "ILK" as used herein generally refers to a reverse search version of the Lucas-Kanade algorithm for optical flow estimation.

[0033] In block 230, the workflow includes relating, for each frame, visual feature matching with a two-view bundle, camera calibration, and automatic focal length. The camera calibration can be used to estimate the depth of points observed as feature matches. As described herein, due to OIS / VCM systems and / or thermal effects, geometry metadata such as camera calibration and focal length may be updated for each frame with significant errors. Therefore, manipulation of specific calibration parameters may be performed to ensure that most points have their depth values ​​close to the geometry focal length.

[0034] Some embodiments may enable smooth transitions between cameras (e.g., wide-angle and telephoto cameras), but may result in FOV jitter issues for the warping source camera. For example, visual features from a downsized image (e.g., 320x240) may not correspond to the same landmarks from frame to frame. Also, for example, due to OIS / VCM updates, camera calibration is updated every frame, and the scene depth of focus is calculated and / or corrected every frame from valid (e.g., inlier) feature matches between the dual cameras (e.g., wide-angle and telephoto). In some embodiments, in the event that the two cameras have different FOVs, the number of inlier visual feature matches may be limited by the smaller FOV (e.g., telephoto), potentially wasting visual information from the margins of the larger FOV (e.g., wide-angle) camera.

[0035] Some approaches to reducing such jitter may include damping control of scene focal length changes, and image-based ContiZoom may be triggered during the zoom process. To further address the jitter issue and utilize the same degree of visual information as that provided by larger FOV cameras, visual feature matching may be performed temporally between adjacent frames t−1 and t for the purpose of temporal consistency, so that each frame takes into account the geometry metadata of the previous frame when determining the warping mesh.

[0036] At block 240, the workflow includes performing bundle adjustment to obtain an optimized set of camera calibration, focal length, and other involved parameters based on the two-view bundle relationship. For example, image-based visual information can be effectively used to correct geometry metadata from upstream modules so that they are more compatible with images displayed as continuous zoom previews.

[0037] In the event that bundle adjustment is performed frame by frame, an optimized (e.g., minimizes pixel error reprojection) solution for each frame of camera calibration and focal length may be determined independently. In some embodiments, misalignment may exist across frames, resulting in a jittery preview when warped frames are displayed in sequential playback.

[0038] In such embodiments, the geometry bundle may be constructed over the entire window of the frame, and camera calibration and other metadata optimizations do not necessarily result in smooth transitions under the warping transformation. Therefore, instead of applying the homography of the most recently optimized data to warp the image, a decay process given by the following equation is introduced for gradual and smooth transitions in the warping transformation:

number

[0039] In some embodiments, prior to extracting image-based visual features, the homography G is determined based on geometry metadata (e.g., camera calibration and focal length) from upstream modules, which may include errors from OIS / VCM updates, thermal effects, and other sources, as discussed above. This can be corrected using image-based data as follows:

[0040] Generally, two sets of camera calibration models are available: one calibration model is updated with OIS / VCM corrections that directly correspond to visual features from the image, and the other calibration model is kept neutral and used as a smooth initial value for further geometry optimization.

[0041] The previously computed geometry-based homography G may be used to perform the aforementioned image-based process to correct any remaining errors, followed by performing a coarse-grained pre-warping of the image features and associated calibration. Such an approach effectively combines geometric and visual information to solve technical problems.

[0042] In block 250, the workflow includes determining a pre-warping transformation of the first camera's image based on a bundle adjustment so that the warped image has only a small disparity with respect to the second camera's image.

[0043] Then, in block 260, the workflow includes precisely warping the image of the first camera and modifying the pre-warping transformation based on image features to further reduce small disparities from the pre-warping transformation.

[0044] Some embodiments include optimizing the detection of one or more visual features and the generation of visual correspondences by performing asynchronous multi-threaded processing that includes receiving one or more images and associated metadata as input and sending visual feature matching and associated metadata as output. For example, image-based continuous zooming may include computationally intensive steps such as visual feature detection and matching. To enable this solution to occur in real time (e.g., at least 30 frames per second (fps)) on consumer-grade devices, the computationally intensive steps may be processed asynchronously by a specific thread that receives images and associated metadata as input and sends visual feature matching and associated metadata as output.

[0045] As used herein, the term "sparse features" generally refers to detected features that are sparsely scattered throughout an image. Feature points can be detected when a pixel and its neighborhood meets a detection threshold. This can include, for example, ArCore features, portrait mode features, and AutoCal features.

[0046] The term "dense features" as used herein generally refers to detected features that may cover the entire image, where feature points may be detected based on predefined image patches and matches may be found for each patch. In some embodiments, ILK may be used to detect dense features. For example, the ILK algorithm may be used to extract "dense" feature points and matches from an image.

[0047] Dense or sparse features may generally have different designs in the ContiZoom pipeline, as described in more detail below.

[0048] Recalibration (Sparse Feature Flow) In contrast to the planar target features typically used during factory calibration, sparse feature calibration uses natural features to recalibrate the geometry information received from the camera sensor. For example, image features are used to update the geometry metadata, and existing geometry-based calculations are leveraged to calculate a revised homography.

[0049] 3 is an exemplary sparse feature workflow 300 for smooth continuous zoom in a multi-camera system, according to an exemplary embodiment. The general flow with sparse features is shown.

[0050] Block 305 receives input images including dual 320x240 images. The input images are from two target cameras. In some examples, the image resolution may range up to 640x480. The quality of feature matching further depends on the quality of the texture of the scene.

[0051] In block 310, sparse feature detection may be performed, as described above. In some embodiments, ArCore features may be used. In some embodiments, FAST features may be used. Generally, scale-invariant features are not required, as the dual images can be rescaled to provide a relatively accurate focal length.

[0052] In block 315, natural feature calibration may be performed. This process recalibrates the camera parameters 325 from a factory calibration, for example, provided by the camera provider 320. In some embodiments, DualCameraCalibrator or AutoCal (each based on a FAST feature detector) may be used. Also, an optical flow-based detector, such as ILK, may be used. However, natural feature calibration generally differs from factory calibration because it is not derived from a planar object. Therefore, a bundle adjustment (BA)-based approach (e.g., block 240 of FIG. 2) may be preferred to perform natural feature calibration.

[0053] In some embodiments, natural feature calibration may not optimize all parameters, but instead may focus on the "principal point" and "external rotation" of wide-angle and telephoto. The following table (Table 1) summarizes parameters that may need to be optimized. Table 1 is for illustrative purposes only and may vary from device to device and may be based on the type and / or characteristics of the cameras involved in the transition process. [Table 1]

[0054] In block 330, delta camera metadata may be obtained. While CPI-based calibration is with respect to active array coordinates, image-based algorithms require calibration with respect to image coordinates. Therefore, a transformation may be determined between the active array and the image. After the camera metadata is corrected, the difference between the factory calibration and the corrected metadata may be stored as delta metadata and may be kept separate from the CPI calibration metadata. In some embodiments, the delta may be a constant offset during transitions from one camera to another.

[0055] Image-based correction may use the same structure of camera metadata, which is the delta between the CPI output and the recalibrated camera metadata. In some embodiments, a delta focal length (e.g., depth) may be used.

[0056] In block 335, features inside a region of interest (ROI) may be detected. ROI, as used herein, is a sub-region within an image frame that is deemed important to the user and, in camera applications, is used as a pilot region for many features, such as autofocus, providing the focal length used for geometry-based methods. In some embodiments, the ROI may be obtained as an ROI rectangle 340 from an algorithm, such as a face detection algorithm, a saliency detection algorithm, or the like. In general, the ROI may be processed differently for sparse and dense features. The features inside the ROI may be based on the sparse features detected in block 310.

[0057] In block 345, the median depth in the ROI is determined, for example, the median depth of the (inlier) image feature points extracted by the sparse / dense feature method as described above may be determined.

[0058] Temporal feature matching and tracking In some embodiments, temporal feature matching and tracking may be performed. Generally speaking, the same landmark or ROI may appear in multiple frames (e.g., three or more frames), resulting in feature tracking. In some embodiments, temporal feature tracking may only be applied to larger FOV (e.g., wide-angle) cameras that are sufficient for the temporal consistency of the ContiZoom mesh.

[0059] FIG. 4A illustrates temporal feature matching and tracking according to an exemplary embodiment. Referring to FIG. 4A, multiple consecutive frames of a wide-angle camera 405 and a telephoto camera 410 are shown. For the wide-angle camera 405, two consecutive frames are shown: a first frame 415 at time t-1 and a second frame 420 at time t. For the telephoto camera 410, two consecutive frames are shown: a third frame 425 at time t-1 and a fourth frame 430 at time t. Intra-frame feature matching is shown, where a first feature A in the first frame 415 at time t-1 is matched to a corresponding feature A' in the third frame 425. Similarly, intra-frame feature matching is shown, where a second feature B in the second frame 420 at time t is matched to a corresponding feature B' in the fourth frame 430. Temporal feature matching and tracking is shown, where a first feature A in a first frame 415 at time t−1 is matched to a second feature B in a second frame 420 at time t. For illustrative purposes, temporal feature tracking is shown for a wide-angle camera 405. In general, it may be desirable to perform temporal feature tracking with a camera that has a large FOV to capture relevant feature tracking.

[0060] 4B shows an example image of temporal feature matching and tracking according to an example embodiment. Referring to FIG. 4B, two images are shown. The first image 435 shows feature matching without focal length attenuation. The second image 440 shows temporal feature matching to reduce jitter.

[0061] In general, temporal feature matching and tracking may involve two tasks to improve ContiZoom quality: 1) determining temporal feature tracking information from images, and 2) applying temporal feature tracking to existing pipelines.

[0062] In some embodiments, the first task may include providing interface functions for feature extraction and feature matching, respectively. As described with respect to FIG. 4A, intra-frame feature matching may be performed on dual images from the lead and follower cameras (e.g., from the wide-angle camera 405 to the telephoto camera 410) to build intra-frame feature matching. In some embodiments, temporal feature tracking may be performed by refactoring the interface functions and caching features extracted from previous frames by enabling feature matching between adjacent frames t−1 and t.

[0063] In some embodiments, the second task may include constructing an index manager for handling visual feature point indices and matching between multiple images. This index manager facilitates the use of temporal feature tracking in addition to intra-frame feature matching. For example, the index manager manages visual feature indices, including feature point indices and feature match indices, and their correspondences with each other. In some embodiments, the index manager may support querying of the feature point index from the feature match index and querying of the feature match index from the first and second feature point indices of the match.

[0064] Some embodiments may include one or more mappings. For example, a first mapping from the indices of feature points from a first image to the indices of matching pairs that include the feature points. As another example, a second mapping from the indices of feature points from a second image to the indices of matching pairs that include the feature points. Also, for example, a vector of feature matches may be determined. For example, each feature match may include two feature points each from the first image and the second image.

[0065] Experimental evidence shows that jitter in wide-angle cameras as a lead can be primarily caused by jitter in the scene focal length estimated per frame. Therefore, temporal feature tracking can be used to estimate the scene focal length at frame t. Utilizing feature tracking across multiple frames allows for improved quality under temporal consistency.

[0066] In some situations, when a user performs zoom-in and / or zoom-out operations on a camera, it may be reasonable to assume that the user is likely not having large movements (e.g., panning, running, walking, or rapid changes in salient objects / ROIs). In events where the user has large movements, small jitter in the zoom or FOV transition is less noticeable. In events where the user does not have large movements, changes in scene focal length between adjacent frames t-1 and t need to be managed to avoid perceptible jitter. One approach to achieving this is to keep the scene focal length unchanged if frames t-1 and t have a sufficient number of inlier matches.

[0067] FIG. 4C illustrates an exemplary application 400 of temporal feature matching and tracking according to an exemplary embodiment. FIG. 4C illustrates multiple consecutive frames of a wide-angle camera 405 and a telephoto camera 410. In general, each frame may estimate scene depth from a set of feature matching between the two cameras (e.g., between the wide-angle camera 405 and the telephoto camera 410) for this same frame. In some embodiments, inlier features may be selected to represent the scene for depth estimation. For illustrative purposes, the presented example includes four (4) scene focal lengths, namely, d1, d2, d3, and d4. Temporal feature tracking between frame t-1 and frame t may be performed as follows:

[0068] In the event that the selected inlier feature has a temporally matched counterpart, the scene focal length d2 may be set to be the same as d1. For example, feature A in frame t1 445 has a corresponding feature A' in frame t1' 465. Also, feature A in frame t1 445 has a temporally matched counterpart feature B in frame t2 450. Therefore, the scene focal length d2 may be set to be the same as d1.

[0069] In the event that a selected inlier feature does not have a temporal match, but other features with a similar depth to the initially selected inlier have a temporal match and an intra-frame feature match, the scene focal length d3 may be set to be the same as d1. For example, feature E in frame t3 455 does not have a temporal match. However, feature D in frame t3 455 is at a similar depth to feature E in frame t3 455. Feature D in frame t3 455 also has a temporally matched counterpart feature C in frame t2 450. Furthermore, feature D in frame t3 455 has a corresponding feature D' in frame t3' 475. Therefore, the scene focal length d3 may be set to be the same as d1.

[0070] In the event that the selected inlier feature and its siblings at close depth do not have both a temporal match and an intra-frame feature match, d4 is recalculated. For example, feature H in frame t4 460 has a corresponding feature H' in frame t4' 480. However, feature H in frame t4 460 does not temporally match any other features. Features G and I in frame t4 460 appear as features with similar depths to feature F. Feature G temporally matches feature F in frame t3 455, but feature G does not have an intra-frame feature match. Feature I also does not have an intra-frame feature match or a temporal match. Therefore, d4 is recalculated.

[0071] In the aforementioned approaches, scene depth (i.e., focal length) estimation is based on intra-frame (i.e., inter-camera) feature matching on a frame-by-frame basis, and temporal tracking is typically used as a post-mortem to determine whether the scene focal length needs to be changed in subsequent frames. However, this approach does not effectively use information from temporal feature tracking.

[0072] Therefore, alternative approaches to reducing and / or eliminating jitter by adjusting the scene focal length may include direct use of temporal feature tracking information, especially if such information is available with sufficiently high quality. Factors that may determine the quality of the temporal feature tracking information may include one or more of: (1) a sufficient number of time matches between adjacent frames at time t-1 and time t, (2) a sufficient number of time matches that allow for time tracking to be constructed across multiple frames in the absence of large panning and / or rotational motion, or (3) available power and latency when running on a mobile device.

[0073] In one approach, temporal feature tracking, along with gyroscope measurements, can be used to triangulate and select "up-to-scale" 3D points as good inlier landmarks. Dual-camera observations of these inlier landmarks can then be used to estimate scene depth per frame.

[0074] In another approach, temporal feature tracking, along with gyroscope and accelerometer measurements, may be used to directly estimate device pose and 3D points as inlier landmarks. Per-frame scene depth may then be based on these inlier landmarks. Given the power and latency aspects of device pose estimation, this approach may be more applicable to offline processing.

[0075] 3, the delta camera metadata from block 330 and the median depth from block 345 are used to update the geometry homography (determine an updated warping transformation) based on image features in block 350. For example, the matched feature points may be used to correct for inaccurate geometry metadata. This may include two approaches: a recalibration-based sparse feature flow approach and an image homography-based dense feature flow approach.

[0076] The homography H is a 3x3 matrix that maps pixels in a plane in a first coordinate system of a first camera to corresponding pixels in the same plane in a second coordinate system of a second camera. The homography can be decomposed as follows:

number

[0077] where K1 and K2 are intrinsic matrices containing the focal length and principal point corresponding to the first and second cameras, respectively, and [R 3×3 |t 3×1 ] is the exterior that can transform a point in the first coordinate system of the first camera into the second coordinate system of the second camera, and d and n are the dimensions of the plane X that point to n. T Define the focused plane of this homography in the coordinates of the first camera so that x + d = 0. In general, n = [0, 0, -1] T is used.

[0078] Given the camera calibration and the distance of the object in focus, there are at least two ways to determine the homography. For example, in a first approach, the homography can be determined by inputting these parameters into an equation such as that shown in Equation 2. Also, for example, in a second approach, the homography can be determined from a set of pixel pairs (e.g., at least four sample point pairs). For example, the homography matrix transforms a plane (at a particular depth) on a telephoto camera to a corresponding plane on a wide-angle camera. In some embodiments, for a pair of cameras (e.g., telephoto and wide-angle camera models), a "four-point" approach may be used to calculate the homography matrix between the telephoto and wide-angle cameras. Inputs may include the telephoto and wide-angle camera models (interior and exterior) and the distance of the target plane (object distance) to the telephoto camera.

[0079] The "four-point" approach is 望遠 This may involve arbitrarily selecting four two-dimensional (2D) points on the telephoto camera, denoted as P. The 2D points may then be deprojected by using the camera internal parameters as 3D rays, denoted as rays. The rays may then intersect with a given plane at a distance Obj_dist away from the telephoto camera. This produces four three-dimensional (3D) points in real space. The 3D points are deprojected as P 広角 The homography matrix may then be determined as follows:

number

[0080] In the formula, [P|1] represents the homogeneous coordinate of P.

[0081] A second approach to determining the homography from a set of pixel pairs may involve in-camera distortion in the estimation process. If image-based visual information is not available, the camera calibration comes from the CPI library and the focal length comes from the autofocus process.

[0082] In block 355 , the protrusion handler performs protrusion processing based on the homography from block 350 .

[0083] In block 360, a mesh transformation function is applied.

[0084] Image-based homography (dense feature flow) The image-based approach directly computes an updated homography based on matched image features and combines the image-based homography with the geometry-based homography. Since the goal is to warp two images (from two cameras) together, the image features can be used directly to compute the warping homography without geometry camera metadata.

[0085] 5 is an example dense feature workflow 500 for smooth continuous zoom in a multi-camera system, according to an example embodiment. The general flow with dense features is shown.

[0086] At block 505, input images including dual 320x240 images are received. The input images are from two target cameras. In some examples, the image resolution may be 640x480 or higher, depending on the computational power of the computing device and latency requirements for various use cases. The quality of the feature matching may further depend on the quality of the texture of the scene.

[0087] At block 510, an aligned ROI region is calculated. Some embodiments may include cropping the original image frame to the ROI region. The ROI rectangle 525 may be obtained (e.g., as described above with respect to the reference ROI rectangle 340). Some embodiments include aligning the two ROI regions with the calculated geometry-based homography so that dense feature detection may have a better initial alignment. To conserve computational resources, in some embodiments, translation components may be extracted from the geometry-based homography, and these translation components may be applied to the ROI. In some embodiments, this translation may be performed by cropping.

[0088] In block 520, dense feature detection may be performed. Such feature detection / matching has been described previously, and the ILK algorithm may be used. This process recalibrates the camera parameters 530 from a factory calibration, for example, provided by the camera provider 535. In some embodiments, DualCameraCalibrator or AutoCal (each based on the FAST feature detector) may be used. Alternatively, an optical flow-based detector, such as ILK, may be used. However, natural feature calibration generally differs from factory calibration because it is not derived from planar objects. Therefore, a bundle adjustment (BA)-based approach (e.g., block 240 of FIG. 2) may be preferred to perform natural feature calibration.

[0089] In block 540, a separate delta homography may be determined. In some embodiments, this may be based on camera parameters 530 from a factory calibration, for example, provided by the camera provider 535. In some embodiments, dense features from the foreground may be separated from features in the background. For example, feature detection may focus on features inside the ROI, but background features may also be used. In some embodiments, a translation-only homography may be applied to ensure that the homography plane is perpendicular to the camera's z-axis. Also, for example, two-cluster k-means may be used, where the features are disparity and the feature location may be chosen to be the center of the ROI.

[0090] Similar k-means processes may be used for their counterparts in sparse feature flows, but sparse features may not have enough feature points to apply these approaches. Therefore, a non-optimal solution based on determining the median depth may be used for sparse feature flows (e.g., in block 345).

[0091] In general, since the original ROI has already been shifted using the geometry-based homography, the foreground homography is a delta homography on top of the geometry-based homography.

[0092] In block 545, an updated warping transform (or combined homography) may be determined as the combination of the image-based delta homography and the geometry-based homography. In some embodiments, the updated warping transform is a concatenation of two homography mappings (the image-based homography applied after the geometry-based homography) with appropriate coordinate transformations.

[0093] Both dense and sparse feature flows may utilize existing components. For example, the term "camera provider" (e.g., camera provider 320, camera provider 535) refers to a module that reads a factory calibration file and generates camera parameters corresponding to each OIS / VCM metadata using CPI.

[0094] The term "CPI camera parameters" (eg, camera parameters 325, camera parameters 530) refers to camera internal and / or external parameters calculated by CPI.

[0095] In block 550 , a protrusion handler performs protrusion processing based on the combined homography from block 545 .

[0096] In block 555, a mesh transformation function is applied.

[0097] Sparse features may generally be more accurate and may provide precise (xy) image coordinates for feature points. However, sparse features may be relatively slow, and the number of features detected may depend on the complexity of the scene. Therefore, if there are not enough feature detections, the performance of the calibration optimization may be adversely affected.

[0098] Dense features may be less accurate because these "feature points" are essentially image patches, but dense feature detection is generally faster and less dependent on scene complexity.

[0099] In some embodiments, a combination of sparse and dense features may be used, for example, sparse features may be detected first, and in the event that the number of detected features is below a threshold, dense feature detection may be performed.

[0100] Also, for example, sparse features may be used to determine more accurate depth, and a better homography may be determined as an initial guess for dense feature matching.

[0101] As another example, dense features may be used to determine area patches of foreground objects, and sparse features may be used to detect precise feature points on the foreground area patches.

[0102] For both sparse and dense features, delta data is stored. For sparse features, the delta data is camera metadata, and for dense features, the delta data is homography. In general, during a zoom operation, both cameras may not be available.

[0103] FIG. 6 illustrates exemplary processing of delta data during a camera transition, according to an exemplary embodiment. The example in FIG. 6 is based on a transition between a wide-angle camera and a telephoto camera. However, a similar approach may apply to transitions between different pairs of cameras. The exemplary zoom ratios, states, number of states, transition points, etc. are for illustrative purposes only. Values ​​such as these may vary depending on the type of device, the type of camera, the distance of the camera from the object, etc. For example, in FIG. 6, "4.2x" is used to mark the zoom ratio at which a transition between a wide-angle camera and a telephoto camera can be made. However, this value may vary depending on the type of device, the cameras involved in the transition, the distance of the camera from the object, etc. Also, for example, the zoom ratios used in the example of FIG. 6 (e.g., 2x, 4x, 4.2x, 4.4x) are example values ​​and may vary depending on device and / or system configuration.

[0104] For example, shown is a transition from wide to telephoto 605 and a reverse transition from telephoto to wide 610. Legend 615 indicates the lead and follower cameras.

[0105] State 1 (initially open at 1.0x) corresponds to when the camera is first activated. The initial camera may be a wide-angle camera, and the homography is the identity operation. Geometry data from the wide-angle camera's CPI may be updated, as it does not affect the homography.

[0106] State 2 (from 2x scale) corresponds to when the homography gradually changes from the identity operation to the target homography. However, because dual cameras are not available in State 2, this homography is geometry-based and does not take image features into account. No delta data is present. To avoid causing any visual distortion during the transition between State 1 and State 2, in some embodiments, updates from the CPI may be attenuated.

[0107] State 3 (4.0x scale to 4.2x scale) corresponds to when two cameras are active simultaneously. In this case, the wide-angle camera is the primary or lead camera, and the telephoto camera is the secondary or follower camera. The image-based approach described above may be used to calculate the delta data, and the updated metadata may be used to calculate the homography.

[0108] State 4 (4.2x scale to 4.4x scale) corresponds to after a switch from the wide-angle camera to the telephoto camera has occurred. In this case, the telephoto camera is the primary or lead camera, but the wide-angle camera is still active. No warping is applied, so the homography is the identity operation. The last calculated delta is stored. Then, the geometry camera metadata for the telephoto and wide-angle cameras is updated.

[0109] In state 5 (above 4.4x), the telephoto camera is the primary or lead camera, and the wide-angle camera is inactive or closed. Other operations remain the same as in state 4.

[0110] State 6 (back to wide angle) is similar to state 2, but now delta data is present. Therefore, the delta data is stored and the homography is calculated with the additional delta.

[0111] State 7 (back to wide angle) is similar to state 1, but now delta data is present, so it is stored.

[0112] During successive transitions, the process may be repeated between states 3 through 7.

[0113] A transition zone is a defined zone during which a smooth transition occurs, where two cameras are simultaneously active within a certain range of zoom scale, so that metadata (e.g., from the OIS / VCM) can be streamed to both cameras simultaneously. For the image-based approach described herein, it may be preferable to set this transition period to be as large as possible to reduce abrupt changes between camera metadata and / or between image-based and geometry-based results.

[0114] In some embodiments, hardware limitations may make it impractical to always stream with two cameras and / or to extend the transition zone as far as might be optimal. In such embodiments, temporary dual streaming may be used. For example, temporary dual streaming means that two cameras are temporarily active simultaneously based on a timer rather than based on a zoom scale. For example, after the camera application is opened, a timer may be set for 10 seconds, and the two cameras may be simultaneously active every 10 seconds. Smooth continuous zooming may be performed periodically based on such a timer.

[0115] FIG. 7 is a table 700 illustrating various cases for switching between a telephoto camera and a wide-angle camera, according to an example embodiment. Column C1 lists the states described with reference to FIG. 6 , column C2 lists the status related to the wide-angle camera's geometry metadata, column C3 lists the status related to the telephoto camera's geometry metadata, column C4 lists the status related to the delta data, and column C5 lists the status related to the homography. Each row, rows R1-R7, provides a status for each state, states 1-7. Table 700 summarizes the information provided with reference to FIG. 6 . For example, row R2 indicates that for state 2, a decay update is applied to the wide-angle camera's geometry metadata, a canonical mapping is used for the telephoto camera's geometry metadata, there is no delta data, and the homography is a combination of a ratio delta and a geometry homography. The other rows present similar information for their respective states.

[0116] The term "delta data" refers to the result of the image-based solution described above, where the delta is the camera geometry metadata in the sparse case and the delta homography in the dense case.

[0117] The status "update" typically indicates a near-real-time update by the OIS / VCM. The status "hold" indicates the status is the same as the previous state. The term "decaying update" refers to a gradual update with a decaying ratio between the data from the previous frame and the data from the current frame. The term "geo" refers to a geometry-based homography (not including image features). The term "delta + geo" refers to a combined homography of the image-based solution and the geometry-based solution. The term "ratio homography" indicates that the strength of the homography may depend on the zoom scale (e.g., for state 2 of row R2), with the homography strength being identity at 2.0x and the homography being full strength at 4.2x. Other scales between 2.0x and 4.2x may be determined as an interpolation between the identity operation and the full-strength homography.

[0118] Delta damping is generally an operation that smooths sharp changes in geometry data, such as sudden OIS / VCM changes. The damping ratio may be based on the change in zoom scale between two consecutive frames. Similar damping may be applied to deltas.

[0119] 8A shows an example geometric relationship 800A for each pair of matched pixels, according to an example embodiment. A first plane 805 in the XYZ plane is shown to include point w with coordinates referenced to an origin O. A second plane 810 corresponds to the first plane 805 in the X'Y'Z' plane. The first coordinate system representing the XYZ plane can be mapped to a second coordinate system representing the X'Y'Z' plane with coordinates referenced to an origin O' by mapping X'=RX+T, where R is rotation and T is translation. For example, point w in the first plane 805 is mapped to point w' in the second plane 810. In some embodiments, a two-view geometric relationship can be established for each pair of matched pixels with the camera interior and exterior, triangulation points from the visual matching, and autofocus distance as the initial scene depth.

[0120] FIG. 8B shows a workflow 800B for determining the geometric relationship for each pair of matched pixels according to an example embodiment. In block 815, a point w (e.g., within the first plane 805) is selected. In block 820, an interior for a first camera (e.g., the wide-angle camera shown in FIGS. 5 and 6) is determined. In block 825, the 2D point w may be deprojected as a 3D ray, denoted as a ray, by using the interior for the first camera. In block 830, depth data may be received. In block 835, a 3D point for the first camera is determined based on the ray and depth data. In block 840, the exterior from the first camera is applied to a second camera (e.g., the telephoto camera shown in FIGS. 5 and 6). In block 845, a 3D point for the second camera (corresponding to the 3D point determined in block 835) is determined. In block 850, an interior for the second camera is determined. In block 855, a reprojection of point w is determined based on the interior with respect to the second camera. Based on the actual position of point w′ of the second camera obtained in block 860 and the reprojection of point w, in block 865, one or more reprojection errors are determined.

[0121] Thus, a vision-based correction of geometry data in individual frames is provided. Workflow 800B enables minimizing the vision correspondence reprojection error. Based on workflow 800B, the geometry-based homography can be re-estimated with partially corrected camera calibration and focused object distance to achieve smoothness across frames.

[0122] 9 shows an example workflow 900 for smooth continuous zoom in a multi-camera system, according to an example embodiment. A continuous zoom frame 905 may include a calibration name file 910, an OIS / VCM pair 915 from two cameras, and a warping grid configuration 920.

[0123] The algorithms described herein with reference to at least Figures 1-8 may be managed by a continuous zoom manager 925. In some embodiments, the continuous zoom manager 925 may include a data trimmer 930, a calibration provider 945, and a homography provider 955. The data trimmer 930 may perform data validation 935 and data dump 940. The calibration provider 945 may provide CPI parameters 950 obtained from a factory calibration file 980.

[0124] The homography provider 955 may determine a geometry-based homography 960 and an image-based homography 965, as described herein. The homography provider 955 may then determine homography compensation 970 and projection processing 975.

[0125] Legend 985 indicates the various classes of components involved, such as container classes, member functions, member variables, and functionality classes.

[0126] Data Network Example 10 illustrates a distributed computing architecture 1000, according to an example embodiment. The distributed computing architecture 1000 includes server devices 1008, 1010 configured to communicate with programmable devices 1004a, 1004b, 1004c, 1004d, and 1004e via a network 1006. The network 1006 may correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WWAN, a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 1006 may also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.

[0127] While FIG. 10 shows only five programmable devices, the distributed application architecture can serve tens, hundreds, or thousands of programmable devices. Furthermore, programmable devices 1004a, 1004b, 1004c, 1004d, and 1004e (or any additional programmable devices) may be any type of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mountable device (HMD), a network terminal, a mobile computing device, etc. In some examples, such as shown by programmable devices 1004a, 1004b, 1004c, and 1004e, the programmable devices may be directly connected to network 1006. In other examples, such as shown by programmable device 1004d, the programmable devices may be indirectly connected to network 1006 through an associated computing device, such as programmable device 1004c. In this example, programmable device 1004c can serve as an associated computing device for passing electronic communications between programmable device 1004d and network 1006. In other examples, such as shown by programmable device 1004e, the computing device may be part of and / or within a vehicle, such as a car, truck, bus, boat or watercraft, airplane, etc. In other examples not shown in Figure 10, the programmable device may be connected both directly and indirectly to the network 1006.

[0128] The server devices 1008, 1010 can be configured to perform one or more services as requested by the programmable devices 1004a-1004e. For example, the server devices 1008 and / or 1010 can provide content to the programmable devices 1004a-1004e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content can include compressed and / or uncompressed content. The content may be encrypted and / or decrypted. Other types of content are possible as well.

[0129] As another example, server devices 1008 and / or 1010 may provide programmable devices 1004a-1004e with access to software for database, search, calculation, graphics, audio, video, World Wide Web / Internet utilization, and / or other functions. Many other examples of server devices are possible as well.

[0130] Computing Device Architecture 11 is a block diagram of an exemplary computing device 1100, according to an exemplary embodiment. In particular, the computing device 1100 shown in FIG. 11 may be configured to perform at least one function of and / or related to method 1200.

[0131] The computing device 1100 may include a user interface module 1101, a network communication module 1102, one or more processors 1103, data storage 1104, one or more cameras 1118, one or more sensors 1120, and a power system 1122, all of which may be linked together via a system bus, network, or other connection mechanism 1105.

[0132] The user interface module 1101 may be operable to transmit data to and / or receive data from external user input / output devices. For example, the user interface module 1101 may be configured to transmit data to and / or receive data from user input devices such as a touchscreen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, and / or other similar devices. The user interface module 1101 may also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices, either now known or later developed. The user interface module 1101 may also be configured to generate audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, earphones, and / or other similar devices. User interface module 1101 may further comprise one or more haptic devices capable of generating haptic outputs, such as vibrations and / or other outputs detectable by touch and / or physical contact with computing device 1100. In some examples, user interface module 1101 may be used to provide a graphical user interface (GUI) for utilizing computing device 1100.

[0133] The network communication module 1102 may include one or more devices providing one or more wireless interfaces 1107 and / or one or more wired interfaces 1108 configurable to communicate over a network. The one or more wireless interfaces 1107 may include one or more wireless transmitters, receivers, and / or transceivers, such as a Bluetooth® transceiver, a Zigbee® transceiver, a Wi-Fi™ transceiver, a WiMAX™ transceiver, an LTE™ transceiver, and / or other types of wireless transceivers configurable to communicate over a wireless network. The one or more wired interfaces 1108 may include one or more wired transmitters, receivers, and / or transceivers, such as an Ethernet transceiver, a Universal Serial Bus (USB) transceiver, or similar transceiver configurable to communicate over a twisted pair wire, coaxial cable, fiber optic link, or similar physical connection to a wired network.

[0134] In some examples, the network communication module 1102 can be configured to provide reliable, secure, and / or authenticated communications. For each communication described herein, information to facilitate reliable communications (e.g., guaranteed message delivery) may be provided, perhaps as part of the message header and / or footer (e.g., packet / message ordering information, encapsulation header and / or footer, size / time information, and transmission verification information such as a cyclic redundancy check (CRC) and / or parity check value). Communications may be protected (e.g., encoded or encrypted) and / or decrypted / decoded using one or more cryptographic protocols and / or algorithms, such as, but not limited to, the Data Encryption Standard (DES), the Advanced Encryption Standard (AES), the Rivest-Shamir-Adelman (RSA) algorithm, the Diffie-Hellman algorithm, a secure socket protocol such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or the Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms may be used to protect (and decrypt / decrypt) communications similarly or in addition to those described herein.

[0135] The one or more processors 1103 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application-specific integrated circuits, etc.) The one or more processors 1103 may be configured to execute computer-readable instructions 1106 contained in data storage 1104 and / or other instructions described herein.

[0136] Data storage 1104 may include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of the one or more processors 1103. The one or more computer-readable storage media can include volatile and / or non-volatile storage components, such as optical, magnetic, organic, or other memory or disk storage, which may be integrated, in whole or in part, with at least one of the one or more processors 1103. In some embodiments, data storage 1104 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other embodiments, data storage 1104 may be implemented using two or more physical devices.

[0137] The data storage 1104 can include computer-readable instructions 1106 and possibly additional data. In some examples, the data storage 1104 can include storage necessary to perform at least some of the methods, scenarios, and techniques described herein and / or at least some of the functionality of the devices and networks described herein. In some examples, the data storage 1104 can include storage for a warping transform module 1112 (e.g., a module that computes geometry-based homographies, image-based homographies, etc.). In particular, in these examples, the computer-readable instructions 1106 can include instructions that, when executed by the one or more processors 1103, enable the computing device 1100 to provide some or all of the functionality of the warping transform module 1112.

[0138] In some examples, computing device 1100 may include one or more cameras 1118. Camera(s) 1118 may include one or more image capture devices, such as still and / or video cameras, equipped to capture light and record the captured light into one or more images. That is, camera(s) 1118 may generate image(s) of the captured light. The one or more images may be one or more still images and / or one or more images utilized in video capture. Camera(s) 1118 may capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or one or more other frequencies of light. Camera(s) 1118 may include wide-angle cameras, telephoto cameras, ultra-wide-angle cameras, etc. Also, for example, camera(s) 1118 may be forward-facing or rear-facing cameras with respect to computing device 1100.

[0139] In some examples, computing device 1100 may include one or more sensors 1120. Sensors 1120 may be configured to measure conditions within computing device 1100 and / or conditions in the environment of computing device 1100 and provide data regarding these conditions.For example, the sensors 1120 may be (i) sensors for acquiring data about the computing device 1100, such as, but not limited to, a thermometer for measuring the temperature of the computing device 1100, a battery sensor for measuring the power of one or more batteries of the power supply system 1122, and / or other sensors for measuring the state of the computing device 1100; (ii) identification sensors for identifying other objects and / or devices, such as, but not limited to, a radio frequency identification (RFID) reader, a proximity sensor, a one-dimensional barcode reader, a two-dimensional barcode (e.g., a quick response (QR) code) reader, and a laser tracker, which may be configured to read identifiers such as RFID tags, barcodes, QR codes, and / or other devices and / or objects configured to read and provide at least identification information; and (iii) tilt sensors, gyroscopes, accelerometers, and the like. The sensors 1120 may include one or more of the following: (i) sensors that measure the position and / or movement of the computing device 1100, such as, but not limited to, a gyroscope, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser displacement sensor, and a compass; (ii) environmental sensors that acquire data indicative of the environment of the computing device 1100, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor, a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a movement sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (iii) force sensors that measure one or more forces (e.g., inertial forces and / or G-forces) acting about the computing device 1100, such as, but not limited to, one or more sensors that measure force, torque, ground force, friction in one or more dimensions, and / or a zero moment point (ZMP) sensor that identifies the ZMP and / or the location of the ZMP. Many other examples of sensors 1120 are possible as well.

[0140] The power supply system 1122 may include one or more batteries 1124 and / or one or more external power interfaces 1126 for providing power to the computing device 1100. Each battery of the one or more batteries 1124, when electrically coupled to the computing device 1100, may serve as a source of stored power for the computing device 1100. The one or more batteries 1124 of the power supply system 1122 may be configured to be portable. Some or all of the one or more batteries 1124 may be easily removable from the computing device 1100. In other examples, some or all of the one or more batteries 1124 may be internal to the computing device 1100 and therefore may not be easily removable from the computing device 1100. Some or all of the one or more batteries 1124 may be rechargeable. For example, rechargeable batteries may be recharged via a wired connection between the battery and another power source, such as by one or more power sources external to the computing device 1100 and connected to the computing device 1100 via one or more external power interfaces. In other examples, some or all of the one or more batteries 1124 may be non-rechargeable batteries.

[0141] The one or more external power interfaces 1126 of the power system 1122 may include one or more wired power interfaces, such as a USB cable and / or a power cord, that enable a wired power connection to one or more power sources external to the computing device 1100. The one or more external power interfaces 1126 may include one or more wireless power interfaces, such as a Qi wireless charger, that enable a wireless power connection to the one or more external power sources, such as via a Qi wireless charger. Once a power connection to an external power source is established using the one or more external power interfaces 1126, the computing device 1100 may draw power from the external power source via the established power connection. In some examples, the power system 1122 may include associated sensors, such as a battery sensor or other type of power sensor associated with one or more batteries.

[0142] Exemplary Methods of Operation 12 illustrates a method 1200 according to an example embodiment. The method 1200 may include various blocks or steps. The blocks or steps may be performed individually or in combination. The blocks or steps may be performed in any order and / or sequentially or in parallel. Additionally, blocks or steps may be omitted or added to the method 1200.

[0143] The blocks of the method 1200 may be performed by various elements of the computing device 1100 as shown and described with reference to FIG.

[0144] Block 1210 includes displaying, by a display screen of the computing device, an initial preview of a scene captured by a first image capture device of the computing device, the first image capture device operating within a first focal length range.

[0145] Block 1220 includes detecting, by the computing device, a zoom operation that is predicted to cause the first image capture device to reach a limit of a first focal length range.

[0146] Block 1230 includes, in response to the detecting, activating a second image capture device of the computing device to capture a zoom preview of the scene, the second image capture device configured to operate within a second focal length range.

[0147] Block 1240 includes updating the geometry-based warping transformation based on a comparison of the respective image features from the initial preview and the zoomed preview.

[0148] Block 1250 includes aligning the zoomed preview with the initial preview by applying an updated warping transform, which reduces one or more visual artifacts caused by changes in field of view when transitioning from the initial preview to the zoomed preview.

[0149] Block 1260 includes displaying, by a display screen of the computing device, an aligned zoom preview of an image captured by the second image capture device while operating within the second focal length range.

[0150] In some embodiments, comparing the respective image features includes detecting one or more visual features in the initial preview and the zoomed preview. These and other embodiments also include generating a visual correspondence between the initial preview and the zoomed preview based on the one or more visual features.

[0151] Some embodiments include optimizing the detection of one or more visual features and the generation of visual correspondences by performing asynchronous multi-threaded processing that includes receiving one or more images and associated metadata as input and sending visual feature matches and associated metadata as output.

[0152] In some embodiments, updating the geometry-based warping transform includes correcting the frame-based geometry metadata based on the visual correspondence.

[0153] In some embodiments, updating the geometry-based warping transform includes estimating a homography from the corrected geometry metadata, the homography mapping pixels in a plane of a first coordinate system associated with a first image capture device to corresponding pixels in the same plane of a second coordinate system associated with a second image capture device.

[0154] In some embodiments, the updating of the geometry-based warping transform utilizes frame-based data including one or more of images associated with the first image capture device and the second image capture device, respectively, a pre-crop of the image, a scene depth, or calibration parameters, hi some embodiments, the calibration parameters include an autofocus distance.

[0155] In some embodiments, application of the updated warping transformation is performed on each frame of the initial preview and the corresponding frame of the zoomed preview in the side-by-side comparison.

[0156] In some embodiments, aligning the zoom preview with the initial preview includes aligning a depth value of a point in image space with a geometric focal length of the point on each frame of the initial preview and the corresponding frame of the zoom preview.

[0157] Some embodiments include generating a bundle adjustment that applies one or more camera calibrations and one or more focal lengths for each frame of the initial preview and the corresponding frame of the zoomed preview.

[0158] Some embodiments include generating, for a collection of consecutive frames, a modified bundle adjustment based on the bundle adjustment of each of the consecutive frames.

[0159] Some embodiments include transitioning from the first image capture device to the second image capture device by the computing device and based on the updated warping transformation.

[0160] In some embodiments, the second focal length range may be larger or smaller than the first focal length range, corresponding to a zoom-in or zoom-out operation on the computing device.

[0161] In some embodiments, the one or more viewing artifacts include binocular disparity.

[0162] In some embodiments, updating the geometry-based warping transform includes reducing jitter by applying temporal feature matching and tracking.

[0163] The particular configuration shown in the figures should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in a given figure. Furthermore, some of the illustrated elements may be combined or omitted. Furthermore, example embodiments may include elements not shown in the figures.

[0164] Steps or blocks representing the processing of information may correspond to circuitry that can be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, steps or blocks representing the processing of information may correspond to modules, segments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor to implement specific logical functions or operations in the method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device, including a disk, hard drive, or other storage medium.

[0165] Computer-readable media may also include non-transitory computer-readable media, such as register memory, processor cache, and computer-readable media that store data for a short period of time, such as random access memory (RAM). Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for a long period of time. Thus, computer-readable media may include, for example, secondary or persistent long-term storage, such as read-only memory (ROM), optical or magnetic disks, compact disk read-only memory (CD-ROM), etc. Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may be considered, for example, to be a computer-readable storage medium or a tangible storage device.

[0166] While various examples and embodiments have been disclosed, other examples and embodiments will be apparent to those skilled in the art. The various examples and embodiments disclosed are for purposes of illustration and are not intended to be limiting, the true scope being indicated by the following claims.

Claims

1. 1. A computer-implemented method, the method comprising: displaying, by a display screen of a computing device, an initial preview of a scene to be captured by a first image capture device of the computing device, the first image capture device operating within a first focal length range; detecting, by the computing device, a zoom operation that is predicted to cause the first image capture device to reach a limit of the first focal length range; and in response to the detecting, activating a second image capture device of the computing device to capture a zoom preview of the scene, the second image capture device configured to operate within a second focal length range, the method further comprising: updating a geometry-based warping transformation based on a comparison of respective image features from the initial preview and the zoomed preview; and aligning the zoomed preview with the initial preview by applying the updated warping transformation, wherein the updated warping transformation reduces one or more visual artifacts caused by a change in field of view upon transition from the initial preview to the zoomed preview, the method further comprising: displaying, by the display screen of the computing device, the aligned zoom preview of the image captured by the second image capture device while operating within the second focal length range.

2. The comparison of each of the image features comprises: detecting one or more visual features within the initial preview and the zoomed preview; generating a visual correspondence between the initial preview and the zoom preview based on the one or more visual features; The method of claim 1 further comprising:

3. 3. The method of claim 2, further comprising optimizing the detection of the one or more visual features and the generation of the visual correspondences by performing asynchronous multi-threaded processing that includes receiving one or more images and associated metadata as input and sending visual feature matches and associated metadata as output.

4. The method of claim 2 , wherein the updating of the geometry-based warping transformation comprises correcting frame-based geometry metadata based on the visual correspondence.

5. 5. The method of claim 4, wherein the updating of the geometry-based warping transform includes estimating a homography from the corrected geometry metadata, the homography mapping pixels in a plane of a first coordinate system associated with the first image capture device to corresponding pixels in the same plane of a second coordinate system associated with the second image capture device.

6. 2. The method of claim 1 , wherein the updating of the geometry-based warping transformation utilizes frame-based data including one or more of images associated with the first image capture device and the second image capture device, respectively, pre-crops of the images, scene depth, or calibration parameters.

7. The method of claim 6 , wherein the calibration parameters include an autofocus distance.

8. The method of claim 1 , wherein the application of the updated warping transformation is performed on each frame of the initial preview and a corresponding frame of the zoomed preview in a side-by-side comparison.

9. 2. The method of claim 1, wherein aligning the zoom preview with the initial preview includes aligning a depth value of a point in image space with a geometric focal length of the point on each frame of the initial preview and a corresponding frame of the zoom preview.

10. The method of claim 1 , further comprising generating, for each frame of the initial preview and a corresponding frame of the zoomed preview, a bundle adjustment applied to one or more camera calibrations and one or more focal lengths.

11. The method of claim 10 , further comprising generating, for a collection of consecutive frames, a modified bundle adjustment based on the bundle adjustment of each of the consecutive frames.

12. The method of claim 1 , further comprising transitioning from the first image capture device to the second image capture device by the computing device and based on the updated warping transformation.

13. The method of claim 1 , wherein the updating of the geometry-based warping transform further comprises reducing jitter by applying temporal feature matching and tracking.

14. 1. A computing device, comprising: A display screen; a first image capture device configured to operate within a first focal length range; a second image capture device configured to operate within a second focal length range; and one or more processors; Data storage; the data storage storing computer-executable instructions that, when executed by the one or more processors, cause the mobile device to perform functions, the functions including: displaying, by the display screen, an initial preview of a scene captured by the first image capture device; detecting, by the computing device, zoom operations that are likely to cause the first image capture device to reach a limit of the first focal length range; activating the second image capture device to capture a zoom preview of the scene in response to the detection; updating a geometry-based warping transformation based on a comparison of respective image features from the initial preview and the zoomed preview; and aligning the zoomed preview with the initial preview by applying the updated warping transformation, the updated warping transformation reducing one or more visual artifacts caused by a change in field of view upon transition from the initial preview to the zoomed preview, the function further comprising: displaying, by the display screen of the computing device, the aligned zoom preview of the image captured by the second image capture device while operating within the second focal length range.

15. The function for the comparison of each of the image features is: detecting one or more visual features within the initial preview and the zoomed preview; generating a visual correspondence between the initial preview and the zoom preview based on the one or more visual features; The computing device of claim 14 further comprising:

16. The computing device of claim 15 , wherein the functionality for updating the geometry-based warping transformation further comprises correcting frame-based geometry metadata based on the visual correspondence.

17. 17. The computing device of claim 16, wherein the functionality for updating the geometry-based warping transformation includes estimating a homography from the corrected geometry metadata, the homography mapping pixels in a plane of a first coordinate system associated with the first image capture device to corresponding pixels in the same plane of a second coordinate system associated with the second image capture device.

18. 15. The computing device of claim 14, wherein the updating of the geometry-based warping transformation utilizes frame-based data including one or more of images associated with the first image capture device and the second image capture device, respectively, pre-crops of the images, scene depth, or calibration parameters.

19. 15. The computing device of claim 14, wherein the function for applying the updated warping transformation is performed on each frame of the initial preview and a corresponding frame of the zoomed preview in a side-by-side comparison.

20. 20. The computing device of claim 19, wherein the functionality for aligning the zoom preview with the initial preview includes aligning a depth value of a point in image space with a geometric focal length of the point on each frame of the initial preview and a corresponding frame of the zoom preview.

21. The computing device of claim 14 , wherein the functionality for updating the geometry-based warping transformation further comprises reducing jitter by applying temporal feature matching and tracking.

22. A non-transitory computer-readable medium containing program instructions executable by one or more processors, the program instructions causing the one or more processors to perform operations, the operations including: displaying, by the display screen, an initial preview of a scene captured by the first image capture device; detecting, by the computing device, zoom operations that are likely to cause the first image capture device to reach a limit of the first focal length range; activating the second image capture device to capture a zoom preview of the scene in response to the detection; updating a geometry-based warping transformation based on a comparison of respective image features from the initial preview and the zoomed preview; and aligning the zoomed preview with the initial preview by applying the updated warping transformation, the updated warping transformation reducing one or more visual artifacts caused by a change in field of view upon transition from the initial preview to the zoomed preview, the operations further comprising: displaying, by the display screen of the computing device, the aligned zoom preview of the image captured by the second image capture device while operating within the second focal length range.

Citation Information

Patent Citations

  • Method and system for obtaining multiple views of an object via real-time video output.

    JP2009528766A

  • Apparatus and method for digital microscope imaging

    JP2014529922A

  • Systems and methods for implementing seamless zoom function using multiple cameras

    US20170230585A1

  • Multi-Camera Video Stabilization

    US20220053133A1