Smooth continuous zoom by image-based visual features and optimized geometric calibration in multi-camera systems

By using geometric metadata and image features for distortion transformation in a multi-camera system, the image distortion problem during camera switching is solved, and smoother and higher-quality image capture is achieved.

CN120226376APending Publication Date: 2025-06-27GOOGLE LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380081715.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-29
Filing Date
2023-09-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In a multi-camera system, image distortion (binocular parallax) problems caused by field of view changes during camera switching affect the smoothness and quality of image capture.

Method used

By estimating the distortion transformation by utilizing available geometric metadata and image features, the image of one camera is distorted to almost align with the image of another camera, thereby reducing perceived image distortion during camera switching.

Benefits of technology

It effectively reduces image distortion during camera switching, improves the smoothness and quality of image capture, and enhances the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

An example method includes displaying an initial preview of a scene captured by a first camera operating within a first focal length range. The method includes detecting a zoom operation predicted to cause the first camera to reach a limit of the first range. The method includes activating a second camera operating within a second focal length range to capture a zoomed preview of the scene. The method includes updating a geometry-based twist transform based on a comparison of respective image features from the initial preview and the zoomed preview. The method includes aligning the zoomed preview with the initial preview by applying the updated twist transform. The method includes displaying an aligned zoom preview of an image captured when the second camera is operating within the second range.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference to related applications is incorporated by reference

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 377,581, filed on September 29, 2022, the entire content of which is hereby incorporated by reference. Background Art

[0002] Many modern computing devices, including mobile phones, personal computers, and tablet computers, include image - capture devices. Some image - capture devices are configured with multi - camera systems. The camera systems are configured to cooperate to meet different image - capture requirements using their respective specifications. A smart phone can integrate multiple types of cameras with multiple focal lengths to handle objects at different distances and scenes in different fields of view (FOV). Summary of the Invention

[0003] The present disclosure generally relates to smooth transitions between multiple cameras. In one aspect, an image - capture device can include multiple cameras. A transition between cameras can cause a perceivable image distortion, such as binocular disparity, due to, for example, a change in the field of view. As described herein, a warping transform is estimated based on available geometric metadata and image - based data to warp an image from one camera to be nearly aligned with an image from another camera, thereby reducing the perceivable image distortion during a camera switch.

[0004] In a first aspect, a computer - implemented method is provided. The method includes: displaying, by a display screen of a computing device, an initial preview of a scene captured by a first image - capture device of the computing device, where the first image - capture device is operating within a first focal - length range. The method further includes: detecting, by the computing device, a zoom operation that is predicted to cause the first image - capture device to reach a limit of the first focal - length range. The method further includes: in response to the detection, activating a second image - capture device of the computing device to capture a zoomed - in preview of the scene, where the second image - capture device is configured to operate within a second focal - length range. The method additionally includes: updating a geometry - based warping transform based on a comparison of corresponding image features from the initial preview and the zoomed - in preview. The method further includes: aligning the zoomed - in preview with the initial preview by applying the updated warping transform, where the updated warping transform reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed - in preview. The method further includes: displaying, by the display screen of the computing device, the aligned zoomed - in preview of the image captured by the second image - capture device while operating within the second focal - length range.

[0005] In a second aspect, a computing device is provided. The computing device includes a display screen, a first image capture device configured to operate within a first focal length range, a second image capture device configured to operate within a second focal length range, one or more processors, and data storage, wherein computer-executable instructions are stored on the data storage, and the computer-executable instructions, when executed by the one or more processors, cause the mobile device to perform functions. The operations include: displaying, by the display screen, an initial preview of a scene captured by the first image capture device; detecting, by the computing device, a zoom operation that may cause the first image capture device to reach the limit of the first focal length range; in response to the detection, activating the second image capture device to capture a zoomed preview of the scene; updating a geometry-based distortion transformation based on a comparison of corresponding image features from the initial preview and the zoomed preview; aligning the zoomed preview with the initial preview by applying the updated distortion transformation, wherein the updated distortion transformation reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed preview; and displaying, by the display screen of the computing device, the aligned zoomed preview of an image captured by the second image capture device while operating within the second focal length range.

[0006] In a third aspect, an article of manufacture is provided. The article of manufacture may include a non-transitory computer-readable medium having program instructions stored thereon, and the program instructions, when executed by one or more processors of a computing device, cause the computing device to perform operations. The operations include: displaying, by the display screen, an initial preview of a scene captured by the first image capture device; detecting, by the computing device, a zoom operation that may cause the first image capture device to reach the limit of the first focal length range; in response to the detection, activating the second image capture device to capture a zoomed preview of the scene; updating a geometry-based distortion transformation based on a comparison of corresponding image features from the initial preview and the zoomed preview; aligning the zoomed preview with the initial preview by applying the updated distortion transformation, wherein the updated distortion transformation reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed preview; and displaying, by the display screen of the computing device, the aligned zoomed preview of an image captured by the second image capture device while operating within the second focal length range.

[0007] In a fourth aspect, a system is provided. The system includes: means for displaying, by a display screen, an initial preview of a scene captured by a first image capture device; means for detecting, by a computing device, a zoom operation that may cause the first image capture device to reach a limit of a first focal length range; in response to the detection, means for activating a second image capture device to capture a zoomed preview of the scene; means for updating a geometry-based distortion transform based on a comparison of corresponding image features from the initial preview and the zoomed preview; means for aligning the zoomed preview with the initial preview by applying the updated distortion transform, wherein the updated distortion transform reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed preview; and means for displaying, by a display screen of the computing device, the aligned zoomed preview of an image captured when the second image capture device operates within a second focal length range.

[0008] Other aspects, embodiments, and implementations will become apparent to those of ordinary skill in the art upon reading the following detailed description with appropriate reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 Shows binocular disparity in a multi-camera system according to an example embodiment.

[0010] Figure 2 Is a flowchart of a workflow for image-based calculation of a distortion transform according to an example embodiment.

[0011] Figure 3 Is an example sparse feature workflow for smooth continuous zoom in a multi-camera system according to an example embodiment.

[0012] Figure 4A Shows temporal feature matching and tracking according to an example embodiment.

[0013] Figure 4B Shows an example image of temporal feature matching and tracking according to an example embodiment.

[0014] Figure 4C Shows an example application 400 of temporal feature matching and tracking according to an example embodiment.

[0015] Figure 5 Is an example dense feature workflow for smooth continuous zoom in a multi-camera system according to an example embodiment.

[0016] Figure 6 Shows an example disposition of incremental data during camera transition according to an example embodiment.

[0017] Figure 7It is a table showing various situations for switching between a telephoto camera and a wide - angle camera according to an exemplary embodiment.

[0018] Figure 8A It depicts an exemplary geometric relationship at each pair of matching pixels according to an exemplary embodiment.

[0019] Figure 8B It depicts a workflow for determining the geometric relationship at each pair of matching pixels according to an exemplary embodiment.

[0020] Figure 9 It depicts an exemplary workflow for smooth continuous zoom in a multi - camera system according to an exemplary embodiment.

[0021] Figure 10 It depicts a distributed computing architecture according to an exemplary embodiment.

[0022] Figure 11 It is a block diagram of a computing device according to an exemplary embodiment.

[0023] Figure 12 It is a flowchart of a method according to an exemplary embodiment. Detailed Description

[0024] Exemplary methods, apparatuses, and systems are described herein. It should be understood that the words "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration". Any embodiment or feature described herein as "example" or "exemplary" is not necessarily to be construed as preferred over or superior to other embodiments or features. Without departing from the scope of the subject matter presented herein, other embodiments may be utilized and other changes may be made.

[0025] Accordingly, the exemplary embodiments described herein are not intended to be limiting. Aspects of the present disclosure generally described herein and illustrated in the figures can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are contemplated herein.

[0026] Further, unless the context otherwise implies, the features shown in each figure can be used in combination with each other. Thus, the figures should generally be considered as constituent aspects of one or more overall embodiments, but it should be understood that not all of the features shown are necessary for each embodiment. Overview

[0027] A smart phone or other mobile device that supports image and / or video capture can be equipped with multiple cameras that cooperate using corresponding specifications to meet different image capture requirements. A smart phone can integrate multiple types of cameras with multiple focal lengths to display and / or capture objects at different distances and scenes in different fields of view (FOV).

[0028] For example, a phone can be configured with a main camera having a medium focal length for meeting normal photo / video capture requirements, a telescopic camera having a longer focal length for capturing distant objects, and an ultra-wide-angle camera having a shorter focal length for capturing a larger FOV. During a photo / video capture session, when the user continuously zooms in to focus on a distant object, a switch from the main camera to the telescopic camera can occur, and when the user continuously zooms out to capture a larger field of view, a switch from the main camera to the ultra-wide-angle camera can occur. A multi-camera system provides a much larger range of focus distances than a single camera. However, a sudden camera switch during zooming can result in a view difference (referred to as "binocular parallax").

[0029] To circumvent binocular parallax, a distortion transformation can be estimated based on available geometric metadata and image features to warp the image of one camera to be nearly aligned with the image of another camera, such that the change during a camera switch is less perceivable. The distortion transformation can include scaling, rotation, reflection, identity mapping, shearing, or various combinations thereof. Additionally, for example, translation, similarity, affine mapping, and / or projective mapping can be used as the distortion transformation. Generally, two planar images can be related by a distortion transformation (such as a homography). For example, computer vision means for calculating a homography can be used, which can warp an image frame from one camera to another camera. As described herein, a homography calculation can be determined to reduce the view difference during a camera switch during zooming. The homography calculation can use geometric information (without image features), which includes metadata such as camera calibration data, focus distance, etc. Although a geometry-based solution can be used, the presence of electrical and / or mechanical components (such as a voice coil motor (VCM), optical image stabilization (OIS) adjustment, and / or thermal effects of the device) can cause changes in dynamic camera calibration and focus distance, which can introduce errors when determining an accurate distortion transformation for a smooth viewing experience, resulting in a sudden transition.

[0030] Some existing means attempt to address this problem. For example, the view of a physical camera can be warped to the same coordinates, and the distortion transformation can depend on camera calibration and focus distance. However, image-based features are not used, and thus, errors caused by VCM / OIS adjustment and / or thermal effects may still not be corrected. Another means can be to blend multiple camera views and apply a fade-in / fade-out style animation to obtain a smooth switch. However, this means depends on the simultaneous display of images from different cameras. Theoretical models for depth estimation using binocular parallax and motion parallax have been proposed, but there is no practical implementation to address the technical problems related to image capture devices.

[0031] This application relates to an “image-based” means (sometimes referred to herein as ContiZoom) for better helping to distort quality to overcome the adverse effects of VCM / OIS adjustment and / or thermal effects. Compared with “geometry-based” means, this new “image-based” means is designed to utilize image information and / or features as additional inputs to improve the distortion transformation used in geometry-based means.

[0032] The means described herein directly utilize image features to provide a more accurate metric for calculating the distortion transformation. This reduces the spatial difference between image frames from two cameras and mitigates the impact of many inaccurate sensor metadata from geometry-based means.

[0033] From a geometric perspective, thermal changes in the device can affect the principal point, which can cause the entire FOV to shift, and this in turn leads to unreliable output from camera parameter interpolation (CPI). Thermal changes affect the focus distance; and thus, with each successive frame and the continuous use of the device, the already inaccurate focus distance can become even more unstable due to additional thermal effects.

[0034] By utilizing image information (features) to adjust inaccurate geometric metadata, these factors can be largely mitigated. For example, image feature matching can be performed between two frames from two different physical cameras. Existing geometric metadata can be corrected based on image features. The geometry-based distortion transformation can be recalculated based on the geometric metadata that has been corrected based on image features.

[0035] As described herein, dual images from a pair of cameras that are being switched are used for the image-based smooth zoom described herein. If the continuous zoom quality is adversely affected by thermal changes or inaccurate estimation of the focus distance, the image-based visual information can effectively solve the resulting problems. Bundle adjustment can be applied to camera calibration and world points from visual feature matching such that the optimized parameters generate a more reliable homography for image distortion. The scene depth can be estimated based on both image-based visual features and phase difference, resulting in improved smoothness of the zoom when the camera is switched.

[0036] In some embodiments, the image-based algorithmic process can be executed at up to 30 frames per second (fps) and can be configured to seamlessly cooperate with other camera features such as image distortion correction, video stabilization, etc. By using multi-threading and DSP solutions, computationally intensive steps such as visual feature extraction can be made less intensive.

[0037] There are several benefits to using image-based visual features, including that images (e.g., in the conventional RGB format) can be readily obtained from a camera system of a device. Since image alignment during camera switching is a desired result of continuous zooming, the distortion estimated from the image itself is more reliable and appropriate. Such distortion effectively combines image-based visual features with geometry-based calibration and focus distance, thereby improving the smoothness and stability of zooming during camera switching.

[0038] Accordingly, the techniques described herein can improve an image capture device equipped with a multi-camera system by reducing and / or eliminating visual differences in images and / or video during camera transitions, thereby enhancing the actual and / or perceived quality of the images and / or video. Enhancement of the actual and / or perceived quality of a photo or video can provide user experience benefits. These techniques are flexible and can thus be applied to a wide variety of videos in both indoor and outdoor environments.

[0039] Hereinafter, the term "homography" is used to refer to one implementation of a distortion transformation. Additionally, terms such as "warped," "warping," etc. may be used, for example, in the context of applying a distortion transformation. Smooth continuous zoom

[0040] Figure 1 Binocular disparity in a multi-camera system according to an example embodiment is shown. For illustrative purposes, in Figure 1In it, both cameras t1 and t2 face the object (in focus and out of focus). In some cases, camera t2 can be physically mounted adjacent to camera t1 (e.g., on its right side, on its left side, etc.). For example, the camera positions can be designed to simulate human left / right eye vision. Generally, in-focus scene objects with the same depth (i.e., the distance to the camera), such as in-focus object 110, can be distorted from one camera t1 to the other camera t2 almost perfectly. For example, the in-focus object 110 in camera t1 is distorted to the in-focus object 110A in camera t2 without any difference. Planar objects with a plane perpendicular to the viewing direction of the camera can exhibit such properties. Zooming in and / or out triggers a camera switch (e.g., between wide-angle and ultra-wide-angle, between wide-angle and telephoto, etc.), resulting in a change in the FOV and a view difference called binocular disparity. For out-of-focus objects, a disparity (jump) between the images in the two cameras can be perceived. For example, the distant object 105 and the near object 115 are out of focus in camera t1. Therefore, when the camera is switched, the distant object 105 in camera t1 is mapped to the distant object 105A in camera t2, and the distant object 105A is displaced relative to the actual position 105B. Similarly, the near object 115 in camera t1 is mapped to the near object 115A in camera t2, and the near object 115A is displaced relative to the actual position 115B.

[0041] As described herein, a warping transformation can be applied to reduce binocular disparity by warping an image from one camera to another. In-focus scene objects with the same depth (e.g., in-focus object 110) can be warped from one camera to another without perceivable disparity.

[0042] For out-of-focus objects (e.g., distant object 105, near object 115, etc.), warping differences can occur, or for planes across multiple depths, warping distortions can occur. Such differences for out-of-focus objects depend on the depth of the camera and the baseline. For example, for in-focus planar objects with a plane tilted relative to the viewing direction of the camera, there can be some "rotation" type differences.

[0043] Figure 2 It is a flowchart of a workflow 200 for image-based calculation of a warping transformation according to an exemplary embodiment.

[0044] At block 210, the workflow involves obtaining frame-based data from a first camera and a second camera. A frame can be regarded as a data processing unit. A frame includes the input data required for continuous zooming based on an image, including the image, pre-cropping of the image, camera calibration including internal and external parameters, autofocus distance, and other metadata.

[0045] At block 220, the workflow involves performing visual feature detection and matching to determine visual correspondence. A variety of visual feature detectors and descriptors can be used, such as for example Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), Fast Retina Keypoints (FREAK), or Features from Accelerated Segment Test (FAST) corner detection algorithms, among others. Additional and / or alternative visual features can be used in the pipeline as long as the algorithm achieves a sufficient quality of visual correspondence.

[0046] Image feature matching can be performed in two steps - such as feature extraction and feature matching. A variety of feature extraction and matching means can be used. For the purposes of this document, existing feature matching methods can be used, such as ArCore features or ILK features. The term "ILK" as used herein generally refers to an inverse search version of the Lucas-Kanade algorithm for optical flow estimation.

[0047] At block 230, the workflow involves correlating visual feature matching, camera calibration, and autofocus distance for each frame via a two-view bundle. Camera calibration can be used to estimate the depth of the points observed as feature matches. As described herein, the OIS / VCM system and / or thermal effects can cause geometric metadata (such as camera calibration and focus distance) to be updated frame-by-frame with significant error. Thus, manipulation of specific calibration parameters can be performed to bring the depth values of most points closer to the geometric focus distance.

[0048] Some embodiments can achieve smooth transitions between cameras (e.g., wide-angle camera and telephoto camera), but result in a FOV jitter problem on the source camera where distortion is occurring. For example, visual features from a downscaled image (e.g., 320×240) may not correspond to the same landmarks frame-by-frame. Additionally, for example, camera calibration is updated frame-by-frame due to OIS / VCM updates, and the scene depth of field is calculated and / or corrected frame-by-frame based on valid (e.g., normal value (inlier)) feature matches between two cameras (e.g., wide-angle and telephoto). In some embodiments, if the two cameras have different FOVs, the number of normal value visual feature matches can be limited by the smaller FOV (e.g., telephoto), resulting in a waste of marginal visual information from the larger FOV camera (e.g., wide-angle).

[0049] Some means of reducing such jitter can involve damping control of the change in scene focus distance, and then triggering image-based ContiZoom during zooming. To further address the jitter problem and use as much of the visual information provided by the larger FOV camera as possible, it can be done temporally between adjacent frames and frame Perform visual feature matching between them to achieve the purpose of temporal consistency, so that each frame will consider the geometric metadata of the previous frame when determining the warping grid.

[0050] At block 240, the workflow involves: performing bundle adjustment based on the two-view bundle relationship to obtain a set of optimized camera calibrations, focus distances, and other parameters involved. For example, image-based visual information can be effectively used to correct the geometric metadata from the upstream module, making the geometric metadata more compatible with the image to be displayed as a continuous zoom preview.

[0051] If bundle adjustment is performed frame by frame, the corresponding per-frame optimized solutions for camera calibration and focus distance can be determined independently (e.g., with the minimum reprojection pixel error). In some embodiments, misalignment can exist across frames, resulting in a jittery preview when the warped frames are played in sequence for display.

[0052] In such embodiments, a geometric bundle can be constructed across a frame window, and the optimization of camera calibration and other metadata may not always result in smooth changes under the warping transformation. Therefore, instead of applying the homography of the latest optimized data to warp the image, a damping process given by the following equation is introduced:

[0053] To achieve a gradual and smooth change of the warping transformation. In Equation 1, the term is a damping ratio with a value between 0 and 1, is the homography to be applied to frame and is the optimized image-based homography at frame .

[0054] In some embodiments, the homography is determined based on the geometric metadata from the upstream module (e.g., camera calibration and focus distance) before extracting the image-based visual features. However, as mentioned before, this can include errors from OIS / VCM updates, thermal effects, and other sources. This can be corrected using image-based data as follows:

[0055] Generally, there are two sets of camera calibration models available. One calibration model has been updated through OIS / VCM correction that directly corresponds to the visual features from the image, and the other calibration model remains in a neutral state and is used as a smooth initial value for further geometric optimization.

[0056] The previously computed geometry-based homography It can be used to perform a rough-level pre-distortion of image features and associated calibration, followed by the aforementioned image-based process to correct the remaining errors. Such means effectively combine geometric information and visual information to solve technical problems.

[0057] At block 250, the workflow involves: determining a pre-distortion transformation of the image of the first camera based on bundle adjustment such that the distorted image has only a small disparity with the image of the second camera.

[0058] Subsequently, at block 260, the workflow involves: modifying the pre-distortion transformation based on image features to finely distort the image of the first camera, thereby further reducing the small disparity from the pre-distortion transformation.

[0059] Some embodiments involve optimizing the detection of one or more visual features and the generation of visual correspondences by performing asynchronous multi-threaded processing, which includes receiving one or more images and associated metadata as input and sending visual feature matches and associated metadata as output. For example, continuous zoom based on images can involve computationally intensive steps such as visual feature detection and matching. To enable the solution to run in real time on consumer-grade devices (e.g., at least 30 frames per second (fps)), the computationally intensive steps can be processed asynchronously by a specific thread that receives the image and associated metadata as input and sends visual feature matches and associated metadata as output.

[0060] The term "sparse feature" as used herein generally refers to detected features that are sparsely distributed throughout the image. Feature points can be detected when a pixel and its neighborhood meet a detection threshold. This can include, for example, ArCore features, portrait mode features, and AutoCal features.

[0061] The term "dense feature" as used herein generally refers to detected features that can cover the entire image, and feature points can be detected based on predefined image patches, and matches can be found for each patch. In some embodiments, ILK can be used for dense feature detection. For example, the ILK algorithm can be used to extract "dense" feature points and matches from an image.

[0062] Dense or sparse features can generally have different designs in the ContiZoom pipeline, as described in further detail below. Recalibration (sparse feature process)

[0063] Contrary to the planar target features typically used during factory calibration, the calibration of sparse features uses natural features to recalibrate the geometric information received from the camera sensor. For example, image features are used to update the geometric metadata, and existing geometry-based calculations are utilized to compute the modified homography.

[0064] Figure 3 Is an example sparse feature workflow 300 for smooth continuous zoom in a multi-camera system according to an example embodiment. A general flow regarding sparse features is shown.

[0065] At block 305, an input image including dual images is received. The input image is from two target cameras. In some examples, the range of the image resolution can be up to . The quality of feature matching depends more on the quality of the texture of the scene.

[0066] At block 310, sparse feature detection can be performed as previously described. In some embodiments, ArCore features can be used. In some embodiments, FAST features can be used. Generally, scale-invariant features need not be used because the dual images can be rescaled to provide the focal length relatively accurately.

[0067] Natural feature calibration can be performed at block 315. This process recalibrates the camera parameters 325 from factory calibration provided, for example, by the camera provider 320. In some embodiments, a DualCameraCalibrator or AutoCal based on the respective FAST feature detector can be used. Additionally, for example, an optical flow-based detector such as ILK can be used. However, natural feature calibration is generally different from factory calibration because natural features do not come from planar objects. Therefore, means based on bundle adjustment (BA) (e.g., Figure 2 block 240) may be more suitable for performing natural feature calibration.

[0068] In some embodiments, natural feature calibration may not optimize all parameters but can focus on the "principal point" and "extrinsic rotation" for wide-angle and telephoto. The following table (Table 1) summarizes the parameters that may need to be optimized. Table 1 is for illustrative purposes only and can vary depending on the device and can be based on the type and / or characteristics of the cameras involved in the transition process.

[0069] At block 330, incremental camera metadata can be obtained. CPI-based calibration is relative to active array coordinates, while image-based algorithms require calibration relative to image coordinates. Thus, a transformation can be determined between the active array and the image. After correcting the camera metadata, the difference between the factory calibration and the corrected metadata can be stored as incremental metadata and saved separately from the CPI calibration metadata. In some embodiments, the increment can be a constant offset during the transition period from one camera to another.

[0070] In image-based correction, camera metadata of the same structure can be used, which is the increment between the CPI output and the recalibrated camera metadata. In some embodiments, an incremental focus distance (e.g., depth) can be used.

[0071] At block 335, features within a region of interest (ROI) can be detected. As used herein, an ROI is a sub-region within an image frame that is considered important to the user and serves as a guiding region for many features (such as autofocus) in camera applications, providing a focus distance for geometry-based methods. In some embodiments, the ROI can be obtained as an ROI rectangle 340 according to algorithms such as a face detection algorithm, a saliency detection algorithm, etc. Generally, for sparse features and dense features, the ROI can be processed in different ways. The features within the ROI can be based on the sparse features detected at block 310.

[0072] At block 345, the median depth in the ROI is determined. For example, the median of the depths of the (normal value) image feature points extracted by the above sparse / dense feature method can be determined. Time Feature Matching and Tracking

[0073] In some embodiments, temporal feature matching and tracking can be performed. Generally, the same landmark or ROI can appear in multiple frames (e.g., three or more frames), resulting in a feature trajectory. In some embodiments, temporal feature tracking can be applied only to a larger FOV (e.g., wide-angle) camera, which is sufficient for the temporal consistency of the ContiZoom grid.

[0074] Figure 4A Temporal feature matching and tracking according to an example embodiment are shown. Referring Figure 4A , multiple consecutive frames are shown for a wide-angle camera 405 and a telephoto camera 410. For the wide-angle camera 405, the first frame 415 at time is shown and the second frame 420 at time is shown as two consecutive frames. For the telephoto camera 410, the time at the third frame 425 and shows the time at the fourth frame 430, as two consecutive frames. It shows the intra-frame feature matching, where at time at the first feature in the first frame 415 matches the corresponding feature in the third frame 425 Similarly, it shows the intra-frame feature matching, where at time at the second feature in the second frame 420 matches the corresponding feature in the fourth frame 430 It shows the temporal feature matching and tracking, where at time the first feature in the first frame 415 at time matches the second feature in the second frame 420 For illustrative purposes, the temporal feature tracking is shown for the wide-angle camera 405. Generally, it may be desirable to perform temporal feature tracking in a camera with a larger FOV in order to capture relevant feature trajectories.

[0075] Figure 4B It shows an example image of temporal feature matching and tracking according to an example embodiment. Referring to Figure 4B , two images are shown. The first image 435 shows the feature matching without damping the focus distance. The second image 440 shows the temporal feature matching for reducing jitter.

[0076] Generally, the temporal feature matching and tracking may involve two tasks: 1) determining the temporal feature tracking information from the images; and 2) applying the temporal feature tracking to the existing pipeline to improve the ContiZoom quality.

[0077] In some embodiments, the first task may involve providing interface functions for feature extraction and feature matching respectively. As described with respect to Figure 4A , the intra-frame feature matching can be performed on the dual images of the lead camera and the follower camera (e.g., the wide-angle camera 405 compared to the telephoto camera 410) to construct the intra-frame feature matching. In some embodiments, the temporal feature tracking can be performed by implementing the feature matching between adjacent frames and where the reconstruction interface function is and the extracted features from the previous frame are cached.

[0078] In some embodiments, the second task may involve building an index manager to handle the indexing of visual feature points and the matching between multiple images. The index manager facilitates the utilization of temporal feature trajectories along with in-frame feature matching. For example, the index manager manages the indexing of visual features (including the indexing of feature points and the indexing of feature matches) along with their mutual correspondences. In some embodiments, the index manager may support querying the feature point index from the feature match index, and querying the feature match index from the indices of the first and second feature points of the match.

[0079] Some embodiments may involve one or more mappings. For example, a first mapping from the index of a feature point in a first image to the index of a matching pair involving that feature point. As another example, a second mapping from the index of a feature point in a second image to the index of a matching pair involving that feature point. Additionally, for example, a vector of feature matches may be determined. For example, each feature match may involve two feature points from a first image and a second image respectively.

[0080] Experimental evidence indicates that jitter on the wide-angle, i.e., the leading camera, may be mainly caused by jitter in the scene focus distance estimated per frame. Therefore, temporal feature trajectories can be used to estimate the scene focus distance at a frame. Utilizing feature trajectories across multiple frames achieves quality improvement under temporal consistency.

[0081] In some cases, it can be reasonably assumed that when a user performs zoom-in and / or zoom-out operations using a camera, the user may not move significantly (e.g., pan, run, walk, or have rapid changes in significant objects / ROIs). If the user has significant movement, small jitter in zoom or FOV transition is less likely to be noticeable. If the user does not move significantly, the change in the scene focus distance between adjacent frames and needs to be managed to avoid perceivable jitter. One means to achieve this is: if frames and have a sufficient number of normal value matches, the scene focus distance is kept unchanged.

[0082] Figure 4C Illustrates an example application 400 of temporal feature matching and tracking according to an example embodiment. Figure 4C Shows multiple consecutive frames of a wide-angle camera 405 and a telephoto camera 410. Generally, each frame can estimate the scene depth based on a set of feature matches between the same frame in two cameras (e.g., the wide-angle camera 405 and the telephoto camera 410). In some embodiments, normal value features can be selected to represent the scene for depth estimation. For illustrative purposes, the presented example involves four ( ) scene focus distances, namely , , and . The time feature tracking between frames and frame can be performed as follows.

[0083] If the selected normal value feature has a time-matched counterpart, the scene focus distance can be set to be the same as . For example, feature A in frame t1 445 has a corresponding feature A' in frame t1' 465. Additionally, feature A in frame t1 445 has a time-matched counterpart feature B in frame t2 450. Therefore, the scene focus distance can be set to be the same as .

[0084] If the selected normal value feature has no time match, but another feature with a depth similar to the initially selected normal value has a time match and an intra-frame feature match, the scene focus distance can be set to be the same as . For example, feature E in frame t3 455 has no time match. However, feature D in frame t3 455 is at a similar depth to feature E in frame t3 455. Additionally, feature D in frame t3 455 has a time-matched counterpart feature C in frame t2 450. Furthermore, feature D in frame t3 455 has a corresponding feature D' in frame t3' 475. Therefore, the scene focus distance can be set to be the same as .

[0085] If the selected normal value feature and its sibling feature with a similar depth have neither a time match nor an intra-frame feature match, recalculation is performed. For example, feature H in frame t4 460 has a corresponding feature H' in frame t4' 480. However, feature H in frame t4 460 does not match another feature in time. Features G and I in frame t4 460 appear to be features with a depth similar to feature F. Feature G matches feature F in frame t3 455 in time; however, feature G has no intra-frame feature match. Feature I also has no intra-frame feature match or time match. Therefore, recalculation is performed.

[0086] In the above method, the estimation of the scene depth (i.e., the focus distance) is based on the intra-frame (i.e., between cameras) feature match for each frame, and time tracking is usually used as a post-verification to determine whether the scene focus distance changes in subsequent frames. However, this method cannot effectively utilize the information from time feature tracking.

[0087] Thus, alternative means of reducing and / or eliminating jitter by adjusting the scene focus distance can involve directly using temporal feature tracking information, especially when such information is available in sufficiently high quality. Factors that can determine the quality of the temporal feature tracking information can include one or more of the following: (1) a sufficient number of temporal matches between adjacent frames at time and time ; (2) a sufficient number of temporal matches that can form a temporal track across multiple frames in the absence of large pan and / or rotational motion; or (3) power and latency that are tolerable when running on a mobile device.

[0088] In one means, temporal feature tracking can be used together with gyroscopic measurements for triangulation and selecting “up-to-scale” 3D points as good normal value landmarks. Subsequently, binocular camera observations of these normal value landmarks can be used to estimate the scene depth for each frame.

[0089] In another means, temporal feature tracking can be used together with gyroscopic and accelerometer measurements to directly estimate the device pose and 3D points as normal value landmarks. Then, the scene depth for each frame can be based on these normal value landmarks. Given the power and latency aspects of device pose estimation, this means can be more suitable for offline processing.

[0090] Referring again to Figure 3 , the incremental camera metadata from block 330 and the median depth from block 345 are used at block 350 to update the geometric homography based on image features (to determine the updated distortion transformation). For example, matched feature points can be used to correct inaccurate geometric metadata. This can involve two means: a sparse feature flow means based on recalibration and a dense feature flow means based on image homography.

[0091] Homography is a matrix that maps pixels on a plane in the first coordinate system of the first camera to corresponding pixels on the same plane in the second coordinate system of the second camera. The homography can be decomposed as follows:

[0092] where and are the intrinsic matrices corresponding to the first camera and the second camera respectively, containing the focal length and the principal point, and where is the extrinsic parameter that can transform a point in the first coordinate system of the first camera to the second coordinate system of the second camera, and Define the epipolar plane of the homography in the coordinates of the first camera such that a point X in the plane satisfies . Generally, use .

[0093] Given camera calibration and the object distance for focusing, there are at least two ways to determine the homography. For example, in the first method, the homography can be determined by inputting these parameters into the formula shown in Equation 2. Additionally, for example, in the second method, the homography can be determined based on a set of pixel pairs (e.g., at least 4 sample point pairs). For example, the homography matrix transforms a plane (with a certain depth) on the telephoto camera to the corresponding plane on the wide-angle camera. In some embodiments, for a pair of cameras (e.g., telephoto and wide-angle camera models), a "four-point" method can be used to calculate the homography matrix between the telephoto camera and the wide-angle camera. The input can include the telephoto and wide-angle camera models (intrinsic and extrinsic parameters) and the distance to the target plane of the telephoto (object distance) camera.

[0094] The "four-point" method can involve arbitrarily selecting four two-dimensional (2D) points on the telephoto camera, denoted as . Next, the 2D points can be back-projected into 3D rays using the camera intrinsic parameters, denoted as Ray. Subsequently, Ray can intersect with a given plane that is at a distance Obj_dist away from the telephoto camera. This generates four three-dimensional (3D) points in the real space. The 3D points can be projected onto the wide-angle camera to obtain . Then, the homography matrix can be determined as:[[]]

[0095] where represents in homogeneous coordinates.

[0096] The second method of determining the homography based on the set of pixel pairs can involve estimating the distortion of the camera intrinsics during the process. If the vision information based on the image is not available, the camera calibration is from the CPI library, and the focusing distance is from the autofocus process.

[0097] At block 355, the highlighting processor performs highlighting processing based on the homography from block 350.

[0098] At block 360, the mesh transformation function is applied. Based on the image-based homography (dense feature flow)

[0099] The image-based approach directly computes the updated homography based on the matched image features and combines the image-based homography with the geometry-based homography. Since the goal is to warp two images (from two cameras) together, the warping homography can be directly computed using the image features without the geometry camera metadata.

[0100] Figure 5 This is an example dense feature workflow 500 for smooth continuous zoom in a multi-camera system according to an example embodiment. A general flow regarding dense features is shown.

[0101] At block 505, an input image including dual images is received. The input image is from two target cameras. In some examples, depending on the computing power of the computing device and the latency requirements of various use cases, the image resolution can be or greater. The quality of feature matching can depend more on the quality of the texture of the scene.

[0102] The aligned ROI regions are computed at block 510. Some embodiments can involve cropping the original image frames to the ROI regions. The ROI rectangle 525 can be obtained (e.g., as described previously with reference to ROI rectangle 340). Some embodiments involve aligning the two ROI regions with the computed geometry-based homography so that dense feature detection can have a better initial placement. In some embodiments, to save computing resources, the translation components can be extracted from the geometry-based homography and these translation components can be applied to the ROI. In some embodiments, this translation can be performed by cropping.

[0103] Dense feature detection can be performed at block 520. Such feature detection / matching has been described previously and the ILK algorithm can be used. This process recalibrates the camera parameters 530 from the factory calibration provided by, for example, the camera provider 535. In some embodiments, a DualCameraCalibrator or AutoCal based on the respective FAST feature detector can be used. Additionally, for example, an optical flow-based detector such as ILK can be used. However, natural feature calibration is usually different from factory calibration because natural features do not come from planar objects. Therefore, a bundle adjustment (BA)-based approach (e.g., Figure 2 block 240) may be more suitable for performing natural feature calibration.

[0104] At block 540, an individual incremental homography can be determined. In some embodiments, this can be based on, for example, camera parameters 530 from factory calibration provided by camera provider 535. In some embodiments, the dense features in the foreground can be separated from those in the background. For example, although feature detection can focus on features within the ROI, background features can also be used. In some embodiments, a translation-only homography can be applied to confirm that the homography plane will be perpendicular to the camera's axis. Additionally, for example, two clusters of means can be used, where the feature is disparity and the feature position is chosen as the center of the ROI.

[0105] Similar -means processes can be used for the counterpart process of the sparse feature flow. However, the sparse features may not have enough feature points to apply these means. Therefore, a sub-optimal solution based on determining the median depth can be used for the sparse feature flow (e.g., at block 345).

[0106] Typically, the foreground homography is an incremental homography on top of the geometry-based homography, since the original ROI has been shifted using the geometry-based homography.

[0107] At block 545, the updated warping transform (or combined homography) can be determined as a combination of the image-based incremental homography and the geometry-based homography. In some embodiments, the updated warping transform is a concatenation of two homography mappings (the image-based homography applied after the geometry-based homography), with appropriate coordinate transformations.

[0108] For both the dense feature flow and the sparse feature flow, existing components can be utilized. For example, the term "camera provider" (e.g., camera provider 320, camera provider 535) refers to a module that reads from a factory calibration file and uses the CPI to generate camera parameters corresponding to each OIS / VCM metadata.

[0109] The term "CPI camera parameters" (e.g., camera parameters 325, camera parameters 530) refers to the camera intrinsic and / or extrinsic parameters calculated by the CPI.

[0110] At block 550, the highlighting processor performs highlighting based on the combined homography from block 545.

[0111] At block 555, a grid transformation function is applied.

[0112] Sparse features can generally be more accurate and provide with precise Feature points of the image coordinates. However, sparse features can be relatively slow, and the number of detected features can depend on the scene complexity. Therefore, in the absence of a sufficient number of feature detections, the performance of calibration optimization can be negatively affected.

[0113] Dense features may not be as accurate because these "feature points" are essentially small image patches. However, dense feature detection is usually fast and less dependent on scene complexity.

[0114] In some embodiments, a combination of sparse features and dense features can be used. For example, sparse features can be detected first, and if the number of detected features is below a threshold, dense feature detection can be performed.

[0115] Additionally, for example, sparse features can be used to determine more accurate depth, and a better homography can be determined as an initial guess for dense feature matching.

[0116] As another example, dense features can be used to determine small regions of foreground objects, and sparse features can be used to detect accurate feature points on the small foreground regions.

[0117] For both sparse features and dense features, incremental data is saved. In the case of sparse features, the incremental data is camera metadata, and in the case of dense features, the incremental data is homography. Typically, during a zoom operation, both cameras may not be available.

[0118] Figure 6 An example disposition of incremental data during a camera transition according to an example embodiment is shown. Figure 6 The example in is based on a transition between a wide-angle camera and a telephoto camera. However, similar means can also be applied to transitions between different pairs of cameras. Example zoom ratios, states, number of states, transition points, etc. are for illustrative purposes only. Such values can vary depending on the device type, camera type, distance between the camera and the object, etc. For example, Figure 6 The "4.2x" is used to mark the zoom ratio at which a transition between a wide-angle camera and a telephoto camera can occur. However, this value can vary depending on the device type, cameras involved in the transition, distance between the camera and the object, etc. Additionally, for example, Figure 6 The zoom ratios (e.g., 2x, 4x, 4.2x, 4.4x) used in the example of are example values and can change depending on the device and / or system configuration.

[0119] For example, a transition from wide-angle to telephoto 605 and a reverse transition from telephoto to wide-angle 610 are shown. Legend 615 indicates the leading camera and the following camera.

[0120] State 1 (initially opened at 1.0x) corresponds to when the camera is first activated. The initial camera can be a wide-angle camera, and the homography is the identity operation. The geometric data of the CPI from the wide-angle camera can be updated as this will not affect the homography.

[0121] State 2 (zoomed in by 2x) corresponds to when the homography gradually changes from the identity operation to the target homography. However, since the dual cameras are not available in State 2, the homography is geometry-based and does not take image features into account. There is no incremental data. In order not to cause any visual distortion during the transition between State 1 and State 2, in some embodiments, the update from the CPI can be damped.

[0122] State 3 (zoomed from 4.0x to 4.2x) corresponds to when both cameras are active simultaneously. In this case, the wide-angle camera is the primary or leading camera, and the telephoto camera is the secondary or following camera. The previously described image-based means can be used to calculate the incremental data, and the updated metadata can be used to calculate the homography.

[0123] State 4 (zoomed from 4.2x to 4.4x) corresponds to after the switch from the wide-angle camera to the telephoto camera occurs. In this case, the telephoto camera is the primary or leading camera, but the wide-angle camera will still be active. No warping is applied, and thus, the homography will be the identity operation. The last calculated increment is stored. Also, the geometric camera metadata of the telephoto camera and the wide-angle camera will be updated.

[0124] In State 5 (beyond 4.4x), the telephoto camera is the primary or leading camera, and the wide-angle camera will be inactive or turned off. Other operations remain the same as in State 4.

[0125] State 6 (back to wide-angle) is similar to State 2. However, now there is incremental data. Therefore, the incremental data is stored, and the homography is calculated using the additional increment.

[0126] State 7 (back to wide-angle) is similar to State 1. However, now there is incremental data. Therefore, the incremental data is stored.

[0127] During consecutive transitions, the process can repeat between States 3 to 7.

[0128] The transition region is the region defined in the smooth transition where two cameras will be active simultaneously within a certain range of zoom ratios such that metadata (such as from OIS / VCM) can be streamed to both cameras simultaneously. For the image-based means described herein, it may be more suitable to configure this transition period to be as large as possible to reduce sudden changes between camera metadata and / or between image-based results and geometry-based results.

[0129] In some embodiments, hardware limitations may make it impractical for two cameras to always stream and / or expand the transition region to an optimal extent. In such embodiments, opportunistic dual streaming may be used. For example, opportunistic dual streaming means that two cameras are opportunistically active simultaneously, which is not based on the zoom ratio but on a timer. For example, after opening the camera application, the timer can be set to 10 seconds, and the two cameras can be active simultaneously every 10 seconds. Based on such a timer, smooth continuous zoom can occur periodically.

[0130] Figure 7 Table 700 shows various situations for switching between a tele camera and a wide-angle camera according to an example embodiment. Column C1 lists the states described with reference to Figure 6 ; column C2 lists the conditions related to the geometric metadata of the wide-angle camera; column C3 lists the conditions related to the geometric metadata of the tele camera; column C4 lists the conditions related to the incremental data; and column C5 lists the conditions related to the homography. Each row (i.e., rows R1 to R7) provides the conditions for each state (states 1 to 7) respectively. Table 700 summarizes the information provided with reference to Figure 6 . For example, row R2 indicates that for state 2, a damping update is applied to the geometric metadata of the wide-angle camera; a regular mapping is used for the geometric metadata of the tele camera; there is no incremental data; and the homography is a combination of a ratio increment and a geometric homography. Other rows present similar information for the corresponding states.

[0131] The term "incremental data" refers to the results of the previously described image-based solution, where the increment is the camera geometric metadata for sparse cases and the incremental homography for dense cases.

[0132] The condition "update" generally indicates a near real-time update according to OIS / VCM. The condition "hold" indicates that the condition is the same as the previous state. The term "damped update" refers to a progressive update that will have a damping ratio between data from a previous frame and data from a current frame. The term "geometry" refers to a geometry-based homography (without image features). The term "increment + geometry" refers to a combined homography of an image-based solution and a geometry-based solution. The term "ratio homography" indicates that the strength of the homography can depend on the zoom ratio (e.g., for state 2 in row R2), the homography strength is identity at 2.0x, and the homography strength will be at full strength at 4.2x. Other ratios between 2.0x and 4.2x can be determined as an interpolation between the identity operation and the full-strength homography.

[0133] Damping for the increment is generally an operation to smooth out sharp changes in geometric data, such as sudden changes in OIS / VCM. The damping ratio can be based on the change in the zoom ratio between two consecutive frames. Similar damping can be applied to the increment.

[0134] Figure 8A An example geometric relationship 800A at each pair of matching pixels according to an example embodiment is depicted. A first plane 805 in a plane is shown as including points , whose coordinates are referenced to an origin . In plane, a second plane 810 corresponds to the first plane 805. A first coordinate system representing plane can be mapped by a mapping to a second coordinate system representing plane, where the coordinate reference origin is , where is a rotation, and is a translation. For example, a point in the first plane 805 is mapped to a point in the second plane 810. In some embodiments, the two-view geometric relationship can be established at each pair of matching pixels using the camera's internal and external parameters, triangulated points from visual matching, and the autofocus distance as the initial scene depth.

[0135] Figure 8B A workflow 800B for determining the geometric relationship at each pair of matching pixels according to an example embodiment is depicted. At block 815, a point (e.g., in the first plane 805) is selected. At block 820, it is determined for a first camera (e.g., with reference to Figure 5 and Figure 6The internal parameters related to the wide - angle camera). At block 825, by using the internal parameters related to the first camera, the 2D points can be back - projected into 3D rays, denoted as Ray. At block 830, depth data can be received. At block 835, based on Ray and the depth data, 3D points for the first camera are determined. At block 840, the external parameters from the first camera are applied to the second camera (e.g., referring to Figure 5 and Figure 6 the tele - camera). At block 845, 3D points for the second camera (corresponding to the 3D points determined at block 835) are determined. At block 850, the internal parameters related to the second camera are determined. At block 855, based on the internal parameters related to the second camera, the re - projection of the point is determined. At block 865, based on the actual position of the point in the second camera obtained at block 860 and the re - projection of the point , one or more re - projection errors are determined.

[0136] Thus, vision - based correction of geometric data at the individual frame level is provided. Workflow 800B achieves a minimized re - projection error for visual correspondence. Based on workflow 800B, the geometry - based homography can be re - estimated through partially - corrected camera calibration and the object distance for focusing to achieve smoothness across frames.

[0137] Figure 9 Depicts an example workflow 900 for smooth continuous zoom in a multi - camera system according to an example embodiment. The continuous zoom frame 905 can include a calibration name file 910, OIS / VCM pairs 915 from two cameras, and a distortion grid configuration 920.

[0138] The algorithms described herein with reference to at least Figure 1 to Figure 8 can be managed by a continuous zoom manager 925. In some embodiments, the continuous zoom manager 925 can include a data trimmer 930, a calibration provider 945, and a homography provider 955. The data trimmer 930 can perform data validation 935 and data dump 940. The calibration provider 945 can provide CPI parameters 950 obtained from a factory calibration file 980.

[0139] The homography provider 955 can determine a geometry - based homography 960 and an image - based homography 965 as described herein. Then, the homography provider 955 can determine a homography compensation 970 and a highlight handling 975.

[0140] Legend 985 indicates the various types of components involved, such as container classes, member functions, member variables, and functional classes. Example data network

[0141] Figure 10 Illustrates a distributed computing architecture 1000 according to an example embodiment. The distributed computing architecture 1000 includes server devices 1008, 1010 configured to communicate with programmable devices 1004a, 1004b, 1004c, 1004d, 1004e via a network 1006. The network 1006 can correspond to a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a wireless wide area network (WWAN), a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 1006 can also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.

[0142] Although Figure 10 only five programmable devices are shown, the distributed application architecture can serve dozens, hundreds, or thousands of programmable devices. Additionally, the programmable devices 1004a, 1004b, 1004c, 1004d, 1004e (or any additional programmable devices) can be any kind of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mounted device (HMD), a network terminal, a mobile computing device, etc. In some examples, as shown by the programmable devices 1004a, 1004b, 1004c, 1004e, the programmable devices can be directly connected to the network 1006. In other examples, as shown by the programmable device 1004d, the programmable device can be indirectly connected to the network 1006 via an associated computing device (such as the programmable device 1004c). In this example, the programmable device 1004c can act as the associated computing device to transfer electronic communications between the programmable device 1004d and the network 1006. In other examples, as shown by the programmable device 1004e, the computing device can be part of and / or inside a vehicle, such as a car, a truck, a bus, a boat or a ship, an airplane, etc. In Figure 10 other examples not shown, the programmable devices can be directly and indirectly connected to the network 1006.

[0143] The server devices 1008, 1010 can be configured to perform one or more services requested by the programmable devices 1004a - 1004e. For example, the server devices 1008 and / or 1010 can provide content to the programmable devices 1004a - 1004e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data, such as compiled software, images, audio, and / or video. The content can include compressed and / or uncompressed content. The content can be encrypted and / or unencrypted. Other types of content are also possible.

[0144] As another example, server device 1008 and / or 1010 can provide programmable devices 1004a - 1004e with access to software for databases, search, computing, graphics, audio, video, World Wide Web / Internet utilization, and / or other functions. Many other examples of server devices are possible. Computing device architecture

[0145] Figure 11 is a block diagram of an example computing device 1100 according to an example embodiment. Specifically, Figure 11 the illustrated computing device 1100 can be configured to perform at least one function of method 1200 and / or at least one function related to method 1200.

[0146] The computing device 1100 can include a user interface module 1101, a network communication module 1102, one or more processors 1103, a data storage 1104, one or more cameras 1118, one or more sensors 1120, and a power system 1122, all of which can be linked together via a system bus, network, or other connection mechanism 1105.

[0147] The user interface module 1101 can be operable to send data to and / or receive data from external user input / output devices. For example, the user interface module 1101 can be configured to send data to and / or receive data from a user input device such as a touch screen, computer mouse, keyboard, keypad, touchpad, trackball, joystick, voice recognition module, and / or other similar devices. The user interface module 1101 can also be configured to provide output to currently known or later developed user display devices such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices. The user interface module 1101 can also be configured to generate audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, headphones, and / or other similar devices. The user interface module 1101 can further be configured with one or more haptic devices that can generate haptic output such as vibrations and / or other outputs detectable by touch and / or physical contact with the computing device 1100. In some examples, the user interface module 1101 can be used to provide a graphical user interface (GUI) for utilizing the computing device 1100.

[0148] The network communication module 1102 may include one or more devices providing one or more wireless interfaces 1107 and / or one or more wired interfaces 1108 configurable to communicate via a network. The one or more wireless interfaces 1107 may include one or more wireless transmitters, receivers, and / or transceivers, such as Bluetooth™ transceivers, Zigbee® transceivers, Wi-Fi™ transceivers, WiMAX™ transceivers, LTE™ transceivers, and / or other types of wireless transceivers configurable to communicate via a wireless network. The one or more wired interfaces 1108 may include one or more wired transmitters, receivers, and / or transceivers, such as Ethernet transceivers, universal serial bus (USB) transceivers, or similar transceivers configurable to communicate via twisted pair, coaxial cable, fiber optic link, or similar physical connections to a wired network.

[0149] In some examples, the network communication module 1102 may be configured to provide reliable, protected, and / or authenticated communication. For each communication described herein, information for facilitating reliable communication (e.g., guaranteed message delivery) may be provided, which may be part of a message header and / or message footer (e.g., packet / message sequencing information, encapsulation headers and / or encapsulation footers, size / time information, and transmission verification information such as cyclic redundancy check (CRC) and / or parity values). One or more cryptographic protocols and / or algorithms may be used to secure (e.g., encode or encrypt) and / or decrypt / decode the communication, such as but not limited to the Data Encryption Standard (DES), Advanced Encryption Standard (AES), Rivest-Shamir-Adelman (RSA) algorithm, Diffie-Hellman algorithm, Secure Sockets Protocol (e.g., Secure Sockets Layer (SSL) or Transport Layer Security (TLS)), and / or Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms may be used to secure (and then decrypt / decode) the communication in addition to those listed herein.

[0150] The one or more processors 1103 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPU), graphics processing units (GPU), application specific integrated circuits, etc.). The one or more processors 1103 may be configured to execute computer-readable instructions 1106 contained in the data storage 1104 and / or other instructions as described herein.

[0151] The data storage 1104 may include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of the one or more processors 1103. The one or more computer-readable storage media may include volatile and / or non-volatile storage components that can be integrated, in whole or in part, with at least one of the one or more processors 1103, such as optical, magnetic, organic, or other memory or disk storage. In some examples, the data storage 1104 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other examples, the data storage 1104 may be implemented using two or more physical devices.

[0152] The data storage 1104 may include computer-readable instructions 1106 and possibly additional data. In some examples, the data storage 1104 may include the storage required to execute at least a portion of the methods, scenarios, and techniques described herein and / or at least a portion of the functionality of the apparatuses and networks described herein. In some examples, the data storage 1104 may include storage for the warping transformation module 1112 (e.g., a module that calculates geometric-based homographies, image-based homographies, etc.). Specifically, in these examples, the computer-readable instructions 1106 may include instructions that, when executed by the one or more processors 1103, enable the computing device 1100 to provide some or all of the functionality of the warping transformation module 1112.

[0153] In some examples, the computing device 1100 may include one or more cameras 1118. The cameras 1118 may include one or more image capture devices, such as still and / or video cameras, that are configured to capture light and record the captured light in one or more images; that is, the cameras 1118 may generate images of the captured light. The one or more images may be one or more still images and / or one or more images used in video footage. The cameras 1118 may capture light and / or electromagnetic radiation that is emitted as visible light, infrared radiation, ultraviolet light, and / or as light at one or more other frequencies. The cameras 1118 may include wide-angle cameras, telephoto cameras, ultra-wide-angle cameras, etc. Additionally, for example, the cameras 1118 may be the front-facing camera or the rear-facing camera of the reference computing device 1100.

[0154] In some examples, computing device 1100 may include one or more sensors 1120. The sensors 1120 may be configured to measure conditions within computing device 1100 and / or conditions in the environment of computing device 1100 and provide data regarding such conditions. For example, sensors 1120 may include one or more of the following: (i) sensors for obtaining data regarding computing device 1100, such as but not limited to a thermometer for measuring the temperature of computing device 1100, a battery sensor for measuring the power of one or more batteries of power system 1122, and / or other sensors for measuring conditions of computing device 1100; (ii) identification sensors for identifying other objects and / or devices, such as but not limited to a radio frequency identification (RFID) reader, a proximity sensor, a one-dimensional barcode reader, a two-dimensional barcode (e.g., Quick Response (QR) code) reader, and a laser tracker, where the identification sensors may be configured to read identifiers, such as RFID tags, barcodes, QR codes, and / or other devices and / or objects configured to be read and provide at least identification information; (iii) sensors for measuring the position and / or movement of computing device 1100, such as but not limited to an inclinometer, a gyroscope, an accelerometer, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser displacement sensor, and a compass; (iv) environmental sensors for obtaining data indicative of the environment of computing device 1100, such as but not limited to an infrared sensor, an optical sensor, a light sensor, a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a motion sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (v) force sensors for measuring one or more forces acting around computing device 1100 (e.g., inertial forces and / or G-forces), such as but not limited to one or more sensors for measuring one or more of the following: force in one or more dimensions, torque, ground force, friction, and / or a zero moment point (ZMP) sensor for identifying and / or the location of the ZMP. Many other examples of sensors 1120 are possible.

[0155] The power supply system 1122 may include one or more batteries 1124 and / or one or more external power interfaces 1126 for supplying power to the computing device 1100. Each of the one or more batteries 1124 can act as a source of stored power for the computing device 1100 when electrically coupled to the computing device 1100. The one or more batteries 1124 of the power supply system 1122 can be configured to be portable. Some or all of the one or more batteries 1124 can be easily removable from the computing device 1100. In other examples, some or all of the one or more batteries 1124 can be inside the computing device 1100 and thus may not be easily removable from the computing device 1100. Some or all of the one or more batteries 1124 can be rechargeable. For example, a rechargeable battery can be recharged via a wired connection between the battery and another power supply, such as by one or more power supplies outside the computing device 1100 and connected to the computing device 1100 via one or more external power interfaces. In other examples, some or all of the one or more batteries 1124 can be non-rechargeable batteries.

[0156] One or more external power interfaces 1126 of the power supply system 1122 can include one or more wired power interfaces, such as a USB cable and / or a power cord, which implement a wired power connection to one or more power supplies located outside the computing device 1100. One or more external power interfaces 1126 can include one or more wireless power interfaces, such as a Qi wireless charger, which implement a wireless power connection to one or more external power supplies, such as via a Qi wireless charger. Once a power connection to an external power supply is established using one or more external power interfaces 1126, the computing device 1100 can draw power from the external power supply via the established power connection. In some examples, the power supply system 1122 can include associated sensors, such as battery sensors associated with one or more batteries or other types of power sensors. Example operating methods

[0157] Figure 12 Method 1200 according to an example embodiment is shown. Method 1200 can include various blocks or steps. The blocks or steps can be performed individually or in combination. The blocks or steps can be performed in any order and / or serially or in parallel. Further, blocks or steps can be omitted or blocks or steps can be added to method 1200.

[0158] The blocks of method 1200 can be performed by various elements of the computing device 1100 as shown and described with reference to Figure 11 as shown.

[0159] Frame 1210 includes: displaying, on a display screen of a computing device, an initial preview of a scene captured by a first image capture device of the computing device, where the first image capture device is operating within a first focal length range.

[0160] Frame 1220 includes: detecting, by the computing device, a zoom operation that is predicted to cause the first image capture device to reach the limit of the first focal length range.

[0161] Frame 1230 includes: in response to the detection, activating a second image capture device of the computing device to capture a zoomed preview of the scene, where the second image capture device is configured to operate within a second focal length range.

[0162] Frame 1240 includes: updating a geometry-based distortion transformation based on a comparison of corresponding image features from the initial preview and the zoomed preview.

[0163] Frame 1250 includes: aligning the zoomed preview with the initial preview by applying the updated distortion transformation, where the updated distortion transformation reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed preview.

[0164] Frame 1260 includes: displaying, on the display screen of the computing device, the aligned zoomed preview of an image captured by the second image capture device while operating within the second focal length range.

[0165] In some embodiments, the comparison of corresponding image features includes detecting one or more visual features in the initial preview and the zoomed preview. Such embodiments also include: generating a visual correspondence between the initial preview and the zoomed preview based on the one or more visual features.

[0166] Some embodiments include: optimizing the detection of one or more visual features and the generation of visual correspondence by performing asynchronous multi-threaded processing that includes receiving one or more images and associated metadata as input and sending visual feature matches and associated metadata as output.

[0167] In some embodiments, the update of the geometry-based distortion transformation includes correcting frame-based geometry metadata based on the visual correspondence.

[0168] In some embodiments, the update of the geometry-based distortion transformation includes estimating a homography based on the corrected geometry metadata, and where the homography maps pixels in a plane of a first coordinate system associated with the first image capture device to corresponding pixels in the same plane of a second coordinate system associated with the second image capture device.

[0169] In some embodiments, the update of the geometry-based warping transformation utilizes frame-based data that includes one or more of an image, a pre-crop of the image, scene depth, or calibration parameters associated with a first image capture device and a second image capture device, respectively. In some embodiments, the calibration parameters include an autofocus distance.

[0170] In some embodiments, the application of the updated warping transformation is performed on each frame of the initial preview and the corresponding frame of the zoomed preview in a side-by-side comparison.

[0171] In some embodiments, the alignment of the zoomed preview with the initial preview includes aligning the depth value of a point in the image space with the geometric focus distance of that point on each frame of the initial preview and the corresponding frame of the zoomed preview.

[0172] Some embodiments include: generating a bundle adjustment to be applied to one or more camera calibrations and one or more focal distances for each frame of the initial preview and the corresponding frame of the zoomed preview.

[0173] Some embodiments include: generating a modified bundle adjustment based on the corresponding bundle adjustments of a series of consecutive frames for a series of consecutive frames.

[0174] Some embodiments include: transitioning from a first image capture device to a second image capture device by a computing device based on the updated warping transformation.

[0175] In some embodiments, the second focal length range can be greater than or less than the first focal length range, which corresponds to a zoom-in or zoom-out operation on the computing device.

[0176] In some embodiments, one or more viewing artifacts include binocular disparity.

[0177] In some embodiments, the update of the geometry-based warping transformation includes reducing jitter by applying temporal feature matching and tracking.

[0178] The specific arrangements shown in the drawings should not be considered restrictive. It should be understood that other embodiments may include more or fewer of each element shown in a given drawing. Further, some of the elements shown may be combined or omitted. Still further, example embodiments may include elements not shown in the figures.

[0179] Steps or blocks representing the processing of information may correspond to circuitry that can be configured to perform the described methods or techniques. Alternatively or additionally, steps or blocks representing the processing of information may correspond to a module, a segment, or a portion of program code, including related data. The program code may include one or more instructions that can be executed by a processor to implement a specific logical function or action in the method or technique. The program code and / or related data may be stored on any type of computer-readable medium, such as a storage device including a disk or hard drive or other storage medium.

[0180] The computer-readable medium may also include non-transitory computer-readable media, such as computer-readable media that store data for a short period of time, such as register memory, processor cache, and random access memory (RAM). The computer-readable medium may also include non-transitory computer-readable media that store program code and / or data for a longer period of time. Thus, the computer-readable medium may include auxiliary or persistent long-term storage, such as, for example, read-only memory (ROM), optical disks or magnetic disks, compact disc read-only memory (CD-ROM). The computer-readable medium may also be any other volatile or non-volatile storage system. The computer-readable medium may be considered, for example, a computer-readable storage medium or a tangible storage device.

[0181] Although various examples and embodiments have been disclosed, other examples and embodiments will be apparent to those skilled in the art. The various examples and embodiments disclosed are for illustrative purposes and are not intended to be limiting, where the true scope is indicated by the appended claims.

Claims

1. A computer-implemented method, comprising: displaying, on a display screen of a computing device, an initial preview of a scene captured by a first image capture device of the computing device, wherein the first image capture device is operating within a first focal length range; detecting, by the computing device, a zoom operation that is predicted to cause the first image capture device to reach a limit of the first focal length range; in response to the detection, activating a second image capture device of the computing device to capture a zoomed preview of the scene, wherein the second image capture device is configured to operate within a second focal length range; updating a geometry-based distortion transform based on a comparison of corresponding image features from the initial preview and the zoomed preview; aligning the zoomed preview with the initial preview by applying the updated distortion transform, wherein the updated distortion transform reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed preview; and displaying, on the display screen of the computing device, the aligned zoomed preview of the image captured by the second image capture device while operating within the second focal length range.

2. The method according to claim 1, wherein The comparison of the corresponding image features further comprises: detecting one or more visual features in the initial preview and the zoomed preview; and generating a visual correspondence between the initial preview and the zoomed preview based on the one or more visual features.

3. The method according to claim 2, further comprising: optimizing the detection of the one or more visual features and the generation of the visual correspondence by performing asynchronous multi-threaded processing, the asynchronous multi-threaded processing comprising receiving one or more images and associated metadata as input and sending visual feature matches and associated metadata as output.

4. The method according to claim 2, wherein, The updating of the geometry-based distortion transform comprises: correcting frame-based geometry metadata based on the visual correspondence.

5. The method according to claim 4, wherein The updating of the geometry-based distortion transform comprises estimating a homography based on the corrected geometry metadata, and wherein the homography maps pixels in a plane of a first coordinate system associated with the first image capture device to corresponding pixels in the same plane of a second coordinate system associated with the second image capture device.

6. The method according to claim 1, wherein The updating of the geometry-based distortion transform utilizes frame-based data, the frame-based data comprising one or more of an image, a pre-crop of the image, scene depth, or calibration parameters respectively associated with the first image capture device and the second image capture device.

7. The method according to claim 6, wherein, The calibration parameters include an auto-focus distance.

8. The method according to claim 1, wherein The application of the updated distortion transform is performed on each frame of the initial preview and the corresponding frame of the zoomed preview in a side-by-side comparison.

9. The method according to claim 1, wherein The alignment of the zoomed preview with the initial preview comprises: aligning a depth value of a point in image space with a geometric focus distance of the point on each frame of the initial preview and the corresponding frame of the zoomed preview.

10. The method according to claim 1, further comprising: For each frame of the initial preview and the corresponding frame of the zoomed preview, generate a bundle adjustment to be applied to one or more camera calibrations and one or more focal distances.

11. The method according to claim 10, further comprising: For a series of consecutive frames, generate a modified bundle adjustment based on the corresponding bundle adjustments of the consecutive frames.

12. The method according to claim 1, further comprising: The computing device transitions from the first image capture device to the second image capture device based on the updated distortion transformation.

13. The method according to claim 1, wherein The updating of the geometric-based distortion transformation further comprises: Reducing jitter by applying temporal feature matching and tracking.

14. A computing device, comprising: A display screen; A first image capture device configured to operate within a first focal length range; A second image capture device configured to operate within a second focal length range; One or more processors; And A data storage, wherein computer-executable instructions are stored on the data storage, and when executed by the one or more processors, the computer-executable instructions cause the mobile device to perform functions, the functions including: Display an initial preview of the scene captured by the first image capture device on the display screen; The computing device detects a zoom operation that may cause the first image capture device to reach the limit of the first focal length range; In response to the detection, activate the second image capture device to capture a zoomed preview of the scene; Update a geometric-based distortion transformation based on a comparison of corresponding image features from the initial preview and the zoomed preview; Align the zoomed preview with the initial preview by applying the updated distortion transformation, wherein the updated distortion transformation reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed preview; and The computing device displays the aligned zoomed preview of the image captured by the second image capture device while operating within the second focal length range on the display screen of the computing device.

15. The computing device according to claim 14, wherein, The function for the comparison of the corresponding image features further comprises: Detect one or more visual features in the initial preview and the zoomed preview; and Based on the one or more visual features, generate a visual correspondence between the initial preview and the zoomed preview.

16. The computing device according to claim 15, wherein, The function for the updating of the geometric-based distortion transformation further comprises: Correct frame-based geometric metadata based on the visual correspondence.

17. The computing device according to claim 16, wherein, The function for the updating of the geometric-based distortion transformation includes estimating a homography based on the corrected geometric metadata, and wherein the homography maps pixels in a plane of a first coordinate system associated with the first image capture device to corresponding pixels at the same plane of a second coordinate system associated with the second image capture device.

18. The computing device according to claim 14, wherein, The update of the geometry-based warping transformation utilizes frame-based data, the frame-based data including one or more of images respectively associated with the first image capture device and the second image capture device, pre-cropping of the images, scene depth, or calibration parameters.

19. The computing device according to claim 14, wherein, The function for the application of the updated warping transformation is performed on each frame of the initial preview and the corresponding frame of the zoomed preview in the side-by-side comparison.

20. The computing device according to claim 19, wherein, The function for the alignment of the zoomed preview with the initial preview includes: Aligning the depth value of a point in image space with the geometric focus distance of the point on each frame of the initial preview and the corresponding frame of the zoomed preview.

21. The computing device according to claim 14, wherein, The function for the update of the geometry-based warping transformation further includes: Reducing jitter by applying temporal feature matching and tracking.

22. A non-transitory computer-readable medium, the non-transitory computer-readable medium including program instructions that can be executed by one or more processors to cause the one or more processors to perform operations, the operations including: Displaying an initial preview of a scene captured by the first image capture device on the display screen; Detecting, by the computing device, a zoom operation that may cause the first image capture device to reach the limit of the first focal length range; In response to the detection, activating the second image capture device to capture a zoomed preview of the scene; Updating a geometry-based warping transformation based on a comparison of corresponding image features from the initial preview and the zoomed preview; Aligning the zoomed preview with the initial preview by applying the updated warping transformation, wherein the updated warping transformation reduces one or more viewing artifacts caused by a change in the field of view when transitioning from the initial preview to the zoomed preview; And Displaying, on the display screen of the computing device, the aligned zoomed preview of the image captured by the second image capture device while operating within the second focal length range.

Citation Information

Cited By

  • Data frame processing method and device, terminal, base station and chip

    CN121357420A