Stabilized object tracking at high magnification
By applying machine learning technology and sensor data in the image capture device, an adjusted preview of the zoomed part is generated, and the problem of stable tracking of objects of interest at high magnification is solved, achieving the stability of image capture and user experience.
Patent Information
- Application Number
- CN202280101547.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-04
- Publication Date
- 2025-05-30
AI Technical Summary
At high magnification, it is difficult for the image capture device to stabilize frame and track the object of interest in a narrow field of view, and the existing image stabilization technology has hardware limitations and noise problems, resulting in a decrease in residual motion and frame stability.
By introducing machine learning techniques, especially convolutional neural networks, in the image capture device, combined with sensor data and optical flow information, an adjusted preview of the zoomed part is generated to ensure that the area of interest is displayed stably in the center of the field of view or near the center of the field of view.
It realizes smooth tracking and stable display of objects of interest at high magnification, overcomes hardware noise and residual motion problems, and improves the stability and user experience of image capture.
Smart Images

Figure CN120077669A_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] Many modern computing devices, including mobile phones, personal computers, and tablet computers, include image capture devices. Some image capture devices are configured with telephoto capabilities. SUMMARY OF THE INVENTION
[0002] The present disclosure generally relates to stabilization of an image in a viewfinder of an image capture device at a high magnification ratio. In one aspect, the image capture device may be configured to frame and track an object of interest in a narrow field of view produced by the high magnification. With the support of a system of machine-learned components, the image capture device may be configured to stabilize and maintain the frame.
[0003] In a first aspect, a computer-implemented method is provided. The method includes displaying, on a display screen of an image capture device, a preview of an image representing a field of view of the image capture device. The method further includes determining a region of interest in the preview of the image. The method further includes: transitioning the image capture device from a normal operation mode to a zoomed operation mode, wherein the zoomed operation mode includes: determining a motion trajectory of the region of interest based on sensor data collected by a sensor associated with the image capture device; and generating an adjusted preview representing a zoomed portion of the field of view based on the determined motion trajectory, wherein the adjusted preview displays the region of interest at or near a center of the zoomed portion. The method additionally includes providing, by the display screen, the adjusted preview of the portion of the field of view.
[0004] In a second aspect, a device is provided. The device includes one or more processors operable to perform operations. The operations include displaying, on a display screen of an image capture device, a preview of an image representing a field of view of the image capture device. The operations further include determining a region of interest in the preview of the image. The operations further include: transitioning the image capture device from a normal operation mode to a zoomed operation mode, wherein the zoomed operation mode includes: determining a motion trajectory of the region of interest based on sensor data collected by a sensor associated with the image capture device; and generating an adjusted preview representing a zoomed portion of the field of view based on the determined motion trajectory, wherein the adjusted preview displays the region of interest at or near a center of the zoomed portion. The operations additionally include providing, by the display screen, the adjusted preview of the portion of the field of view.
[0005] In a third aspect, an article is provided. The article may include a non-transitory computer-readable medium having program instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to perform operations. The operations include displaying, on a display screen of an image capture device, a preview of an image representing a field of view of the image capture device. The operations further include determining a region of interest in the preview of the image. The operations further include: transitioning the image capture device from a normal operation mode to a zoomed operation mode, wherein the zoomed operation mode includes: determining a motion trajectory of the region of interest based on sensor data collected by a sensor associated with the image capture device; and generating, based on the determined motion trajectory, an adjusted preview representing a zoomed portion of the field of view, wherein the adjusted preview displays the region of interest at or near a center of the zoomed portion. The operations additionally include providing, by the display screen, the adjusted preview of the portion of the field of view.
[0006] In a fourth aspect, a system is provided. The system includes: means for displaying, on a display screen of an image capture device, a preview of an image representing a field of view of the image capture device; means for determining a region of interest in the preview of the image; means for transitioning the image capture device from a normal operation mode to a zoomed operation mode, wherein the zoomed operation mode includes: means for determining a motion trajectory of the region of interest based on sensor data collected by a sensor associated with the image capture device; and means for generating, based on the determined motion trajectory, an adjusted preview representing a zoomed portion of the field of view, wherein the adjusted preview displays the region of interest at or near a center of the zoomed portion; and means for providing, by the display screen, the adjusted preview of the portion of the field of view.
[0007] Other aspects, embodiments, and implementations will become apparent to those of ordinary skill in the art upon reading the following detailed description with appropriate reference to the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a diagram showing an adjusted preview of a portion of a field of view according to an example embodiment.
[0009] Figure 2 is a diagram showing a warning notification for stabilized object tracking according to an example embodiment.
[0010] Figure 3A is an example workflow for stabilized object tracking according to an example embodiment.
[0011] Figure 3B is an example workflow for applying zoom stabilization according to an example embodiment.
[0012] Figure 4Is an example workflow for processing consecutive frames in a hybrid tracker according to an example embodiment.
[0013] Figure 5 Depicts an example tracking optimization process according to an example embodiment.
[0014] Figure 6 Depicts another example tracking optimization process according to an example embodiment.
[0015] Figure 7 Depicts another example tracking optimization process according to an example embodiment.
[0016] Figure 8 Depicts an example tracking optimization process for two regions of interest according to an example embodiment.
[0017] Figure 9 Shows an example image with stabilized object tracking according to an example embodiment.
[0018] Figure 10 Shows another example image with stabilized object tracking according to an example embodiment.
[0019] Figure 11 Shows another example image with stabilized object tracking according to an example embodiment.
[0020] Figure 12 Is a diagram showing the training phase and inference phase of a machine learning model according to an example embodiment.
[0021] Figure 13 Depicts a distributed computing architecture according to an example embodiment.
[0022] Figure 14 Is a block diagram of a computing device according to an example embodiment.
[0023] Figure 15 Depicts a network of a computing cluster arranged as a cloud-based server system according to an example embodiment.
[0024] Figure 16 Is a flowchart of a method according to an example embodiment. Detailed Description
[0025] Example methods, apparatuses, and systems are described herein. It should be understood that the terms "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration". Any embodiment or feature described herein as "example" or "exemplary" is not necessarily to be construed as preferred or superior to other embodiments or features. Other embodiments may be utilized and other changes may be made without departing from the scope of the subject matter presented herein.
[0026] Accordingly, the example embodiments described herein are not intended to be limiting. Aspects of the present disclosure, as generally described herein and shown in the accompanying drawings, can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are contemplated herein.
[0027] Further, unless the context otherwise implies, features shown in each of the drawings can be used in combination with one another. Accordingly, the drawings are generally to be regarded as constituent aspects of one or more overall embodiments, but it is understood that not all of the features shown are necessary for each embodiment. Overview
[0028] This application relates to image stabilization using machine learning techniques such as, but not limited to, neural network techniques. In the case where a user of an image capture device previews an image at a high magnification, the resulting image may be unstable, and there may be challenges in framing an object of interest in the image. Further, for example, after framing the object of interest, the small field of view (FOV) resulting from the high magnification may pose additional challenges in maintaining a moving frame in a smooth, continuous, and / or stable manner. Accordingly, there arise the following technical problems related to image processing, which relate to stabilizing the object of interest in the preview and maintaining smooth movement of the object of interest within the moving frame. Further, for example, there arise the following technical problems related to image processing, which relate to smoothly transitioning the operation of the image capture device between different modes (e.g., corresponding to different magnifications).
[0029] Telephoto cameras are becoming increasingly popular in flagship devices. Higher and higher optical zoom lenses combined with higher resolution image sensors have been used to increase the maximum magnification in each successive device release. Image quality at high magnifications has been continuously improved. However, the extremely narrow FOV at high magnifications (e.g., an FOV of approximately at a magnification of ) makes it challenging to frame an object of interest within the FOV, and particularly challenging to maintain such framing while pressing the shutter button to capture an image. This can lead to a "hide-and-seek" game as the user tries to find the object of interest in the narrow field of view. Although existing optical image stabilization (OIS) and / or electronic image stabilization (EIS) algorithms attempt to improve this situation to a limited extent, there remains residual motion that is magnified at higher magnifications.
[0030] For example, where OIS is limited to approximately In the case of, and in the absence of EIS, it is possible to have "what you see is what you get" (WYSIWYG); however, it can be challenging to maintain the region of interest (ROI) within the frame, and it can also be challenging to find the ROI in a zoomed frame. Additionally, in the case of applying EIS (e.g., limited by noise caused by gyroscopes and / or sensors), the WYSIWYG feature may be lost, and a stuttering effect may appear in the image. Therefore, some smart devices configured with cameras attempt to address this issue by reducing the maximum magnification in video mode compared to photo mode.
[0031] Generally, baseline EIS can compensate for hand shake without the need for tracking capabilities. However, it can be challenging to keep a moving object within the viewfinder at high magnifications. Additionally, for example, the trajectory of a moving object may not be smooth.
[0032] EIS based on gyroscopes and / or OIS can be used as a stabilization technique in some cases. This stabilization technique is sensor-based, such as gyroscope sensing and / or OIS sensing. While this can compensate for camera pose changes without relying on image content, it may result in limiting the field of view (FOV) of the stabilized frame of the output, as the edges are used to generate a stabilized virtual pose. At high magnifications, hardware limitations such as gyroscope noise, OIS sensing noise, OIS calibration error, signal latency, etc. may also introduce visible residual motion.
[0033] In addition to the above EIS techniques, some techniques may also involve detecting the position of a face. While this approach can produce stabilized frames that move along with the face, this technique is limited to faces and not general objects of interest. Additionally, like the EIS approach, since the frame is cropped after stabilization, this method can cause the FOV of the output stabilized frame to be limited.
[0034] The techniques described herein address these challenges by enabling smooth photo and / or video capture and an easy framing experience at high magnifications. This is achieved by reliably tracking the ROI and locking it at or near the center of the viewfinder or by maintaining smooth movement within the frame. Gyroscope and / or OIS sensor information is used to determine the motion trajectory, but image-based ROI tracking effectively overcomes hardware limitations (such as gyroscope noise, OIS sensing noise, OIS calibration error, signal latency, etc.).
[0035] The techniques described herein may include aspects that combine image-based techniques with techniques based on motion data and optical image stabilization (OIS) data. Neural networks such as convolutional neural networks can be trained and applied to perform one or more aspects described herein. In some examples, the neural network may be arranged as an encoder / decoder neural network.
[0036] This method is for high magnification and uses digital zoom to crop the edge part for stabilization without any additional cropping. The zoom stabilization pipeline can be configured to determine the saliency map and object tracker with reference to the entire image sensor area or the ROI center via the cropped sensor area, such that stabilization can be achieved along the ROI as long as the ROI is within the image sensor. As described herein, saliency detection, object tracking, and optical flow are used jointly to propose the ROI to be stabilized. The potential frame delay caused by the camera pipeline depth (e.g., frames) is also addressed. A new user interface (UI) and / or user experience (UX) design is also described, which enables a frame-in-frame viewfinder when the zoom stabilization mode is active, and the bounding box moves relative to the frame to indicate the stabilized and magnified FOV relative to the sensor area or the entire FOV. The techniques described also maintain the same FOV defined by the user without sacrificing additional edge parts to stabilize the frame.
[0037] In one example, a copy of a trained neural network for detecting salient objects may reside on a mobile computing device. The mobile computing device may include a camera that can capture an input photo or video. A user of the mobile computing device may view the input photo or video and determine that an object in the input photo or video should be tracked. The input photo or video and motion data may be provided to the trained neural network residing on the mobile computing device. In response, the trained neural network may generate a predicted output showing the ROI with a bounding box. In other examples, the trained neural network does not reside on the mobile computing device; instead, the mobile computing device (e.g., via the Internet or another data network) provides the input photo or video and motion data to a trained neural network located remotely. The remotely located neural network may process the input photo or video and motion data as indicated above and provide an output. In other examples, non-mobile computing devices may also use the trained neural network to stabilize object tracking in images and videos at high magnification - including photos or videos not captured by the camera of the computing device.
[0038] Thus, the techniques described herein can improve image capture devices by stabilizing images and providing a closer view, thereby enhancing the actual and / or perceived quality of the images. The enhancement of the actual and / or perceived quality of photos or videos can provide user experience benefits. These techniques are flexible and can thus be applied to a wide variety of videos in both indoor and outdoor environments. Techniques for image stabilization using neural networks
[0039] One of the main features of image stabilization is to maintain an accurate and reliable tracker for the region of interest (ROI). Minimizing tracking noise and reducing noise caused by the gyroscope and / or OIS can significantly contribute to image stabilization. Additional challenges include pose changes, occlusions, and objects moving into and / or out of the sensor area. There may also be latency issues related to the delay between image processing at the hardware layer and subsequent changes at the software or application layer. For example, the pipeline depth (e.g., five frames) may cause a delay. For mobile camera applications, additional image cropping may not be allowed in photo mode. In addition, the stabilization needs to smoothly transition between multiple modes of the camera.
[0040] Thus, as described herein, the zoom stabilization mode of a mobile device can capture smooth photos and / or videos at high magnifications by reliably tracking the user's region of interest (ROI) (or object of interest) and locking it at the center of the field of view (FOV) or smoothly moving the ROI within the frame to simplify the user's framing experience.
[0041] Noise from the gyroscope and / or OIS or from calibration errors can affect smooth tracking. There may also be challenges due to the complex integration of the camera application with the underlying hardware layer, saliency nodes, rectiface, and / or EIS nodes, etc. In addition, for example, the image quality can be low at high magnifications, especially in the remosaic mode. Thus, the zoom stabilization mode can also be configured to support binning transitions. In some embodiments, in strong light conditions (e.g., outdoors), the remosaic mode can result in higher image quality due to higher resolution compared to the binning mode. However, in low light conditions, considering the poor noise performance in the remosaic mode, the binning mode can result in higher image quality.
[0042] As described herein, zoom stabilization can have several technical advantages, such as limited power and low latency budget for real-time photo preview. Zoom stabilization can also be configured for multi-object handling (e.g., flocks of animals). In addition, for example, zoom stabilization is compatible with existing features (e.g., HDR+, long-range video, etc.).
[0043] Figure 1 is a diagram showing an adjusted preview of a portion of the field of view according to an example embodiment. In some embodiments, the display screen of an image capture device can display a preview of an image representing the field of view of the image capture device. For example, the display screen 100A can display the field of view 105, which may include an image of an object of interest, such as a crescent moon 110. When the image capture device is operating at a high magnification, the field of view 105 can be very narrow, and small hand movements can cause the crescent moon 110 to fall outside the field of view 105.
[0044] Some embodiments relate to determining a region of interest in a preview of an image. For example, there may be no ROI within the field of view, or the ROI may have moved out of the field of view. In such embodiments, the background motion within the field of view can be tracked. In some embodiments, a new ROI can be detected within the field of view. For example, an ROI tracker can identify a new object. An important feature of the ROI tracker is to reliably predict what the user of the camera is trying to capture. This can be achieved individually or in combination with user indications and machine learning-based algorithms. For example, a Tap ROI tracker in a camera application can enable a user to tap on the display screen and indicate an object of interest and / or a region of interest. Additionally, for example, a machine learning model can be used to generate a saliency map, where the saliency map indicates the user's region of interest. Although existing saliency maps output a bounding box of a fixed size for the ROI, the zoom stabilization described herein is configured to estimate the size of the ROI and output an appropriate bounding box for the ROI. In some embodiments, the confidence level of the ROI may be low. In such embodiments, background motion, the central ROI, or a combination of both can be used to maintain smooth framing.
[0045] Some embodiments relate to transitioning an image capture device from a normal operation mode to a zoomed operation mode. For example, an image can be captured at different zoom levels. At high magnifications, the field of view can be considerably narrower, and small movements of the camera can cause a sudden change in the image being captured by the field of view. In some embodiments, a threshold magnification can be used to determine whether an image stabilization algorithm such as the zoom stabilization algorithm described herein needs to be turned on or off. For example, some cameras can be configured such that the mode switch occurs at a magnification of. For example, the zoom stabilization mode can be turned on for magnifications greater than and turned off for magnifications less than Additional and / or alternative magnifications can be utilized.
[0046] In some embodiments, the zoom operation mode may involve determining an adjusted preview representing a zoomed portion of the field of view based on sensor data collected by a sensor associated with the image capture device, where the adjusted preview displays the region of interest at or near the center of the zoomed portion. As shown in display screen 100B, the field of view 105 may be displayed within a view defined by an outer frame 140. An inner frame 145 may be displayed within the outer frame 140. The inner frame 145 may include the object of interest, i.e., the crescent 110. Thus, the display screen 100B may display an adjusted preview 115 (e.g., a magnified view) that represents a zoomed portion (e.g., within the inner frame 145) of the field of view (e.g., the field of view 105 displayed within the outer frame 140). As shown, the display screen 100B displays a magnified view 115 of a portion of the field of view, including a magnified view of the crescent 110A. To the extent that the image of the crescent 110 may display a range of non-smooth motion within the original field of view 105, the inner frame 145 is a stabilized image, and the magnified view 115 is a stabilized and zoomed view of the crescent 110A.
[0047] The display screen 100B may include additional features related to the camera application. For example, multiple modes may be available to the user, including a motion mode 120, a portrait mode 122, a video mode 126, and a video bokeh mode 128. As shown, the camera application may be in a camera mode 124. The camera mode 124 may provide additional features such as a reverse icon 130 for activating a reverse camera view, a trigger button 132 for capturing a preview image, and a photo stream icon 134 for accessing a database of captured images. As another example, a magnification slider 138 may be displayed, and the user may move a virtual object along the magnification slider 138 to select a magnification. In some embodiments, the user may use the display screen to adjust the magnification (e.g., by moving two fingers away from each other in an outward motion on the display screen 100B), and the magnification slider 138 may automatically display the magnification.
[0048] As indicated, the magnification slider 138 may be at . For a camera configured to switch modes at a magnification, in the case where the magnification slider 138 moves beyond , the camera may switch from a normal mode to a zoom stabilization mode and may automatically activate image stabilization. In such instances, the object of interest may be determined, and the magnified view 115, the outer frame 140, the inner frame 145, etc. may be displayed.
[0049] The camera application can also provide various user-adjustable features to adjust one or more image characteristics (e.g., brightness, hue, contrast, shadows, highlights, global brightness adjustment of the entire image, local brightness adjustment of the ROI, etc.). For example, in some embodiments, a slider 136A can be provided to adjust characteristic A, a slider 136B can be provided to adjust characteristic B, and in an embodiment, a slider 136C can be provided to adjust characteristic C.
[0050] In some embodiments, the zoom operation mode can involve: determining a motion trajectory of a region of interest based on sensor data collected by a sensor associated with the image capture device; and generating an adjusted preview representing a zoomed portion of the field of view based on the determined motion trajectory, wherein the adjusted preview displays the region of interest at or near the center of the zoomed portion.
[0051] Figure 2 FIG. is a diagram showing a warning notification for stabilized object tracking according to an example embodiment. The display screen 200A (respectively, display screen 200B) can include additional features related to the camera application. For example, multiple modes can be available to the user, including a motion mode 120, a portrait mode 122, a video mode 126, and a video bokeh mode 128. As shown, the camera application can be in a camera mode 124. The camera mode 124 can provide additional features, such as a reverse icon 130 for activating a reverse camera view, a trigger button 132 for capturing a preview image, and a photo stream icon 134 for accessing a database of captured images. Also for example, a magnification slider 138 can be displayed, and the user can move a virtual object along the magnification slider 138 to select a magnification. In some embodiments, the user can use the display screen 200A (respectively, display screen 200B) to adjust the magnification (e.g., by moving two fingers away from each other in an outward motion on the display screen 200A (respectively, display screen 200B)), and the magnification slider 138 can automatically display the magnification.
[0052] As indicated, the magnification slider 138 can be at . For a camera configured to switch modes at a magnification of , when the magnification slider 138 moves to , the camera can switch from a normal mode to a zoom stabilization mode, and image stabilization can be automatically activated. In such instances, an object of interest can be determined, and a magnified view 115, an outer frame 140, an inner frame 145, etc. can be displayed. The zoom ratio is for illustrative purposes only and can vary depending on the device and / or system configuration.
[0053] At high magnification, operating the camera to track a moving object may result in a stop-move-stop type of motion, resulting in large residual motion under traditional EIS. This can be caused, for example, by the prominent term as described in Equation 1 below. Some embodiments include providing an image overlay by the display screen, the image overlay showing a representation of the zoomed portion relative to the field of view. For example, the frame within frame stabilizes the image and guides the user. Some embodiments include determining a bounding box of the region of interest, and wherein providing the image overlay includes providing the region of interest framed within the bounding box. As shown, the display screen 200A shows the field of view within the outer frame 140, and the inner frame 145 within the outer frame 140 frames the ROI, i.e., the image of the moon 150. The adjusted preview corresponding to the inner frame 145 (such as the magnified view 115) is Figure 1 displayed together with the magnified view of the moon 150. As shown, the image of the moon 150 is centered within the inner frame 145.
[0054] Some embodiments include detecting that the region of interest is approaching the boundary of the image overlay. For example, as the camera moves, and / or due to the movement of the object, the image of the moon 150 can move within the field of view. In such embodiments, the zoom stabilization mode is capable of maintaining the stabilized magnified view 115, wherein the image of the moon is centered within the magnified view 115. In some embodiments, the movement of the camera and / or the object of interest can cause the object of interest to approach the boundary of the inner frame 145, as shown in the display screen 200B. Although the image of the moon is centered within the magnified view 115, the image can be closer to the boundary of the inner frame 145.
[0055] Such embodiments also include providing a notification to the user that indicates that the region of interest is approaching the boundary of the image overlay. For example, the boundary of the inner frame 145 can turn red, can start flashing, the device can vibrate, an audio notification can be provided, a voice command can be generated, an arrow indicating the direction of movement of the camera to keep the image of the moon away from the boundary of the image overlay can be displayed. Thus, the frame within frame can be configured to guide the user to find their object of interest, and the zoom stabilization mode can be configured to notify the user in the case where the object of interest is closer to the boundary.
[0056] Typically, the inner frame 145 can be configured to automatically slide as an object within the field of view moves, to detect and track the salient object, without the user having to center the object. In some embodiments, a zoom stabilization algorithm can stabilize and track the object while maintaining the object at or near the center of the magnified view 115. Typically, although the outer frame 140 shows the field of view 105, the area defined by the inner frame 145 can be cropped and shown as the magnified view 115. As described herein, an object of interest can be recognized and tracked, and the inner frame 145 can be determined and cropped to generate the magnified view 115, where a closer view of the object of interest is shown at or near the center of the magnified view 115 while maintaining a stabilized image with smooth movement. Thus, the object of interest can be locked at or near the center, or can be shown to move smoothly within the frame. Typically, the object of interest is locked at the center when the ROI is static and / or moving at a constant speed. In some embodiments, the motion trajectory of the region of interest indicates a changing speed of movement between consecutive frames. In such embodiments, the adjustment of the preview includes maintaining the smooth movement of the region of interest at or near the center of the zoomed portion between consecutive frames of the preview. For example, when the ROI is moving at a varying speed, the object of interest is shown to move smoothly within the frame. The smooth movement tracks the ROI as the ROI moves. This approach enables high-zoom photos and / or videos without having to be mounted on a tripod.
[0057] Typically, the relative position of the image of the moon 150 remains stable within the magnified preview 115, while the frames 140 and 145 track the moon, and thus the preview shown to the user is well centered and / or the moon moves smoothly and remains relatively stable within the frames 140 and 145. This is a significant improvement to the display function of the image capture device, as the stabilized image of the moon could potentially move out of the magnified preview 115, and the moon could disappear from the field of view when the user tries to move the device to manually track the moon. However, the zoom stabilization algorithm is able to track the moon and alert the user when the moon is at or near the boundary of the inner frame 145. This is particularly useful at high magnifications, as the preview FOV can be very narrow. For example, at which, the object of interest can easily move out of the preview, and the zoom stabilization tracking is able to detect the object of interest, frame the object of interest, smoothly track the object of interest, and also alert the user.
[0058] Figure 3A is an example workflow 300A for stabilized object tracking according to an example embodiment. Some image capture devices include a hardware abstraction layer (HAL) that connects higher-level camera framework application programming interfaces (APIs) in the camera application (APP) layer to the underlying camera driver and hardware.
[0059] At 302, a telephoto preview (or zoom stabilization mode) can be activated when the magnification exceeds a threshold magnification (e.g., ). The zoom ratio is for illustrative purposes only and can vary depending on the device and / or system configuration.
[0060] At 304, an input tracker can be initialized. At step 1, the system can move to block 306 to determine if a tap ROI is being tracked. In some embodiments, determining the region of interest includes receiving a user indication of the region of interest. At step 2, the system can determine that a tap ROI is being tracked, and at block 308, a user indication of the ROI can be detected, and an initial ROI center can be extracted from the user tap.
[0061] At step 3, the system can determine that the tap ROI is not being tracked. Some embodiments include generating a saliency map through a neural network. For example, at block 310, a saliency detection algorithm and / or a face detection algorithm can be activated to identify the ROI. For example, at block 310, a machine learning (ML)-based saliency map can be determined in the absence of a detected user tap. The saliency can be directly applied to the sensor area without further cropping. This allows for potential detection of new salient ROIs outside the user's final zoomed FOV. The algorithm can then extract the initial ROI center from the ML-based saliency map. In some embodiments, the system can estimate the ROI size through motion vectors from adjacent frames.
[0062] In some embodiments, the algorithm tracks the ROI by: (i) a combination of motion vectors and optical flow, or (ii) using an ML hybrid tracker. Motion vector processing can provide better accuracy when the ROI transformation is rigid and unoccluded, and the hybrid tracker can be more reliable when the transformation is non-rigid or occlusion occurs.
[0063] In some embodiments, the hybrid tracker uses the ROI center with the cropped frame to make the ROI trackable after being reduced in size. In some embodiments, the cropping ratio can be set to the target digital magnification, and / or a slightly more intelligent magnification to keep the ROI within the cropped frame. Then, at step 4, the process continues to the ROI region to be tracked.
[0064] At block 312, the hybrid tracker can use combined saliency detection, object tracker, and optical flow to jointly stabilize the ROI to obtain a reliable and accurate ROI. If the difference in the ROI center between the hybrid tracker and the non-hybrid tracker is small, the algorithm uses the non-hybrid tracker, and if the difference between the two trackers is large, the algorithm can use additional weights to smoothly weight in the hybrid tracker. Generally, the non-hybrid tracker or the ILK tracker uses a motion vector map (a motion vector map similar to optical flow but with a tile size of 64×64) to find the shift in the ROI between adjacent frames. As used herein, the term "ILK" generally refers to the inverse search version of the Lucas-Kanade algorithm for optical flow estimation. The term "ILK tracker" refers to an optical flow-based tracker. The joint stabilization results in the stabilization of the potential frame delay caused by the camera pipeline depth (e.g., 5 frames). For example, information from the hybrid tracker can be combined with the motion vectors from the ILK to predict the ROI. This addresses the potential frame delay caused by the camera pipeline depth. Then, at step 5, the process continues to determine the ROI center and ROI confidence.
[0065] At block 314, EIS inputs such as gyroscope and / or OIS data, as well as frame metadata, are provided to the zoom stabilization algorithm. For example, real-time filtering and lightweight optimization are performed to stabilize the frame using the gyroscope and / or OIS data and the ROI input. Then, at step 6, the process continues to the zoom stabilization algorithm 316.
[0066] At block 316, the algorithm can generate a stabilized frame using the EIS input (e.g., gyroscope sensor, OIS sensor, etc.) based on real-time filtering and lightweight optimization, and obtain the motion trajectory while overcoming hardware limitations (such as gyroscope noise, OIS sensing noise, OIS calibration error, signal latency, etc.). In some embodiments, a small-resolution full sensor frame, the stabilized frame center coordinates, and the crop ratio can be provided to the user interface to generate the frame-in-frame viewfinder as described above.
[0067] Figure 3B is an example workflow 300B for applying zoom stabilization according to an example embodiment. In particular, the features of the Figure 3A zoom stabilization algorithm 316 are described herein. The zoom stabilization algorithm is based on the following relationship: (Equation 1)
[0068] The gyroscope / OIS noise 340 is provided to the camera motion analysis 342. The camera motion analysis 342 can determine the motion trajectory of the region of interest. For example, spatial information about the position of one or more objects of interest (e.g., face, bounding box, etc.) can be extracted from each captured image frame. Some embodiments include determining motion vectors associated with a previous image frame and a current image frame. For example, a motion vector can be generated from the spatial information by averaging two adjacent frames. For example, motion vectors can be extracted between consecutive frames under each tile. This results in enhanced tracking capabilities.
[0069] For magnifications above a threshold (e.g., ), a user tap indicating the ROI can take precedence over automatic tracking. In the absence of a user tap, the system can use a face detection algorithm to detect a face or a saliency model to detect an object of interest. In the absence of a face or an object of interest, the motion vector can be based on the center of the frame. The motion vector between the previous frame and the current frame is determined to obtain an approximate model of the frame-by-frame movement.
[0070] The motion vector information can be combined with the output of the saliency model or the face detection model to determine the tracking of the object of interest. For example, the true camera pose 344 is provided to the video stabilization 346. This process controls the virtual rotational stabilization (e.g., the roll of the camera). For example, the term reduces the pitch / yaw weight to make it less sensitive to gyro noise. Additionally, for example, the term reduces the weight to make it less sensitive to OIS noise.
[0071] Furthermore, for example, the ROI center and confidence 348 provide virtual translational stabilization to the video stabilization 346. The term is introduced to stabilize the ROI at the center. The term corresponds to a stop-move-stop type of motion, resulting in large residual motion under traditional EIS. Based on the true camera pose 344, the actual ROI position can be determined, and optimization can be performed based on the combined image information. Based on the inputs from the camera motion analysis 342 and the ROI center and confidence 348, the video stabilization 346 provides the virtual camera pose 350 to the warping block 354. The image 352 is also provided to the warping block 354.
[0072] Referring again to Figure 3A , at step 7, frame warping from the zoom stabilization algorithm 316 is used to generate the stabilized frame 318. Additionally, at step 8, the bounding box of the stabilized region is provided to the frame-in-frame UI feedback 320. In the absence of a reliable ROI available from the previous frame, a center crop can be performed on the frame. Otherwise, an ROI-centered crop is performed.
[0073] At step 9, the algorithm provides the small-resolution full sensor frame (before stabilization) to block 320. At block 320, the algorithm provides the small-resolution full sensor frame, the stabilized frame center coordinates, and the cropping ratio to the UI to generate an in-frame viewfinder, thereby enabling dynamic preview bounding box visualization.
[0074] Generally, non-hybrid trackers can offer better accuracy. Hybrid trackers can provide a better trade-off between occlusion handling and accuracy. In some embodiments, non-hybrid and hybrid trackers can be combined to achieve an optimal reliable and accurate ROI.
[0075] To maintain the stabilization quality, gyroscope and / or OIS noise that scales with the magnification can be suppressed, and rotation effect correction, seamless transition, etc. can be achieved.
[0076] In the case where the preview includes multiple objects, the most prominent object among the objects can be identified. For example, in the case where there are multiple objects in the preview, the attention can be focused on one object instead of switching between multiple objects. For example, face detection type matching can be performed to identify the face of interest among several faces in the image. Additionally, for example, a machine learning model (e.g., a visual saliency model) can generate saliency scores for multiple candidate salient objects detected in the image. The zoom stabilization algorithm can select the object with a high saliency score as the salient object. The saliency model can be trained on training data indicating user interests and / or preferences, and the trained saliency model can predict the user's object of interest.
[0077] For example, the visual saliency model can be trained based on a training data set including training scenes, sequences, and / or events. For example, the training data set can include images (e.g., digital photos) that include user-drawn bounding boxes containing visual saliency regions (e.g., regions where one or more objects of particular user interest can reside). Based on the training data set, the visual saliency model can predict the visual saliency regions within the image. For example, as a result of the training, the visual saliency model can generate a visual saliency heat map for a given image and produce a bounding box enclosing the region with the highest visual saliency probability (e.g., the highest saliency score). One or more processors can calculate the visual saliency heat map in the background operation of the device. In some embodiments, the visual saliency heat map can indicate the magnitude of the visual saliency probability on a scale from black to white, where white indicates a high saliency probability and black indicates a low saliency probability.
[0078] In some embodiments, the visual saliency heatmap includes a bounding box surrounding the region within the image that contains the highest visual saliency probability. In cases where there are multiple objects of interest in a photo scene, causing the visual saliency model to identify multiple salient regions within the captured image, the visual saliency model can be trained to produce a bounding box surrounding the salient region closest to the center of the captured image. This trained technique assumes that the user is interested in the most central object in the image. Alternatively, the visual saliency model can be trained to produce a bounding box surrounding all the objects of interest in the captured image.
[0079] The image capture device can operate under the guidance of an auto - zoom manager that implements various aspects of the zoom - stabilization mode. In some embodiments, automatically or in response to a received trigger signal, including for example a user performing a predefined gesture (e.g., tap, press) on an input / output device, the auto - zoom manager can implement several steps to calibrate the image capture device. For example, the auto - zoom manager can receive one or more captured images from the image sensor of the image capture device and use a visual saliency model to generate a visual saliency heatmap using the one or more captured images. The visual saliency model can also output a bounding box surrounding the region with the highest visual saliency probability.
[0080] In cases where there are multiple objects of interest in the preview, causing the visual saliency model to identify multiple salient regions within the preview, the visual saliency model can be trained to produce a bounding box surrounding the salient region closest to the center of the preview. This trained technique assumes that the user is interested in the most central object in the image. Alternatively, the visual saliency model can be trained to produce a bounding box surrounding all the objects of interest in the preview or the object of interest with the highest saliency score.
[0081] Some embodiments include determining that the adjusted preview is at a magnification below a threshold magnification. For example, the magnification can be below Such embodiments include transitioning from a zoom - operation mode to a normal operation mode. Thus, the zoom - stabilization mode can be deactivated, and the normal operation mode can be activated. In some embodiments, in the normal operation mode, object tracking may no longer be performed. In some embodiments, object tracking may continue; however, the frame - in - frame view and / or an enlarged view of a portion of the entire field of view (e.g., a cropped portion of the entire FOV corresponding to the framing of the ROI) may no longer be generated and / or displayed.
[0082] Figure 4 is an example workflow for processing consecutive frames in a hybrid tracker according to an example embodiment. It can be in the frame The detection at the location is based on the ROI tapped by the user, i.e., the touch ROI 405. The hybrid tracker 410 can track the ROI to determine the ROI at the frame at the location indicated by the square 415 .
[0083] The hybrid tracker path can be determined as: (Equation 2)
[0084] where , represents the composition of the function with itself, and represents the coordinates of the ROI of the hybrid tracker based on predicting the position of the ROI at the frame using the ILK motion vector.
[0085] A non - hybrid (e.g., optical - flow - based ILK) tracker 430 can determine the full - frame motion vector (MV) of the frame , and determine the motion vector of the ROI in the frame at the square 435 as . In some embodiments, voting can be applied at step 1 to determine at the square 440. The non - hybrid tracker path can be determined as: (Equation 3)
[0086] where represents the coordinates of the ROI based on the non - hybrid tracker or the ILK tracker. To determine the final ROI, a threshold condition 420 can be checked: (Equation 4)
[0087] where represents the final ROI coordinates. In some embodiments, a selection can be made between and . For example, when it is determined that the threshold condition 420 is not met, the system can select the provided by Equation 2 as the selected ROI. As indicated in the square 425, this ROI can be based on and the combination of the ILK from the non - hybrid tracker. For example, when the results of the two tracker methods - the hybrid tracker compounded with the ILK motion vector and the non - hybrid (or ILK tracker) - are significantly different, this indicates the presence of occlusion and / or non - rigid transformation. Therefore, IH(t) is selected because the hybrid tracker is more robust in handling occlusion / non - rigid transformation. In such cases, the selection is used.
[0088] In addition, for example, when it is determined that the threshold condition 420 is satisfied, the system may select the ROI as , as determined by Equation 3. For example, when the results of two tracker methods - a hybrid tracker combined with ILK motion vectors and a non-hybrid (or ILK tracker) - are close to each other, the ILK tracker provides higher accuracy and is used for selection .
[0089] The selected ROI can be set as the new ROI for the iterative process.
[0090] Blocks 450 - 460 illustrate the processing of an object of interest represented by the letter "A". At block 450, the letter "A" is shown in the frame . At block 455, the letter "A" is shown at a new position in the next frame . Thus, the hybrid tracker calculates: (Equation 5)
[0091] At block 460, the letter "A" is shown as having moved further to the right in the next frame . Thus, the hybrid tracker calculates: (Equation 6)
[0092] Although this example shows the calculation for three consecutive frames, a similar iterative approach also applies to consecutive frames. In some embodiments, the hybrid tracker may downscale the frame to , and the ROI center of the cropped frame can be initialized as the hybrid tracker input so that the object of interest can be tracked after downscaling.
[0093] As described herein, an inner frame (e.g., inner frame 145) can be cropped from the entire FOV to determine an enlarged and stabilized view. The cropping ratio for such cropping can be set to the target digital magnification or a more intelligent desired magnification to ensure that the ROI is within the cropped frame.
[0094] Returning to reference Figure 3B , based on the true camera pose 344, the actual ROI position can be determined, and optimization can be performed based on the combined image information as follows: (Equation 7)
[0095] where is the tracking point in the real domain, and is the target point in the virtual domain. The term reduces the pitch / deflection weight to make the stabilization less sensitive to gyro noise. In addition, for example, the term Reduce the weight so that the stabilization is less sensitive to OIS noise. The term corresponds to the pause-move-pause type of motion, thus resulting in large residual motion under traditional EIS. The virtual pose can be determined by the term as follows: (Equation 8)
[0096] where denotes the inverse of the matrix . Here, denotes the internal parameter matrix of the camera corresponding to the virtual camera pose, is the predicted rotation for the virtual camera pose, denotes the internal parameter matrix of the camera corresponding to the real camera pose, is the predicted rotation for the real camera pose. The weight term can be smaller on the pitch / roll axes than in traditional EIS, but the same weight can be maintained for roll. Thus, the tracking term can then dominate the pitch / roll compensation to reduce the residual motion caused by gyro / OIS noise.
[0097] In some embodiments, a two-step optimization can be performed.
[0098] Step 1: Find the target point where the ROI will be located in the stabilized frame .
[0099] Step 2: Find the virtual camera pose.
[0100] Figure 5 depicts an example tracking optimization process according to an example embodiment. Three consecutive input frames are shown as images of a cat. Input frame 1 505 shows the cat on the left side of the display, with a bounding box on the cat's face. Input frame 2510 shows the cat at the center of the display, with a bounding box on the cat's face. The initial motion vector is generated based on the positions of the consecutive bounding boxes in input frame 1505 and input frame 2 510, as indicated by the dashed line.
[0101] Input frame 3 515 shows the cat on the right side of the display, with a bounding box on the cat's face. The motion vector is generated based on the motion vector in input frame 2 510 and the positions of the consecutive bounding boxes in input frame 2 510 and input frame 3 515, as indicated by the two dashed lines. The point is the tracking point in the real domain. The stabilized frame 520 can be generated based on input frames 1 to 3, and is displayed as the target point in the virtual domain. The point can be solved from the following closed-form equation: (Equation 9)
[0102] where the operation "." represents multiplication, and and are positive weight coefficients satisfying .
[0103] Figure 5 illustrates how to determine the first term in Equation 9 , where represents the coordinates of the virtual target center in the previous frame. The term constrains the coordinates of to be close to the coordinates of , and this in turn ensures that the center of the virtual ROI is stabilized.
[0104] Figure 6 depicts another example tracking optimization process according to an example embodiment. Input frames 605, 615, and 625 are shown as images of a cat moving from the left, towards the center, and towards the right, respectively. Points are determined at consecutive frames to generate stabilized frames 610, 620, and 630 corresponding to input frames 605, 615, and 625, respectively. Figure 6 illustrates the effect of the middle term in Equation 9. Here, represents the position of the ROI in the true pose (e.g., the unstabilized frame), and thus the term can be adjusted to control how closely the virtual target follows the true position. For example, input frame 615 corresponds to , and the target point can be determined. For example, if is less than a second threshold thresh2, then , where "diff" represents the difference or distance. Input frame 625 corresponds to the case , and the target point can be determined. As indicated by the stabilized frames 610, 620, and 630, even though the position of the cat is shifted in input frames 605, 615, and 620, the position of the cat indicated by the stabilized frames 610, 620, and 630 remains at or near the center of the frame. The point can be solved from the closed-form equation Equation 9.
[0105] Figure 7 depicts another example tracking optimization process according to an example embodiment. Input frames 705, 715, and 725 are shown as images with a cat on the left, and the stabilized frames 710, 720, and 730 correspond to input frames 705, 715, and 725, respectively. Again, the point It can be solved from the closed-form equation, Equation 9. This example shows the third term in Equation 9 . In this example, if is less than the third threshold thresh3, then . The term represents the distance between and the absolute center of the stabilized frame. Thus, the condition that this distance is less than the threshold ensures that the virtual ROI is not at or near the boundaries of the stabilized frame. The stabilized frame 2720 shows the situation , while the stabilized frame 3730 shows the situation
[0106] Figure 8 depicts an example tracking optimization process for two regions of interest according to an example embodiment. For example, input frame 1805 has an image of a cat, and a new object of interest represented by a dog is detected in input frame 2815. The new ROI can be detected by a touch tap event (e.g., the user taps on the display to indicate the ROI). In the case of a new ROI, the hybrid tracker described herein replaces the non-hybrid (also referred to herein as ILK) tracker. The point can be solved from the closed-form equation: (Equation 10)
[0107] where, represents the position of the ROI in the true pose, represents the position of the ROI in the virtual pose, represents the inverse of the matrix . Here, represents the intrinsic matrix of the camera corresponding to the virtual camera pose, is the predicted rotation for the virtual camera pose, represents the intrinsic matrix of the camera corresponding to the true camera pose, is the predicted rotation for the true camera pose.
[0108] Figure 9An example image with stabilized object tracking according to an example embodiment is shown. The initial FOV 905 is shown with a bounding box 915 indicating an object of interest (e.g., an airplane). A warping grid 910 is shown. For example, after determining the virtual camera pose, a stabilized grid from the physical camera to the virtual camera, such as the warping grid 910, can be generated by determining the source quadrilateral and the target quadrilateral for each horizontal stripe. In some embodiments, the warping grid 910 can be cropped from the entire initial FOV 905 to generate an enlarged FOV 920. The object of interest, in this example an airplane 925, is shown in a closer view in the enlarged FOV 920. As shown, the airplane 925 is clearly visible and can be smoothly tracked based on motion vectors in consecutive frames.
[0109] Figure 10 Another example image with stabilized object tracking according to an example embodiment is shown. The initial FOV 1005 is shown with a bounding box 1015 indicating an object of interest (e.g., a mailbox displaying a house number). A warping grid 1010 is shown. For example, after determining the virtual camera pose, a stabilized grid from the physical camera to the virtual camera, such as the warping grid 1010, can be generated by determining the source quadrilateral and the target quadrilateral for each horizontal stripe. In some embodiments, the warping grid 1010 can be cropped from the entire initial FOV 1005 to generate an enlarged FOV 1020. The object of interest, in this example a mailbox displaying a house number 1025, is shown in a closer view in the enlarged FOV 1020. As shown, the mailbox 1025 is clearly visible and the house number "705" can be discerned.
[0110] Figure 11 Another example image with stabilized object tracking according to an example embodiment is shown. The initial FOV 1105 is shown with a bounding box 1115 indicating an object of interest (e.g., the moon). A warping grid 1110 is shown. For example, after determining the virtual camera pose, a stabilized grid from the physical camera to the virtual camera, such as the warping grid 1110, can be generated by determining the source quadrilateral and the target quadrilateral for each horizontal stripe. In some embodiments, the warping grid 1110 can be cropped from the entire initial FOV 1105 to generate an enlarged FOV 1120. The object of interest, in this example the moon 1125, is shown in a closer view in the enlarged FOV 1120. As shown, the moon 1125 is clearly visible and can be smoothly tracked based on motion vectors in consecutive frames.
[0111] As described herein, the magnification for telephoto, Tele RM, can be times, and the magnification for zoom stabilization can be The zoom ratio is for illustrative purposes only and may vary depending on the device and / or system configuration. The image capture device can transition between two modes - a normal mode without zoom stabilization and a zoom stabilization mode. In some embodiments, the zoom stabilization mode can be controlled by a combination of magnification and grid interpolation. In some embodiments, the transition can be between a baseline EIS mode and a center ROI-based zoom stabilization mode. Generally, the transition is seamless, with or without an ROI tracking item. For example, the transition between the center ROI and an ROI source can be seamless by adjusting the virtual ROI target and re-tracking the transition between different ROI sources. Additionally, for example, the transition between a pixel merge mode and a remosaic mode can be seamless by performing frame-by-frame cropping of YUV in a camera application and adjusting the EIS margins accordingly. Training a machine learning model for generating inferences / predictions
[0112] Figure 12 FIG. 1200 according to an example embodiment is shown, which illustrates a training phase 1202 and an inference phase 1204 of a trained machine learning model 1232. Some machine learning techniques involve training one or more machine learning algorithms on an input training data set to identify patterns in the training data and provide output inferences and / or predictions regarding the patterns in the training data. The resulting trained machine learning algorithms can be referred to as trained machine learning models. For example, Figure 12 The training phase 1202 is shown, during which one or more machine learning algorithms 1220 are being trained on training data 1210 to become a trained machine learning model 1232. Then, during the inference phase 1204, the trained machine learning model 1232 can receive input data 1230 and one or more inference / prediction requests 1240 (possibly as part of the input data 1230), and provide one or more inferences and / or predictions 1250 as output in response.
[0113] Thus, the trained machine learning model 1232 can include one or more models of one or more machine learning algorithms 1220. The machine learning algorithms 1220 can include, but are not limited to: artificial neural networks (e.g., convolutional neural networks, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems as described herein). The machine learning algorithms 1220 can be supervised or unsupervised and can implement any suitable combination of online learning and offline learning.
[0114] In some examples, the machine learning algorithm 1220 and / or the trained machine learning model 1232 can be accelerated using an on-device co-processor, such as a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), and / or an application specific integrated circuit (ASIC). Such on-device co-processors can be used to accelerate the machine learning algorithm 1220 and / or the trained machine learning model 1232. In some examples, the trained machine learning model 1232 can be trained, resident, and executed to provide inference on a specific computing device, and / or otherwise can make inferences for a specific computing device.
[0115] During the training phase 1202, the machine learning algorithm 1220 can be trained using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques by at least providing training data 1210 as training input. Unsupervised learning involves providing a portion (or all) of the training data 1210 to the machine learning algorithm 1220, and the machine learning algorithm 1220 determining one or more output inferences based on the provided portion (or all) of the training data 1210. Supervised learning involves providing a portion of the training data 1210 to the machine learning algorithm 1220, where the machine learning algorithm 1220 determines one or more output inferences based on the provided portion of the training data 1210, and accepting or correcting the output inferences based on the correct results associated with the training data 1210. In some examples, the supervised learning of the machine learning algorithm 1220 can be controlled by a rule set and / or a tag set for the training input, and the rule set and / or the tag set can be used to correct the inferences of the machine learning algorithm 1220.
[0116] Semi-supervised learning involves having correct results for a portion (but not all) of the training data 1210. During semi-supervised learning, supervised learning is used for the portion of the training data 1210 that has correct results, and unsupervised learning is used for the portion of the training data 1210 that does not have correct results. Reinforcement learning involves the machine learning algorithm 1220 receiving a reward signal regarding a previous inference, where the reward signal can be a numerical value. During reinforcement learning, the machine learning algorithm 1220 can output an inference and receive the reward signal as a response, where the machine learning algorithm 1220 is configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values provided by the reward signal over time. In some examples, the machine learning algorithm 1220 and / or the trained machine learning model 1232 can be trained using other machine learning techniques, including but not limited to incremental learning and curriculum learning.
[0117] In some examples, the machine learning algorithm 1220 and / or the trained machine learning model 1232 may use transfer learning techniques. For example, transfer learning techniques may involve pre-training the trained machine learning model 1232 on one data set and additionally training it using the training data 1210. More specifically, the machine learning algorithm 1220 may be pre-trained on data from one or more computing devices, and the resulting trained machine learning model is provided to the computing device CD1, where CD1 is intended to execute the trained machine learning model during the inference phase 1204. Then, during the training phase 1202, the pre-trained machine learning model may be additionally trained using the training data 1210, where the training data 1210 may be derived from the kernel data and non-kernel data of the computing device CD1. This further training of the machine learning algorithm 1220 and / or the pre-trained machine learning model using the training data 1210 from CD1's data may be performed using supervised learning or unsupervised learning. Once the machine learning algorithm 1220 and / or the pre-trained machine learning model has been trained at least on the training data 1210, the training phase 1202 may be completed. The resulting trained machine learning model may be used as at least one of the trained machine learning models 1232.
[0118] In particular, once the training phase 1202 has been completed, if there is no trained machine learning model 1232 on the computing device, the trained machine learning model may be provided to the computing device. After the trained machine learning model 1232 has been provided to the computing device CD1, the inference phase 1204 may begin.
[0119] During the inference phase 1204, the trained machine learning model 1232 may receive the input data 1230 and generate and output one or more corresponding inferences and / or predictions 1250 regarding the input data 1230. Thus, the input data 1230 may be used as an input to the trained machine learning model 1232 for providing the corresponding inferences and / or predictions 1250 to the kernel components and non-kernel components. For example, the trained machine learning model 1232 may generate inferences and / or predictions 1250 in response to one or more inference / prediction requests 1240. In some examples, the trained machine learning model 1232 may be executed as part of other software. For example, the trained machine learning model 1232 may be executed by an inference or prediction daemon to be readily available for providing inferences and / or predictions upon request. The input data 1230 may include data from the computing device CD1 on which the trained machine learning model 1232 is executed and / or input data from one or more computing devices other than CD1.
[0120] The training data 1210 may include an image (e.g., a digital photograph) that includes a user-drawn bounding box that encloses a visually significant region (e.g., a region in which one or more objects of particular interest to the user may reside).
[0121] The input data 1230 may include one or more captured images, or a preview of an image. Other types of input data are possible.
[0122] The inference and / or prediction 1250 may include an output image, a bounding box that encloses a region of interest having the highest visual significance probability, and / or other output data generated by a trained machine learning model 1232 that operates on the input data 1230 (and the training data 1210). In some examples, the trained machine learning model 1232 may use the output inference and / or prediction 1250 as input feedback 1260. The trained machine learning model 1232 may also rely on past inferences as input for generating new inferences.
[0123] A convolutional neural network, such as a visual significance model, may be an example of the machine learning algorithm 1220. After training, a trained version of the convolutional neural network may be an example of the trained machine learning model 1232. In this manner, an example of the inference / prediction request 1240 may be a request to predict a region of interest in a preview of an image, and a corresponding example of the inference and / or prediction 1250 may be an output image having a bounding box that encloses a visually significant region.
[0124] In some examples, a computing device CD_SOLO may include a trained version of a convolutional neural network 100, possibly after training the convolutional neural network. The computing device CD_SOLO may then receive a request to predict a region of interest in a preview of an image and use the trained version of the convolutional neural network to generate an output image having a bounding box that encloses a visually significant region.
[0125] In some examples, two or more computing devices CD_CLI and CD_SRV may be used to provide the output image; for example, a first computing device CD_CLI may generate and send a request to a second computing device CD_SRV to predict a region of interest in a preview of an image. The CD_SRV may then use a trained version of a convolutional neural network (possibly after training the convolutional neural network) to generate an output image having a bounding box that encloses a visually significant region and respond to the request from the CD_CLI to predict a region of interest in a preview of an image. Then, upon receiving a response to the request, the CD_CLI may provide the requested region of interest (e.g., using a user interface and / or a display). Example data network
[0126] Figure 13 Depicts a distributed computing architecture 1300 according to an example embodiment. The distributed computing architecture 1300 includes server devices 1308, 1310 configured to communicate with programmable devices 1304a, 1304b, 1304c, 1304d, 1304e via a network 1306. The network 1306 may correspond to a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a wireless wide area network (WWAN), a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 1306 may also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.
[0127] Although Figure 13 only five programmable devices are shown, the distributed application architecture may serve dozens, hundreds, or thousands of programmable devices. Additionally, the programmable devices 1304a, 1304b, 1304c, 1304d, 1304e (or any additional programmable devices) may be any kind of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mounted device (HMD), a network terminal, a mobile computing device, and so on. In some examples, as shown for programmable devices 1304a, 1304b, 1304c, 1304e, the programmable devices may be directly connected to the network 1306. In other examples, as shown for programmable device 1304d, the programmable device may be indirectly connected to the network 1306 via an associated computing device (such as programmable device 1304c). In this example, programmable device 1304c may act as the associated computing device for relaying electronic communications between programmable device 1304d and the network 1306. In other examples, as shown for programmable device 1304e, the computing device may be part of a vehicle (such as a car, truck, bus, boat, or ship, an airplane, etc.) and / or located inside the vehicle. In Figure 13 other examples not shown, the programmable device may be both directly and indirectly connected to the network 1306.
[0128] The server devices 1308, 1310 may be configured to perform one or more services as requested by the programmable devices 1304a - 1304e. For example, the server device 1308 and / or 1310 may provide content to the programmable devices 1304a - 1304e. The content may include, but is not limited to, web pages, hypertext, scripts, binary data (such as compiled software), images, audio, and / or video. The content may include compressed and / or uncompressed content. The content may be encrypted and / or unencrypted. Other types of content are also possible.
[0129] As another example, server devices 1308 and / or 1310 may provide access to software for databases, search, computing, graphics, audio, video, World Wide Web / Internet utilization, and / or other functions to programmable devices 1304a - 1304e. Many other examples of server devices are possible. Computing device architecture
[0130] Figure 14 is a block diagram of an example computing device 1400 according to an example embodiment. In particular, Figure 14 the illustrated computing device 1400 may be configured to perform at least one function of method 1600 and / or at least one function related to the method.
[0131] The computing device 1400 may include a user interface module 1401, a network communication module 1402, one or more processors 1403, a data store 1404, one or more cameras 1418, one or more sensors 1420, and a power system 1422, all of which may be linked together via a system bus, network, or other connection mechanism 1405.
[0132] The user interface module 1401 may be operable to send data to and / or receive data from an external user input / output device. For example, the user interface module 1401 may be configured to send data to and / or receive data from a user input device such as a touch screen, computer mouse, keyboard, keypad, touchpad, trackball, joystick, voice recognition module, and / or other similar devices. The user interface module 1401 may also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices, whether now known or later developed. The user interface module 1401 may also be configured to generate an audible output using a device such as a speaker, speaker jack, audio output port, audio output device, headphones, and / or other similar devices. The user interface module 1401 may further be configured with one or more haptic devices that may generate a haptic output such as vibration and / or other output detectable by touch and / or physical contact with the computing device 1400. In some examples, the user interface module 1401 may be used to provide a graphical user interface (GUI) for utilizing the computing device 1400.
[0133] The network communication module 1402 may include one or more devices that provide one or more wireless interfaces 1407 and / or one or more wired interfaces 1408 that are configurable to communicate via a network. The wireless interface 1407 may include one or more wireless transmitters, receivers, and / or transceivers, such as Bluetooth™ transceivers, Zigbee® transceivers, Wi-Fi™ transceivers, WiMAX™ transceivers, LTE™ transceivers, and / or other types of wireless transceivers that are configurable to communicate via a wireless network. The wired interface 1408 may include one or more wired transmitters, receivers, and / or transceivers, such as Ethernet transceivers, Universal Serial Bus (USB) transceivers, or similar transceivers that may be configured to communicate via twisted pair, coaxial cable, fiber optic link, or similar physical connections to a wired network.
[0134] In some examples, the network communication module 1402 may be configured to provide reliable, protected, and / or authenticated communication. For each communication described herein, information may be provided to facilitate reliable communication (e.g., guaranteed message delivery), which may be part of a message header and / or message footer (e.g., packet / message sequencing information, encapsulation headers and / or encapsulation footers, size / time information, and transmission verification information, such as cyclic redundancy check (CRC) and / or parity values). One or more cryptographic protocols and / or algorithms may be used to secure (e.g., encode or encrypt) and / or decrypt / decode the communication, such as but not limited to the Data Encryption Standard (DES), Advanced Encryption Standard (AES), Rivest-Shamir-Adelman (RSA) algorithm, Diffie-Hellman algorithm, Secure Sockets Protocol (e.g., Secure Sockets Layer (SSL) or Transport Layer Security (TLS)), and / or Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms may be used to secure (and then decrypt / decode) the communication in addition to those listed herein.
[0135] The one or more processors 1403 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPU), graphics processing units (GPU), application specific integrated circuits, etc.). The one or more processors 1403 may be configured to execute computer-readable instructions 1406 contained in the data storage 1404 and / or other instructions described herein.
[0136] The data storage 1404 may include one or more non-transitory computer-readable storage media that may be read and / or accessed by at least one of the one or more processors 1403. The one or more computer-readable storage media may include volatile and / or non-volatile storage components that may be integrated, in whole or in part, with at least one of the one or more processors 1403, such as optical, magnetic, organic, or other memory or disk storage. In some examples, the data storage 1404 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other examples, the data storage 1404 may be implemented using two or more physical devices.
[0137] The data storage 1404 may include computer-readable instructions 1406 and possibly additional data. In some examples, the data storage 1404 may include storage required to execute at least a portion of the methods, scenarios, and techniques described herein and / or at least a portion of the functionality of the devices and networks described herein. In some examples, the data storage 1404 may include storage for a trained neural network model 1412 (e.g., a model of a trained convolutional neural network). Specifically, in these examples, the computer-readable instructions 1406 may include instructions that, when executed by the processor 1403, cause the computing device 1400 to be able to provide some or all of the functionality of the trained neural network model 1412.
[0138] In some examples, the computing device 1400 may include one or more cameras 1418. The cameras 1418 may include one or more image capture devices, such as a still camera and / or a video camera, that are configured to capture light and record the captured light in one or more images; that is, the cameras 1418 may generate images of the captured light. The one or more images may be one or more still images and / or one or more images utilized in a video stream. The cameras 1418 may capture light and / or electromagnetic radiation that is emitted as visible light, infrared radiation, ultraviolet light, and / or as light and / or electromagnetic radiation at one or more other frequencies.
[0139] In some examples, computing device 1400 may include one or more sensors 1420. The sensors 1420 may be configured to measure conditions within the computing device 1400 and / or conditions in the environment of the computing device 1400 and provide data regarding such conditions. For example, sensors 1420 may include one or more of the following: (i) sensors for obtaining data regarding the computing device 1400, such as but not limited to a thermometer for measuring the temperature of the computing device 1400, a battery sensor for measuring the power of one or more batteries of the power system 1422, and / or other sensors for measuring conditions of the computing device 1400; (ii) identification sensors for identifying other objects and / or devices, such as but not limited to radio frequency identification (RFID) readers, proximity sensors, one-dimensional barcode readers, two-dimensional barcode (e.g., quick response (QR) code) readers, and laser trackers, where the identification sensors may be configured to read identifiers such as RFID tags, barcodes, QR codes, and / or other devices and / or objects configured to be read and provide at least identification information; (iii) sensors for measuring the position and / or movement of the computing device 1400, such as but not limited to tilt sensors, gyroscopes, accelerometers, Doppler sensors, GPS devices, sonar sensors, radar devices, laser displacement sensors, and compasses; (iv) environmental sensors for obtaining data indicative of the environment of the computing device 1400, such as but not limited to infrared sensors, optical sensors, light sensors, biosensors, capacitive sensors, touch sensors, temperature sensors, wireless sensors, radio sensors, motion sensors, microphones, sound sensors, ultrasonic sensors, and / or smoke sensors; and / or (v) force sensors for measuring one or more forces acting around the computing device 1400 (e.g., inertial forces and / or G-forces), such as but not limited to one or more sensors for measuring one or more of the following: force in one or more dimensions, torque, ground force, friction, and / or a zero moment point (ZMP) sensor for identifying the ZMP and / or the location of the ZMP. Many other examples of sensors 1420 are possible.
[0140] Power system 1422 may include one or more batteries 1424 and / or one or more external power interfaces 1426 for supplying power to computing device 1400. Each of the one or more batteries 1424 may act as a source of stored power for computing device 1400 when electrically coupled to computing device 1400. The one or more batteries 1424 of power system 1422 may be configured to be portable. Some or all of the one or more batteries 1424 may be easily removable from computing device 1400. In other examples, some or all of the one or more batteries 1424 may be located inside computing device 1400 and may thus be not easily removable from computing device 1400. Some or all of the one or more batteries 1424 may be rechargeable. For example, a rechargeable battery may be charged via a wired connection between the battery and another power supply (such as via one or more power supplies located outside computing device 1400 and connected to computing device 1400 via one or more external power interfaces). In other examples, some or all of the one or more batteries 1424 may be non-rechargeable batteries.
[0141] One or more external power interfaces 1426 of power system 1422 may include one or more wired power interfaces that implement a wired power connection to one or more power supplies located outside computing device 1400, such as a USB cable and / or a power line. One or more external power interfaces 1426 may include one or more wireless power interfaces that implement a wireless power connection to one or more external power supplies, such as via a Qi wireless charger, such as a Qi wireless charger. Once a power connection to an external power source is established using one or more external power interfaces 1426, computing device 1400 may draw power from the external power source via the established power connection. In some examples, power system 1422 may include associated sensors, such as battery sensors associated with one or more batteries or other types of power sensors. Cloud-based server
[0142] Figure 15 Depicts a cloud-based server system according to an example embodiment. In Figure 15Among them, the functionality of the convolutional neural network and / or the computing device may be distributed among the computing clusters 1509a, 1509b, and 1509c. The computing cluster 1509a may include one or more computing devices 1500a, a cluster storage array 1510a, and a cluster router 1511a connected through a local cluster network 1512a. Similarly, the computing cluster 1509b may include one or more computing devices 1500b, a cluster storage array 1510b, and a cluster router 1511b connected through a local cluster network 1512b. Likewise, the computing cluster 1509c may include one or more computing devices 1500c, a cluster storage array 1510c, and a cluster router 1511c connected through a local cluster network 1512c.
[0143] In some embodiments, each of the computing clusters 1509a, 1509b, and 1509c may have an equal number of computing devices, an equal number of cluster storage arrays, and an equal number of cluster routers. However, in other embodiments, each computing cluster may have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each computing cluster may depend on one or more computing tasks assigned to each computing cluster.
[0144] For example, in the computing cluster 1509a, the computing device 1500a may be configured to execute various computing tasks of the convolutional neural network, confidence learning, and / or the computing device. In one embodiment, the various functionality of the convolutional neural network, confidence learning, and / or the computing device may be distributed among one or more of the computing devices 1500a, 1500b, and 1500c. The computing devices 1500b and 1500c in the corresponding computing clusters 1509b and 1509c may be configured similarly to the computing device 1500a in the computing cluster 1509a. On the other hand, in some embodiments, the computing devices 1500a, 1500b, and 1500c may be configured to perform different functions.
[0145] In some embodiments, the computing tasks and the stored data associated with the convolutional neural network and / or the computing device may be distributed across the computing devices 1500a, 1500b, and 1500c at least in part based on the following: the processing requirements of the convolutional neural network and / or the computing device, the processing capabilities of the computing devices 1500a, 1500b, 1500c, the latency of the network links between the computing devices in each computing cluster and between the computing clusters themselves, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the overall system architecture.
[0146] The cluster storage arrays 1510a, 1510b, and 1510c of the compute clusters 1509a, 1509b, and 1509c can be data storage arrays that include disk array controllers configured to manage read and write access to groups of hard disk drives. The disk array controllers, alone or in conjunction with their respective computing devices, can also be configured to manage backup or redundant copies of data stored in the cluster storage arrays to protect against disk drive or other cluster storage array failures and / or network failures that prevent one or more computing devices from accessing one or more cluster storage arrays.
[0147] Similar to the way in which the functions of a convolutional neural network and / or computing devices can be distributed across the computing devices 1500a, 1500b, 1500c of the compute clusters 1509a, 1509b, 1509c, the respective active and / or backup portions of these components can be distributed across the cluster storage arrays 1510a, 1510b, 1510c. For example, some cluster storage arrays can be configured to store a portion of the data of a convolutional neural network and / or computing devices, while other cluster storage arrays can store other portions of the data of the convolutional neural network and / or computing devices. Additionally, for example, some cluster storage arrays can be configured to store the data of a first convolutional neural network, while other cluster storage arrays can store the data of a second convolutional neural network and / or a third convolutional neural network. Additionally, some cluster storage arrays can be configured to store backup versions of the data stored in other cluster storage arrays.
[0148] The cluster routers 1511a, 1511b, and 1511c in the compute clusters 1509a, 1509b, and 1509c can include networking equipment configured to provide internal and external communication for the compute clusters. For example, the cluster router 1511a in the compute cluster 1509a can include one or more Internet switching and routing devices configured to (i) provide local area network communication between the computing device 1500a and the cluster storage array 1510a via the local cluster network 1512a, and (ii) provide wide area network communication between the compute cluster 1509a and the compute clusters 1509b and 1509c via the wide area network link 1513a to the network 1306. The cluster routers 1511b and 1511c can include network equipment similar to the cluster router 1511a, and the cluster routers 1511b and 1511c can perform networking functions similar to the networking functions performed by the cluster router 1511a for the compute cluster 1509a for the compute clusters 1509b and 1509b.
[0149] In some embodiments, the configuration of cluster routers 1511a, 1511b, 1511c may be at least partially based on the data communication requirements of the computing devices and the cluster storage array, the data communication capabilities of the network equipment in cluster routers 1511a, 1511b, 1511c, the latency and throughput of local cluster networks 1512a, 1512b, 1512c, the latency, throughput, and cost of wide area network links 1513a, 1513b, 1513c, and / or other factors that may contribute to adjusting the cost, speed, fault tolerance, resilience, efficiency, and / or other design criteria of the system architecture. Example method of operation
[0150] Figure 16 Method 1600 according to an example embodiment is shown. Method 1600 may include various blocks or steps. These blocks or steps may be performed individually or in combination. These blocks or steps may be performed in any order and / or serially or in parallel. Further, blocks or steps may be omitted from or added to method 1600.
[0151] The blocks of method 1600 may be performed by respective elements of computing device 1400 as shown and described with reference Figure 8 shown and described.
[0152] Block 1610 includes displaying, on a display screen of an image capture device, a preview of an image representing the field of view of the image capture device.
[0153] Block 1620 includes determining a region of interest in the preview of the image.
[0154] Block 1630 includes transitioning the image capture device from a normal operation mode to a zoomed operation mode, where the zoomed operation mode includes: determining a motion trajectory of the region of interest based on sensor data collected by a sensor associated with the image capture device; and generating an adjusted preview representing a zoomed portion of the field of view based on the determined motion trajectory, where the adjusted preview displays the region of interest at or near the center of the zoomed portion.
[0155] Block 1640 includes providing the adjusted preview of the portion of the field of view by the display screen.
[0156] Some embodiments include providing an image overlay by the display screen, the image overlay showing a representation of the zoomed portion relative to the field of view.
[0157] Some embodiments include determining a bounding box of the region of interest, and where providing the image overlay includes providing the region of interest framed within the bounding box.
[0158] Some embodiments include determining one or more of the following: (i) a low-resolution version of the displayed image, (ii) coordinates of an adjusted region of interest within an adjusted preview, or (iii) a cropping ratio. Such embodiments also include generating an image overlay to enable dynamic visualization of the bounding box.
[0159] Some embodiments include detecting that a region of interest is approaching a boundary of an image overlay. Such embodiments also include providing a notification to a user that indicates that the region of interest is approaching the boundary of the image overlay.
[0160] Some embodiments include determining a motion vector associated with a previous image frame and a current image frame. Such embodiments also include determining a size of a region of interest based on the determined motion vector.
[0161] Some embodiments include determining an optical flow corresponding to a region of interest. Such embodiments also include tracking the region of interest within a portion of the field of view based on the determined optical flow, and wherein an adjustment of a preview of the image is based on the tracking of the region of interest.
[0162] Some embodiments include tracking a region of interest within a field of view based on a combination of motion vector processing and optical flow. For example, when it is determined that a transformation associated with the region of interest is rigid and there is no occlusion, a combination of motion vector processing and optical flow can be used to track the region of interest.
[0163] Some embodiments include tracking a region of interest within a field of view based on a hybrid tracker. For example, when it is determined that a transformation associated with the region of interest is non-rigid or has occlusion, a hybrid tracker can be used to track the region of interest. In some embodiments, the hybrid tracker is based on a centered cropped frame to track the region of interest after a downscaling operation. In some embodiments, the hybrid tracker includes: (a) one or more motion vectors associated with a current image frame, and (b) a saliency map indicating the region of interest.
[0164] Some embodiments include generating a saliency map by a neural network.
[0165] In some embodiments, a preview includes multiple objects, and the method includes: using a saliency map to select an object among the multiple objects, wherein a determination of the region of interest is based on the selected object.
[0166] In some embodiments, the sensor is one of a gyroscope or an optical image stabilization (OIS) sensor.
[0167] In some embodiments, a motion trajectory of a region of interest indicates a moving speed of a change between consecutive frames. In such embodiments, an adjustment of the preview includes maintaining a smooth movement of the region of interest at or near a center of a zoomed portion between consecutive frames of the preview.
[0168] In some embodiments, the motion trajectory of the region of interest indicates a nearly constant moving speed between consecutive frames. In such embodiments, the adjustment to the preview includes locking the position of the region of interest at or near the center of the zoomed portion between consecutive frames of the preview.
[0169] In some embodiments, determining the region of interest includes receiving a user indication of the region of interest.
[0170] In some embodiments, determining the region of interest includes determining a saliency map indicating the region of interest based on a neural network.
[0171] Some embodiments include determining that the adjusted preview is at a magnification below a threshold magnification. Such embodiments include transitioning from a zoom operation mode to a normal operation mode.
[0172] The particular arrangements shown in the figures should not be considered restrictive. It should be understood that other embodiments may include more or fewer of each element shown in a given figure. Further, some of the elements shown may be combined or omitted. Still further, illustrative embodiments may include elements not shown in the figures.
[0173] The steps or blocks representing the processing of information may correspond to circuitry that can be configured to perform the particular logical functions described herein for the method or technique. Alternatively or additionally, the steps or blocks representing the processing of information may correspond to a module, segment, or portion of program code, including related data. The program code may include one or more instructions that may be executed by a processor to implement the particular logical function or action in the method or technique. The program code and / or related data may be stored on any type of computer-readable medium, such as a storage device including a disk, hard drive, or other storage medium.
[0174] The computer-readable medium may also include non-transitory computer-readable media, such as computer-readable media that stores data for a short period of time, such as register memory, processor cache, and random access memory (RAM). The computer-readable medium may also include non-transitory computer-readable media that stores program code and / or data for a longer period of time. Thus, the computer-readable medium can include auxiliary or persistent long-term storage, such as, for example, read-only memory (ROM), optical or magnetic disks, compact disc read-only memory (CD-ROM). The computer-readable medium can also be any other volatile or non-volatile storage system. The computer-readable medium can be considered, for example, a computer-readable storage medium or a tangible storage device.
[0175] Although various examples and embodiments have been disclosed, other examples and embodiments will be apparent to those skilled in the art. The various examples and embodiments disclosed are for illustrative purposes and are not intended to be limiting, with the true scope being indicated by the appended claims.
Claims
1. A computer-implemented method, comprising: displaying, on a display screen of an image capture device, a preview of an image representing a field of view of the image capture device; determining a region of interest in the preview of the image; transitioning the image capture device from a normal operation mode to a zoomed operation mode, wherein the zoomed operation mode comprises: determining, based on sensor data collected by a sensor associated with the image capture device, a motion trajectory of the region of interest, and generating, based on the determined motion trajectory, an adjusted preview representing a zoomed portion of the field of view, wherein the adjusted preview displays the region of interest at or near the center of the zoomed portion; and providing, by the display screen, the adjusted preview of the portion of the field of view.
2. The method according to claim 1, further comprising: providing, by the display screen, an image overlay that displays a representation of the zoomed portion relative to the field of view.
3. The method according to claim 2, further comprising: determining a bounding box of the region of interest, and wherein providing the image overlay comprises providing the region of interest framed within the bounding box.
4. The method according to claim 3, further comprising: determining one or more of: (i) a low-resolution version of the displayed image, (ii) coordinates of the adjusted region of interest within the adjusted preview, or (iii) a cropping ratio; and generating the image overlay to effect dynamic visualization of the bounding box.
5. The method according to claim 2, further comprising: detecting that the region of interest is approaching a boundary of the image overlay; and providing a user notification that indicates that the region of interest is approaching the boundary of the image overlay.
6. The method according to claim 1, further comprising: determining a motion vector associated with a previous image frame and a current image frame; and determining the size of the region of interest based on the determined motion vector.
7. The method according to claim 1, further comprising: determining an optical flow corresponding to the region of interest; and tracking the region of interest within the portion of the field of view based on the determined optical flow, and wherein the adjustment of the preview of the image is based on the tracking of the region of interest.
8. The method according to claim 1, further comprising: tracking the region of interest within the field of view based on a combination of motion vector processing and optical flow.
9. The method according to claim 1, further comprising: tracking the region of interest within the field of view based on a hybrid tracker.
10. The method according to claim 9, wherein, the hybrid tracker tracks the region of interest after a downscaling operation based on a centered cropped frame.
11. The method according to claim 9, wherein, the hybrid tracker comprises: (a) one or more motion vectors associated with a current image frame, and (b) a saliency map indicating the region of interest.
12. The method according to claim 11, further Comprising: Generating the saliency map through a neural network.
13. The method according to claim 1, wherein, the preview includes a plurality of objects, and the method further includes: using the saliency map to select an object among the plurality of objects, and wherein the determination of the region of interest is based on the selected object.
14. The method according to claim 1, wherein, the sensor is one of a gyroscope or an optical image stabilization (OIS) sensor.
15. The method according to claim 1, wherein, the motion trajectory of the region of interest indicates the moving speed of the change between consecutive frames, and wherein the adjustment of the preview further includes: maintaining a smooth movement of the region of interest at or near the center of the zoomed portion between the consecutive frames of the preview.
16. The method according to claim 1, wherein, the motion trajectory of the region of interest indicates a nearly constant moving speed between consecutive frames, and wherein the adjustment of the preview of the image further includes: locking the position of the region of interest at or near the center of the zoomed portion between the consecutive frames of the preview.
17. The method according to claim 1, wherein, the determination of the region of interest further includes: receiving a user indication of the region of interest.
18. The method according to claim 1, wherein, the determination of the region of interest further includes: determining a saliency map indicating the region of interest based on a neural network.
19. The method according to claim 1, further comprising: determining that the adjusted preview is at a magnification below a threshold magnification; and transitioning from the zoom operation mode to the normal operation mode.
20. The method according to claim 1, wherein, the region of interest includes a face, and wherein the determination of the region of interest is based on a face detection algorithm.
21. A mobile device, comprising: an image capture device, the image capture device including a display screen; one or more processors; and a data storage, wherein computer-executable instructions are stored on the data storage, and when executed by the one or more processors, the computer-executable instructions cause the mobile device to perform functions including the computer-implemented method according to any one of claims 1 to 20.
22. A non-transitory computer-readable medium, the non-transitory computer-readable medium including program instructions that can be executed by one or more processors to cause the one or more processors to perform operations including the computer-implemented method according to any one of claims 1 to 20.