Stabilized object tracking at high magnification

A neural network-based image stabilization method addresses the challenges of maintaining a stable frame at high magnification by combining saliency detection and optical flow to track the region of interest, enhancing user experience through smooth and stable framing.

JP2025533832APending Publication Date: 2025-10-09GOOGLE LLC
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2025519597
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-10-04
Publication Date
2025-10-09

AI Technical Summary

Technical Problem

Image stabilization at high magnification is challenging due to the narrow field of view and difficulty in maintaining a stable frame, especially with existing optical and electronic image stabilization techniques, which suffer from hardware limitations and residual motion.

Method used

A neural network-based image stabilization method that combines saliency detection, object tracking, and optical flow to stabilize the region of interest at high magnification, using a zoom stabilization pipeline that maintains the same field of view without additional cropping, and provides a frame-in-frame viewfinder for user guidance.

Benefits of technology

Enables smooth photo and video capture at high magnification by reliably tracking and locking the region of interest, overcoming hardware limitations and maintaining stable framing without sacrificing field of view, thus improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025533832000001_ABST
    Figure 2025533832000001_ABST
Patent Text Reader

Abstract

An example method includes displaying, via a display screen of an image capture device, a preview of an image representing a field of view of the image capture device. The method includes determining a region of interest in the preview. The method includes transitioning the image capture device from a normal operating mode to a zoom operating mode. The zoom operating mode includes determining a trajectory of movement of the region of interest based on sensor data collected by a sensor associated with the image capture device, and generating an adjusted preview representing a zoomed portion of the field of view based on the determined trajectory of movement. The adjusted preview displays the region of interest at or near the center of the zoomed portion. The method includes providing the adjusted preview of the portion of the field of view.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Many modern computing devices, including mobile phones, personal computers, and tablets, include image capture devices, some of which are configured with telephoto capabilities. Summary of the Invention

[0002] The present disclosure generally relates to stabilizing an image in a viewfinder of an image capture device at high magnification. In one aspect, the image capture device may be configured to frame and track an object of interest in the narrow field of view resulting from the high magnification. Powered by a system of machine-learned components, the image capture device may be configured to stabilize and maintain the frame.

[0003] In a first aspect, a computer-implemented method is provided. The method includes displaying, via a display screen of an image capture device, a preview of an image representing a field of view of the image capture device. The method also includes determining a region of interest in the preview image. The method further includes transitioning the image capture device from a normal mode of operation to a zoom mode of operation, the zoom mode of operation including determining a trajectory of movement of the region of interest based on sensor data collected by a sensor associated with the image capture device, and generating, based on the determined trajectory of movement, an adjusted preview representing a zoomed portion of the field of view, the adjusted preview displaying the region of interest at or near the center of the zoomed portion. The method further includes providing, via the display screen, the adjusted preview of the portion of the field of view.

[0004] In a second aspect, a device is provided. The device includes one or more processors operable to perform operations. The operations include displaying, via a display screen of an image capture device, a preview of an image representing a field of view of the image capture device. The operations also include determining a region of interest in the preview image. The operations further include transitioning the image capture device from a normal mode of operation to a zoom mode of operation, the zoom mode of operation including determining a trajectory of movement of the region of interest based on sensor data collected by a sensor associated with the image capture device, and generating, based on the determined trajectory of movement, an adjusted preview representing a zoomed portion of the field of view, the adjusted preview displaying the region of interest at or near a center of the zoomed portion. The operations further include providing, via the display screen, the adjusted preview of the portion of the field of view.

[0005] In a third aspect, an article of manufacture is provided. The article of manufacture may include a non-transitory computer-readable medium having stored thereon program instructions that, when executed by one or more processors of a computing device, cause the computing device to perform operations. The operations include displaying, via a display screen of an image capture device, a preview of an image representing a field of view of the image capture device. The operations also include determining a region of interest in the preview image. The operations further include transitioning the image capture device from a normal mode of operation to a zoom mode of operation, the zoom mode of operation including determining a trajectory of movement of the region of interest based on sensor data collected by a sensor associated with the image capture device, and generating, based on the determined trajectory of movement, an adjusted preview representing a zoomed portion of the field of view, the adjusted preview displaying the region of interest at or near the center of the zoomed portion. The operations further include providing, via the display screen, the adjusted preview of the portion of the field of view.

[0006] In a fourth aspect, a system is provided, the system including: means for displaying, via a display screen of an image capture device, a preview of an image representing a field of view of the image capture device; means for determining a region of interest within the preview image; means for transitioning the image capture device from a normal mode of operation to a zoom mode of operation, the zoom mode of operation including: means for determining a trajectory of movement of the region of interest based on sensor data collected by a sensor associated with the image capture device; and means for generating, based on the determined trajectory of movement, an adjusted preview representing a zoomed portion of the field of view, the adjusted preview displaying the region of interest at or near the center of the zoomed portion; and means for providing, via the display screen, the adjusted preview of the portion of the field of view.

[0007] Other aspects, embodiments, and implementations will become apparent to those skilled in the art from a reading of the following detailed description, with appropriate reference to the accompanying drawings. [Brief explanation of the drawings]

[0008] [Figure 1] 10 is an illustration of an adjusted preview of a portion of a field of view, according to an example embodiment; [Figure 2] FIG. 10 illustrates a warning notification for stabilized object tracking, according to an example embodiment. [Figure 3A] 1 is an exemplary workflow for stabilized object tracking, according to an exemplary embodiment. [Figure 3B] 1 is an exemplary workflow for applying zoom stabilization, according to an exemplary embodiment. [Figure 4] 1 is an exemplary workflow for processing successive frames in a hybrid tracker, according to an exemplary embodiment. [Figure 5] 1 illustrates an exemplary tracking optimization process, according to an exemplary embodiment. [Figure 6] 10 illustrates another exemplary tracking optimization process, according to an exemplary embodiment. [Figure 7] 10 illustrates another exemplary tracking optimization process, according to an exemplary embodiment. [Figure 8] 1 illustrates an exemplary tracking optimization process for two regions of interest, according to an exemplary embodiment. [Figure 9] 1 illustrates an example image with stabilized object tracking, according to an example embodiment. [Figure 10] 10 illustrates another example image with stabilized object tracking, according to an example embodiment. [Figure 11] 10 illustrates another example image with stabilized object tracking, according to an example embodiment. [Figure 12] FIG. 1 illustrates the training and inference stages of a machine learning model, according to an example embodiment. [Figure 13] 1 illustrates a distributed computing architecture in accordance with an exemplary embodiment. [Figure 14] FIG. 1 is a block diagram of a computing device in accordance with an exemplary embodiment. [Figure 15] 1 illustrates a network of computing clusters arranged as a cloud-based server system, according to an example embodiment. [Figure 16] 1 is a flowchart of a method according to an example embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0009] Exemplary methods, devices, and systems are described herein. It should be understood that the words "example" and "exemplary" are used herein to mean "serving as an example, instance, or illustration." Any embodiment or feature described herein as "example" or "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments or features. Other embodiments may be utilized, and other changes may be made, without departing from the scope of the subject matter presented herein.

[0010] Accordingly, the exemplary embodiments described herein are not intended to be limiting. The aspects of the present disclosure, as generally described and illustrated in the Figures herein, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are contemplated herein.

[0011] Furthermore, unless the context suggests otherwise, the features illustrated in each of the figures may be used in combination with one another. Thus, the figures should be viewed generally as component aspects of one or more overall embodiments, with the understanding that not all of the illustrated features are required for each embodiment.

[0012] overview This application relates to image stabilization using machine learning techniques, such as, but not limited to, neural network techniques. When a user of an image capture device previews an image at a high magnification, the resulting image may not be stable, and framing an object of interest in the image may be a challenge. Furthermore, for example, after framing the object of interest, maintaining a smooth, continuous, and / or stable moving frame may be even more difficult due to the small field of view (FOV) resulting from the high magnification. Therefore, image processing-related technical challenges arise, including stabilizing the object of interest in the preview and maintaining smooth movement of the object of interest in the moving frame. Furthermore, image processing-related technical challenges arise, including, for example, smoothly transitioning the operation of the image capture device between different modes (e.g., modes corresponding to different magnifications).

[0013] Telephoto cameras are becoming increasingly popular in flagship devices. The combination of high-magnification optical zoom lenses and high-resolution image sensors has led to increased maximum magnification with each new device model. Image quality at high magnifications continues to improve. However, the extremely narrow FOV at high magnifications (e.g., approximately 1° FOV at 80x magnification) makes it difficult to frame an object of interest within the FOV, and even more so to maintain such framing while pressing the shutter button to capture an image. This can sometimes lead to a "hide-and-seek" situation, where users must search for the object of interest within the narrow field of view. Existing optical image stabilization (OIS) and / or electronic image stabilization (EIS) algorithms attempt to ameliorate this situation to some extent, but residual motion is magnified at high magnifications.

[0014] For example, if OIS is limited to approximately ±0.9° and there is no EIS, it is likely that "what you see is what you get" (WYSIWYG). However, it can be difficult to maintain a region of interest (ROI) within the frame, and it can also be difficult to find the ROI within a zoomed frame. Also, if EIS (e.g., limited by gyro and / or sensor noise) is applied, the WYSIWYG characteristics can be compromised, and a dragging effect can appear in the image. Therefore, some smart devices with cameras address this issue by reducing the maximum magnification in video mode compared to photo mode.

[0015] Generally, baseline EIS may compensate for camera shake without tracking capabilities. However, under high magnification, it may be difficult to keep a moving object in the viewfinder. Also, for example, the trajectory of the moving object may not be smooth.

[0016] Gyro- and / or OIS-based EIS can be used as a stabilization technique in some situations. This stabilization technique is based on sensors such as gyro sensing and / or OIS sensing. This allows for compensation for changes in camera pose without relying on image content, but the FOV of the output stabilized frame may be limited due to the margin used to generate a stable virtual pose. At high magnifications, hardware limitations such as gyro noise, OIS sensing noise, OIS calibration errors, and signal delays may also result in visible residual motion.

[0017] Some techniques may include detecting the position of a face in addition to the EIS technique described above. This approach may result in a stabilized frame that follows the movement of the face, but this technique is limited to faces and does not apply to general objects of interest. Also, as with the EIS approach, the frame is cropped after stabilization, which may limit the FOV of the output stabilized frame.

[0018] The technology described herein addresses these challenges by enabling smooth photo and / or video capture and an easy framing experience at high magnification. This is achieved by reliably tracking an ROI and locking the ROI at or near the center of the viewfinder or by maintaining smooth movement within the frame. While gyro and / or OIS sensor information is used to determine the movement trajectory, image-based ROI tracking effectively overcomes hardware limitations (such as gyro noise, OIS sensing noise, OIS calibration error, and signal delay).

[0019] The techniques described herein may include aspects of image-based techniques combined with techniques based on motion data and optical image stabilization (OIS) data. A neural network, such as a convolutional neural network, may be trained and applied to perform one or more aspects described herein. In some examples, the neural network may be configured as an encoder / decoder neural network.

[0020] This method targets high magnifications and utilizes the digital zoom crop margin for stabilization without additional cropping. The zoom stabilization pipeline can be configured to enable saliency map and object tracker determination with reference to the entire image sensor area or a cropped sensor area centered on the ROI, thereby achieving stabilization along the ROI as long as the ROI is located within the image sensor. As described herein, a stabilizing ROI is jointly proposed using saliency detection, object tracking, and optical flow. Potential frame delays due to the depth of the camera pipeline (e.g., 5 frames) are also addressed. We also describe a novel user interface (UI) and / or user experience (UX) design that enables a frame-in-frame viewfinder when zoom stabilization mode is active, with a bounding box moving relative to the frame to indicate a stabilized, expanded FOV relative to the entire sensor area or FOV. The described technique also maintains the same FOV defined by the user without sacrificing additional margins to stabilize the frame.

[0021] In one example, a (copy of) a neural network trained to detect salient objects may be located within a mobile computing device. The mobile computing device may include a camera capable of capturing input photographs or videos. A user of the mobile computing device may view the input photographs or videos and determine that an object in the input photographs or videos should be tracked. The input photographs or videos and motion data may be provided to a trained neural network located on the mobile computing device. In response, the trained neural network may generate a predicted output indicating an ROI having a bounding box. In another example, the trained neural network does not reside on the mobile computing device; instead, the mobile computing device provides the input photographs or videos and motion data to a remotely located trained neural network (e.g., via the Internet or other data network). The remotely located neural network may process the input photographs or videos and motion data as described above and provide an output. In another example, a non-mobile computing device may also use the trained neural network to stabilize object tracking in images and videos at high magnification, including photographs or videos not captured by the computing device's camera.

[0022] In this manner, the techniques described herein can improve image capture devices by stabilizing images and providing zoomed-in views, thereby improving the actual and / or apparent quality of the image. Improving the actual and / or apparent quality of a photo or video can be beneficial to the user experience. The flexibility of these techniques allows them to be applied to a wide variety of video types in both indoor and outdoor settings.

[0023] Techniques for image stabilization using neural networks One of the primary functions of image stabilization is to maintain an accurate and reliable tracker of the region of interest (ROI). Minimizing tracking noise and reducing noise from the gyro and / or OIS can significantly contribute to image stabilization. Additional challenges include pose changes, occlusions, and object movement into and / or out of the sensor area. Latency issues can also arise, related to the delay between image processing at the hardware layer and subsequent changes at the software or application layer. For example, pipeline depth (e.g., 5 frames) can introduce delays. For mobile camera applications, photo mode may not allow additional image cropping. Furthermore, stabilization must be able to smoothly transition between multiple camera modes.

[0024] Thus, as described herein, the zoom stabilization mode of a mobile device can capture smooth photos and / or videos at high magnification and facilitate the user's framing experience by reliably tracking and locking a user's region of interest (ROI) (or object of interest) in the center of the field of view (FOV) or smoothly moving the ROI within the frame.

[0025] Noise from the gyro and / or OIS, or noise from calibration errors, can affect smooth tracking. Challenges can also arise from the complex integration of the camera application with underlying hardware layers, such as the saliency node, rectification node, and / or EIS node. For example, at high magnifications, particularly in remosaic mode, image quality can be poor. Therefore, the zoom stabilization mode can also be configured to support binning transitions. In some embodiments, under bright light conditions (e.g., outdoors), remosaic mode may provide higher image quality than binning mode due to its higher resolution. However, under low light conditions, binning mode may provide better image quality due to poor noise performance in remosaic mode.

[0026] As described herein, zoom stabilization can have several technical advantages, such as limited power and low latency budgets for real-time photo preview. Zoom stabilization can also be configured to handle multiple objects (e.g., a herd of animals). Also, for example, zoom stabilization is compatible with existing features (e.g., HDR+, long shot video, etc.).

[0027] 1 is a diagram illustrating an adjusted preview of a portion of a field of view, according to an example embodiment. In some embodiments, the display screen of an image capture device may display a preview of an image representing the field of view of the image capture device. For example, display screen 100A may display field of view 105, which may include an object of interest, such as an image of a crescent moon 110. While the image capture device is operating at a high magnification, field of view 105 may be narrow, and small movements of the hand may cause crescent moon 110 to move out of field of view 105.

[0028] Some embodiments include determining a region of interest in a preview image. For example, there may be no ROI within the field of view, or the ROI may have moved outside the field of view. In such embodiments, background motion within the field of view may be tracked. In some embodiments, a new ROI may be detected within the field of view. For example, an ROI tracker may identify a new object. One important function of an ROI tracker is to reliably predict what a camera user intends to capture. This may be achieved individually or with a combination of user prompts and machine learning-based algorithms. For example, a tap ROI tracker in a camera application may allow a user to tap on the display screen to indicate an object and / or region of interest. Also, for example, a saliency map may be generated using a machine learning model, indicating regions of interest to the user. While existing saliency maps output a fixed-size bounding box for the ROI, the zoom stabilization described herein is configured to estimate the size of the ROI and output an appropriate bounding box for the ROI. In some embodiments, the ROI may have low confidence. In such embodiments, background motion, a central ROI, or a combination of both may be used to maintain smooth framing.

[0029] Some embodiments include transitioning an image capture device from a normal operating mode to a zoom operating mode. For example, images may be captured at different levels of zoom. At high magnifications, the field of view may be significantly narrower, and small movements of the camera may cause the field of view to suddenly change the image being captured. In some embodiments, a threshold magnification may be used to determine whether an image stabilization algorithm, such as the zoom stabilization algorithm described herein, needs to be turned on or off. For example, some cameras may be configured so that a mode switch occurs at a 15x magnification. For example, for magnifications greater than 15x, zoom stabilization mode may be turned on, and for magnifications less than 15x, zoom stabilization mode may be turned off. Additional and / or alternative magnifications may be utilized.

[0030] In some embodiments, the zoom mode of operation may include determining an adjusted preview representing a zoomed portion of the field of view based on sensor data collected by a sensor associated with the image capture device, with the adjusted preview displaying a region of interest at or near the center of the zoomed portion. As shown on display screen 100B, field of view 105 may be displayed within a view bounded by outer frame 140. Inner frame 145 may be displayed within outer frame 140. Inner frame 145 may include an object of interest, a crescent moon 110. Accordingly, display screen 100B may display adjusted preview 115 (e.g., a magnified view) representing a zoomed portion (e.g., within inner frame 145) of the field of view (e.g., field of view 105 displayed within outer frame 140). As shown, display screen 100B displays magnified view 115 of the portion of the field of view, including a magnified view of crescent moon 110A. Even though the image of crescent moon 110 exhibits non-smooth motion within the original field of view 105, inner frame 145 is a stabilized image and magnified view 115 is a stabilized zoom view of crescent moon 110A.

[0031] Display screen 100B may include additional functionality associated with a camera application. For example, a user may have available multiple modes, including motion mode 120, portrait mode 122, video mode 126, and video bokeh mode 128. As shown, the camera application may be in camera mode 124. Camera mode 124 may provide additional functionality, such as a reverse icon 130 for activating a reverse camera view, a trigger button 132 for capturing a preview image, and a photo stream icon 134 for accessing a database of captured images. Also, for example, a magnification slider 138 may be displayed, and a user may select a magnification by moving a virtual object along magnification slider 138. In some embodiments, a user may adjust the magnification using the display screen (e.g., by moving two fingers outward away from each other on display screen 100B), and magnification slider 138 may automatically display the magnification.

[0032] As shown, the magnification slider 138 may be at 30x. For a camera configured to switch modes at 15x magnification, if the magnification slider 138 is moved beyond 15x, the camera may switch from normal mode to zoom stabilization mode and image stabilization may be automatically activated. In such a case, an object of interest may be determined and a magnified view 115, an outer frame 140, an inner frame 145, etc. may be displayed.

[0033] The camera application may also provide various user adjustment functions to adjust one or more image characteristics (e.g., brightness, hue, contrast, shadows, highlights, global brightness adjustment for the entire image, local brightness adjustment for a ROI, etc.) For example, in some embodiments, slider 136A may be provided to adjust characteristic A, slider 136B may be provided to adjust characteristic B, and in some embodiments, slider 136C may be provided to adjust characteristic C.

[0034] In some embodiments, the zoom mode of operation may include determining a trajectory of movement of a region of interest based on sensor data collected by a sensor associated with the image capture device, and generating an adjusted preview representing the zoomed portion of the field of view based on the determined trajectory of movement, the adjusted preview displaying the region of interest at or near the center of the zoomed portion.

[0035] 2 illustrates a warning notification for stabilized object tracking, according to an example embodiment. Display screen 200A (or display screen 200B) may include additional functionality related to a camera application. For example, a user may have available multiple modes, including motion mode 120, portrait mode 122, video mode 126, and video bokeh mode 128. As shown, the camera application may be in camera mode 124. Camera mode 124 may provide additional functionality, such as a reverse icon 130 for activating a reverse camera view, a trigger button 132 for capturing a preview image, and a photo stream icon 134 for accessing a database of captured images. For example, a magnification slider 138 may also be displayed, allowing a user to select a magnification by moving a virtual object along the magnification slider 138. In some embodiments, the user may adjust the magnification using display screen 200A (or display screen 200B) (e.g., by moving two fingers outward away from each other on display screen 200A (or display screen 200B)), and magnification slider 138 may automatically display the magnification.

[0036] As shown, magnification slider 138 may be 15x. For cameras configured to switch modes at 15x magnification, when magnification slider 138 reaches 15x, the camera may switch from normal mode to zoom stabilization mode and image stabilization may be automatically activated. In such a case, an object of interest may be determined and a magnified view 115, outer frame 140, inner frame 145, etc. may be displayed. Zoom ratios are for illustrative purposes only and may vary depending on device and / or system configuration.

[0037] At high magnifications, manipulating the camera to track a moving object can result in pause-move-pause type motion, which leads to large residual motion in conventional EIS. This is due to the large salient term E, as explained in Equation 1 below: Protrusion This may be triggered by the zoomed-in portion of the field of view. Some embodiments include providing an image overlay via the display screen that displays a representation of the zoomed-in portion relative to the field of view. For example, a frame-in-frame function stabilizes the image and guides the user. Some embodiments include determining a bounding box for the region of interest, and providing the image overlay includes providing the region of interest contained within the bounding box. As shown, the display screen 200A displays the field of view within an outer frame 140, and an inner frame 145 within the outer frame 140 surrounds an image of the ROI, moon 150. An adjusted preview, e.g., zoomed view 115, corresponds to the inner frame 145 and displays a zoomed-in view of moon 150. As shown, the image of moon 150 is centered within the inner frame 145.

[0038] Some embodiments include detecting when the region of interest is approaching a boundary of the image overlay. For example, when the camera moves and / or due to object movement, the image of the moon 150 may move within the field of view. In such embodiments, the zoom stabilization mode may maintain a stabilized enlarged view 115 in which the image of the moon is centered within the enlarged view 115. In some embodiments, movement of the camera and / or the object of interest may cause the object of interest to approach a boundary of the inner frame 145, as shown on the display screen 200B. The image of the moon may be centered within the enlarged view 115, but the image may be closer to the boundary of the inner frame 145.

[0039] Such embodiments also include providing a notification to the user indicating that the region of interest is approaching a boundary of the image overlay. For example, the boundary of inner frame 145 may turn red, begin flashing, the device may vibrate, an audio notification may be provided, audio instructions may be generated, or an arrow may be displayed indicating the direction in which to move the camera to move the image of the moon away from the boundary of the image overlay. Thus, the frame-in-frame may be configured to guide the user in finding the object of interest, and the zoom stabilization mode may be configured to issue a notification to the user when the object of interest is closer to the boundary.

[0040] Generally, the inner frame 145 may be configured to automatically slide with the object's movement within the field of view to detect and track salient objects without the user having to center the object. In some embodiments, a zoom stabilization algorithm may stabilize and track the object while maintaining it at or near the center of the enlarged view 115. Generally, the outer frame 140 displays the field of view 105, while the region defined by the inner frame 145 can be cropped and displayed as the enlarged view 115. As described herein, an object of interest can be identified and tracked, the inner frame 145 can be determined and cropped to generate the enlarged view 115, and a zoomed-in view of the object of interest can be displayed at or near the center of the enlarged view 115 while maintaining a stabilized image with smooth motion. Thus, the object of interest can be locked at or near the center or can be displayed as moving smoothly within the frame. Generally, if the ROI is stationary and / or moving at a constant speed, the object of interest is locked to the center. In some embodiments, the trajectory of the region of interest's motion exhibits variable speeds of movement between successive frames. In such embodiments, adjusting the preview includes maintaining smooth motion of the region of interest at or near the center of the zoomed portion between successive frames of the preview. For example, the object of interest may appear to move smoothly within the frame as the ROI moves at variable speeds. The smooth motion tracks the movement of the ROI. Such an approach allows for high magnification photography and / or video without the need for a tripod mount.

[0041] Generally, the relative position of the image of the moon 150 remains stable within the enlarged preview 115, while the frames 140 and 145 track the moon, so that the preview displayed to the user remains well-centered and / or the moon moves smoothly within the frames 140 and 145 while remaining relatively stable. This is a significant improvement over the display capabilities of the image capture device, as the stabilized image of the moon may move out of the enlarged preview 115, causing the moon to disappear from view when the user attempts to manually track the moon by moving the device. However, the zoom stabilization algorithm tracks the moon and can alert the user when the moon is at or near the boundary of the inner frame 145. This is particularly useful at high magnifications, where the preview FOV can be very narrow. For example, at 30X, an object of interest may easily move out of the preview, but zoom stabilization tracking can detect the object of interest, frame it, smoothly track it, and alert the user.

[0042] 3A is an example workflow 300A for stabilized object tracking, according to an example embodiment. Some image capture devices include a hardware abstraction layer (HAL) that connects a higher-level camera framework application programming interface (API) at the camera application (APP) layer to the underlying camera driver and hardware.

[0043] At 302, telepreview (or zoom stabilization mode) may be activated if the magnification exceeds a threshold magnification (e.g., 15x). Zoom ratios are for illustrative purposes only and may vary depending on device and / or system configuration.

[0044] At 304, an input tracker may be initialized. At step 1, the system may proceed to block 306 to determine whether a tap ROI is being tracked. In some embodiments, determining the region of interest includes receiving a user designation of the region of interest. At step 2, the system may determine that a tap ROI is being tracked, and at block 308, the user designation of the ROI may be detected and an initial ROI center may be extracted from the user's tap.

[0045] In step 3, the system may determine that the tap ROI is not being tracked. Some embodiments include generating a saliency map using a neural network. For example, in block 310, a saliency detection algorithm and / or a face detection algorithm may be activated to identify the ROI. For example, if a user tap is not detected, a machine learning (ML)-based saliency map may be determined in block 310. The saliency may be applied directly to the sensor area without further cropping. This allows for new salient ROI detection that may exist outside the user's final zoom FOV. The algorithm may then extract an initial ROI center from the ML-based saliency map. In some embodiments, the system may estimate the ROI size by motion vectors from neighboring frames.

[0046] In some embodiments, the algorithm tracks the ROI by using (i) a combination of motion vectors and optical flow, or (ii) an ML hybrid tracker. The motion vector process may provide better accuracy when the ROI transformation is rigid without occlusion, while the hybrid tracker may have higher reliability when the transformation is non-rigid or when occlusion occurs.

[0047] In some embodiments, the hybrid tracker uses a cropped frame centered on the ROI to enable tracking of the ROI after downsizing. In some embodiments, the crop ratio can be set to a target digital magnification and / or a more adaptive ratio that maintains the ROI within the cropped frame. The process then proceeds to step 4, the ROI region to be tracked.

[0048] At block 312, the hybrid tracker may jointly stabilize the ROI using the combined saliency detection, object tracker, and optical flow to obtain a reliable and accurate ROI. The algorithm uses the non-hybrid tracker when the delta difference in the ROI center between the hybrid and non-hybrid trackers is small, and can smoothly weight the hybrid tracker using additional weights when the delta difference between the two trackers is large. Generally speaking, the non-hybrid tracker, or ILK tracker, uses a motion vector map (similar to that for optical flow, but with a patch size of 64x64) to find the ROI displacement between adjacent frames. As used herein, the term "ILK" generally refers to a reverse search version of the Lucas-Kanade algorithm for optical flow estimation. The term "ILK tracker" refers to an optical flow-based tracker. Joint stabilization leads to stabilization of potential frame delays due to the depth of the camera pipeline (e.g., 5 frames). For example, information from the hybrid tracker may be combined with motion vectors from ILK to predict the ROI. This resolves potential frame delays due to the depth of the camera pipeline. The process then proceeds to step 5 where the ROI center and ROI confidence are determined.

[0049] At block 314, EIS input, such as gyro and / or OIS data, and frame metadata are provided to a zoom stabilization algorithm. For example, real-time filtering and lightweight optimization is performed to stabilize the frame with the gyro and / or OIS data and ROI input. The process then proceeds to step 6, zoom stabilization algorithm 316.

[0050] In block 316, the algorithm may generate stabilized frames with EIS inputs (e.g., gyro sensors, OIS sensors, etc.) based on real-time filtering and lightweight optimization to obtain motion trajectories while overcoming hardware limitations (gyro noise, OIS sensing noise, OIS calibration errors, signal delay, etc.). In some embodiments, as described above, the low-resolution full sensor frame, stabilized frame center coordinates, and crop ratio may be provided to a user interface to generate a frame-in-frame viewfinder.

[0051]

number

[0052] The gyro / OIS noise 340 is provided to camera motion analysis 342. The camera motion analysis 342 may determine a trajectory of movement of the region of interest. For example, spatial information regarding the position of one or more objects of interest (e.g., faces, bounding boxes, etc.) may be extracted from each captured image frame. Some embodiments include determining a motion vector associated with the previous image frame and the current image frame. For example, the motion vector may be generated from the spatial information by averaging two adjacent frames. For example, a motion vector may be extracted between successive frames for each 64x64 patch. This results in enhanced tracking capabilities.

[0053] For magnifications above a threshold (e.g., 15x), a user tap indicating an ROI may override automatic tracking. In the absence of a user tap, the system may use a face detection algorithm to detect a face or a saliency model to detect an object of interest. In the absence of a face or object of interest, motion vectors may be based on the center of the frame. Motion vectors between the previous and current frames are determined to obtain an approximate model of movement for each frame.

[0054]

number

[0055] Also, for example, the ROI center and confidence 348 provide virtual transform stabilization to the video stabilization 346. The term w2E ROI Center is introduced to stabilize the central ROI. The term w3EP rotrusion corresponds to a pause-move-pause type of motion, which results in large residual motion in traditional EIS. Based on the actual camera pose 344, the actual ROI position can be determined, and optimization can be performed based on the combined image information. Based on inputs from camera motion analysis 342 and ROI center and confidence 348, video stabilization 346 provides a virtual camera pose 350 to warping block 354. Image 352 is also provided to warping block 354.

[0056] 3A, in step 7, a stabilized frame 318 is generated using frame warping from the zoom stabilization algorithm 316. Also in step 8, the bounding box of the stabilized region is provided to frame-in-frame UI feedback 320. If a reliable ROI is not available from the previous frame, center cropping may be performed on the frame; otherwise, ROI-centered cropping is performed.

[0057] In step 9, the algorithm provides the low-resolution full sensor frame (before stabilization) to block 320. In block 320, the algorithm provides the low-resolution full sensor frame, stabilized frame center coordinates, and crop ratio to the UI to generate a frame-in-frame viewfinder and enable dynamic preview bounding box visualization.

[0058] In general, non-hybrid trackers may be more accurate. Hybrid trackers may offer a better tradeoff between occlusion handling and accuracy. In some embodiments, non-hybrid and hybrid trackers may be combined to achieve optimal reliability and accuracy ROI.

[0059] To maintain stabilization quality, gyro noise and / or OIS noise that scales with magnification may be suppressed, rotation effect correction, seamless transitions, etc. may be achieved.

[0060] If the preview includes multiple objects, the most salient object among those objects may be identified. For example, if there are multiple objects in the preview, attention may be focused on one object instead of switching between the multiple objects. For example, face detection-type matching may be performed to identify a face of interest among several faces in the image. Also, for example, a machine learning model (e.g., a visual saliency model) may generate saliency scores for multiple candidate salient objects detected in the image. A zoom stabilization algorithm may select objects with high saliency scores as salient objects. The saliency model may be trained with training data indicative of a user's interests and / or preferences, and the trained saliency model may predict objects of interest to the user.

[0061] For example, the visual saliency model may be trained based on a training dataset including training scenes, training sequences, and / or events. For example, the training dataset may include images (e.g., digital photographs) including user-drawn bounding boxes containing visually salient regions (e.g., regions where one or more objects of particular interest to the user may reside). Based on the training dataset, the visual saliency model may predict visually salient regions within the images. For example, as a result of training, the visual saliency model may generate a visual saliency heatmap for a given image and generate bounding boxes that enclose regions with the greatest probability of visual saliency (e.g., the highest saliency score). One or more processors may calculate the visual saliency heatmap during background operation of the device. In some embodiments, the visual saliency heatmap may indicate the magnitude of the probability of visual saliency on a scale from black to white, with white indicating a high probability of saliency and black indicating a low probability of saliency.

[0062] In some embodiments, the visual saliency heatmap includes a bounding box that encloses the region in the image that has the greatest probability of visual saliency. When multiple objects of interest are present in the scene, causing the visual saliency model to identify multiple salient regions in the captured image, the visual saliency model can be trained to generate a bounding box that encloses the salient region closest to the center of the captured image. This training technique assumes that the user is interested in the most centrally located object in the image. Alternatively, the visual saliency model can be trained to generate a bounding box that encloses all objects of interest in the captured image.

[0063] The image capture device may perform operations under the direction of an autozoom manager that implements various aspects of the zoom stabilization mode. In some embodiments, the autozoom manager may perform several steps to calibrate the image capture device, either automatically or in response to a received trigger signal, including, for example, a gesture (e.g., tap, press) performed by a user on an input / output device. For example, the autozoom manager may receive one or more captured images from an image sensor of the image capture device and, utilizing a visual saliency model, generate a visual saliency heatmap using the one or more captured images. The visual saliency model may also output a bounding box that encloses areas with the highest probability of visual saliency.

[0064] If there are multiple objects of interest in the preview, causing the visual saliency model to identify multiple salient regions in the preview, the visual saliency model can be trained to generate a bounding box that encloses the salient region closest to the center of the preview. This training technique assumes that the user is interested in the most centrally located object in the image. Alternatively, the visual saliency model can be trained to generate a bounding box that encloses all objects of interest in the preview, or the object of interest with the highest saliency score.

[0065] Some embodiments include determining that the adjusted preview is at a magnification below a threshold magnification. For example, the magnification may be less than 15x. Such embodiments include transitioning from a zoom mode of operation to a normal mode of operation. Accordingly, zoom stabilization mode may be deactivated and normal mode of operation may be activated. In some embodiments, object tracking may no longer be performed in the normal mode of operation. In some embodiments, object tracking may continue to be performed, but frame-in-frame views and / or magnified views of portions of the full field of view (e.g., a cropped portion of the full FOV corresponding to the framing of the ROI) may no longer be generated and / or displayed.

[0066] 4 is an example workflow for processing consecutive frames in a hybrid tracker, according to an example embodiment. A touch ROI 405, which is an ROI based on a user's tap, can be detected in frame tN. The hybrid tracker 410 tracks the ROI and generates an ROI 415 shown in block 415. t-N The ROI at frame tN can be determined as (IH(tN)).

[0067] The path of the hybrid tracker is determined as follows. IH(t)=f(f(f(f(H,tN),t-N1),…t-1) (Equation 2)

[0068] where f(H,t) = H(t) + ILK(t+1), f(f(x)) denotes the self-composition of function f, and ILK(t) represents the coordinates of the ROI based on the hybrid tracker with ILK motion vectors to predict the position of the ROI in frame t.

[0069] In block 435, the non-hybrid (e.g., optical flow-based ILK) tracker 430 calculates the full-frame motion vectors (MVs) of frame t and the motion vectors of the ROI in frame t−1 as a function of the ROI. t-1 In some embodiments, the ROI may be determined at block 440 as tTo determine (IO(t)), voting may be applied in step 1. The path of the non-hybrid tracker is determined as follows: IO(t)=O(t-1)+ILK(t) (Equation 3)

[0070] where IO(t) represents the coordinates of the ROI based on the non-hybrid or ILK tracker. To determine the final ROI, a threshold condition 420 may be checked. O(t)=|IO(tN)-IH(tN)|>Threshold? (Equation 4)

[0071] where O(t) represents the finalized ROI coordinates. In some embodiments, a selection between IH(t) and I0(t) may be made. For example, if it is determined that the threshold condition 420 is not met, the system may select IH(t) provided by Equation 2 as the selected ROI. As shown in block 425, this ROI may be based on a combination of IH(t) and the ILK from a non-hybrid tracker. For example, if the results of the two tracker methods, the hybrid tracker combined with the ILK motion vectors, and the non-hybrid (or ILK tracker), differ significantly, this indicates occlusion and / or non-rigid transformations. Therefore, IH(t) is selected because the hybrid tracker is more robust when dealing with occlusion / non-rigid transformations. In such situations, the selection O(t) = IH(t) is used.

[0072] Also, for example, if it is determined that the threshold condition 420 is met, the system may select the ROI as IO(t) as determined by Equation 3. For example, if the results of two tracker methods, i.e., a hybrid tracker combined with ILK motion vectors and a non-hybrid tracker (or ILK tracker), are close to each other, the accuracy of the ILK tracker is higher and its selection O(t)=IO(t) is used.

[0073] The selected ROI can be set as a new ROI for the iterative process.

[0074] Blocks 450-460 show the process with an object of interest represented by the letter "A." In block 450, the letter "A" is shown in frame t-2. In block 455, the letter "A" is shown in a new position in the next frame t-1. Thus, the hybrid tracker calculates: H(t-1)=H(t-2)+ILK(t-1) (Equation 5)

[0075] At block 460, it is shown that the letter "A" has moved further to the right in the next frame t. Therefore, the hybrid tracker calculates: H(t)=H(t-1)+ILK(t) (Equation 6)

[0076] In this example, calculations using three consecutive frames are shown, but a similar iterative approach applies to N consecutive frames. In some embodiments, the hybrid tracker may downsize the frame to 320x240, and a cropped frame with the ROI center may be initialized as the hybrid tracker input to enable tracking of the object of interest after downsizing.

[0077] As described herein, an inner frame (e.g., inner frame 145) may be cropped from the entire FOV to determine a magnified stabilization view. The crop ratio of such a crop may be set to a target digital magnification or a more adaptive desired ratio to ensure that the ROI is within the cropped frame.

[0078]

number

[0079]

number

[0080] In the formula, A -1 represents the inverse matrix of matrix A. Here, Kv represents the intrinsic parameter matrix of the camera corresponding to the virtual camera pose, and R v is the predicted rotation of the virtual camera pose, and K p denotes the intrinsic parameter matrix of the camera corresponding to the actual camera pose, and R p is the predicted rotation of the actual camera pose. The weighting term may be smaller in the pitch / yaw axis than traditional EIS, but may maintain the same weight on roll. Thus, the tracking term can then dominate the pitch / yaw compensation to reduce residual motion caused by gyro / OIS noise.

[0081] In some embodiments, a two-step optimization may be performed.

[0082] Step 1: A target point t at which the ROI is located within the stabilized frame v Find.

[0083] Step 2: Find the virtual camera pose.

[0084] 5 illustrates an exemplary tracking optimization process, according to an exemplary embodiment. Three consecutive input frames are shown with an image of a cat. Input Frame 1 505 shows the cat on the left side of the display with a bounding box around the cat's face. Input Frame 2 510 shows the cat in the center of the display with a bounding box around the cat's face. An initial motion vector is generated based on the positions of the consecutive bounding boxes in Input Frame 1 505 and Input Frame 2 510, as indicated by the dashed lines.

[0085]

number

[0086] Here, the operation "." represents multiplication, and c1, c2, and c3 are positive weight coefficients such that c1+c2+c3=1.

[0087] Figure 5 shows the first term c1.t in Equation 9. v,prev Show how is determined, and t v,prev represents the coordinates of the virtual target center in the previous frame. v,prev is t v The coordinate of t v,prev This ensures that the center of the virtual ROI is stable.

[0088] 6 illustrates another exemplary tracking optimization process, according to an exemplary embodiment. Input frames 605, 615, and 625 show images of a cat moving from left to center to right, respectively. Point t v is determined in successive frames to generate stabilized frames 610, 620, and 630 corresponding to input frames 605, 615, and 625, respectively. p Here, x p represents the position of the ROI in the real pose (e.g., the unstabilized frame), so the term c2.x p can be adjusted to control how closely the virtual target follows the real position. For example, input frame 615 corresponds to c2=0, and target point t v,1 may be determined. For example, diff(t( v,prev ),x p ) is less than a second threshold thresh2, then c2=0, where "diff" stands for difference or distance. The input frame 625 corresponds to the case c2>0, and the target point t v,2 As shown by stabilized frames 610, 620, and 630, the cat's position is maintained at or near the center of the frame, even though the cat's position moves within input frames 605, 615, and 620. Point t v can be solved from the closed-form equation Equation 9.

[0089] 7 illustrates another exemplary tracking optimization process, according to an exemplary embodiment. Input frames 705, 715, and 725 are shown with an image of a cat on the left, and stabilized frames 710, 720, and 730 corresponding to input frames 705, 715, and 725, respectively. Again, point t v can be solved from Equation 9, which is a closed-form equation. This example shows the third term c3.center in Equation 9. In this example, diff( v,prev , Center)*digital_zoom_factor is less than the third threshold thresh3, then c3=0. v ,Center) is t v and the absolute center of the stabilized frame. Therefore, the condition that this distance is less than a threshold ensures that the virtual ROI is not at or near the boundary of the stabilized frame. Stabilized Frame 2 720 illustrates the case where c3>0, and Stabilized Frame 3 730 illustrates the case where c3=0.

[0090]

number

[0091] In the formula, x p represents the position of the ROI in the real pose, and t v represents the position of the ROI in the virtual posture, and A -1 represents the inverse matrix of matrix A. Here, K v represents the intrinsic parameter matrix of the camera corresponding to the virtual camera pose, and R v is the predicted rotation of the virtual camera pose, and K p denotes the intrinsic parameter matrix of the camera corresponding to the actual camera pose, and R p is the predicted rotation of the actual camera pose.

[0092] FIG. 9 shows an example image with stabilized object tracking, according to an example embodiment. An initial FOV 905 is shown along with a bounding box 915 showing an object of interest (e.g., an airplane). A warping mesh 910 is shown. For example, after determining the virtual camera pose, a stabilization mesh, such as the physical-to-virtual camera warping mesh 910, can be generated by determining source and destination quadrilaterals for each horizontal stripe. In some embodiments, the warping mesh 910 may be cropped from the entire initial FOV 905 to generate an enlarged FOV 920. The object of interest, in this example, the airplane 925, is shown in zoomed-in view within the enlarged FOV 920. As shown, the airplane 925 is clearly visible and can be smoothly tracked based on its motion vectors in successive frames.

[0093] FIG. 10 shows another example image with stabilized object tracking, according to an example embodiment. An initial FOV 1005 is shown with a bounding box 1015 depicting an object of interest (e.g., a post displaying a house number). A warping mesh 1010 is shown. For example, after determining the virtual camera pose, a stabilization mesh, such as the physical-to-virtual camera warping mesh 1010, can be generated by determining source and destination quadrilaterals for each horizontal stripe. In some embodiments, the warping mesh 1010 may be cropped from the entire initial FOV 1005 to generate the enlarged FOV 1020. The object of interest, in this example, a post 1025 displaying a house number, is shown in zoomed-in view within the enlarged FOV 1020. As shown, the post 1025 is clearly visible, and the house number "705" can be discerned.

[0094] FIG. 11 shows another example image with stabilized object tracking, according to an example embodiment. An initial FOV 1105 is shown with a bounding box 1115 showing an object of interest (e.g., the moon). A warping mesh 1110 is shown. For example, after determining the virtual camera pose, a stabilization mesh, such as the physical-to-virtual camera warping mesh 1110, can be generated by determining source and destination quadrilaterals for each horizontal stripe. In some embodiments, the warping mesh 1110 may be cropped from the entire initial FOV 1105 to generate the enlarged FOV 1120. The object of interest, in this example, the moon 1125, is shown in zoomed-in view within the enlarged FOV 1120. As shown, the moon 1125 is clearly visible and can be smoothly tracked based on its motion vectors in successive frames.

[0095] As described herein, the telephoto magnification (Tele RM) may be 9.4x, and the magnification for zoom stabilization may be 15x. Zoom ratios are for illustrative purposes only and may vary depending on the device and / or system configuration. The image capture device may transition between two modes: a normal mode without zoom stabilization and a zoom stabilization mode. In some embodiments, the zoom stabilization mode may be controlled by a combination of magnification and mesh interpolation. In some embodiments, the transition may be between a basic EIS mode and a centered ROI-based zoom stabilization mode. In general, the transition is seamless with or without an ROI tracking term. For example, the transition between a centered ROI and an ROI source may be seamless by adjusting the virtual ROI target and retracking the transition between different ROI sources. Also, the transition between a binning mode and a remosaic mode may be seamless, for example, by performing YUV frame-in-frame cropping in the camera application and adjusting the EIS margins accordingly.

[0096] Training a machine learning model to generate inferences / predictions FIG. 12 shows a diagram 1200 illustrating the training phase 1202 and inference phase 1204 of trained machine learning model(s) 1232, according to an example embodiment. Some machine learning techniques involve training one or more machine learning algorithms with an input set of training data to recognize patterns in the training data and provide output inferences and / or predictions regarding (the patterns in) the training data. The resulting trained machine learning algorithms may be referred to as trained machine learning models. For example, FIG. 12 shows a training phase 1202 in which one or more machine learning algorithms 1220 are trained with training data 1210 to become trained machine learning models 1232. Then, during the inference phase 1204, the trained machine learning model 1232 may receive input data 1230 and one or more inference / prediction requests 1240 (perhaps as part of the input data 1230) and, in response, provide one or more inferences and / or predictions 1250 as output.

[0097] Thus, trained machine learning model(s) 1232 may include one or more models of one or more machine learning algorithms 1220. The machine learning algorithm(s) 1220 may include, but are not limited to, artificial neural networks (e.g., convolutional neural networks described herein, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems). The machine learning algorithm(s) 1220 may be supervised or unsupervised and may implement any suitable combination of online and offline learning.

[0098] In some examples, the machine learning algorithm(s) 1220 and / or the trained machine learning model(s) 1232 may be accelerated using an on-device coprocessor, such as a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), and / or an application-specific integrated circuit (ASIC). Such on-device coprocessors may be used to accelerate the machine learning algorithm(s) 1220 and / or the trained machine learning model(s) 1232. In some examples, the trained machine learning model(s) 1232 may be trained to provide inferences on, reside on, execute, and / or otherwise perform inferences for a particular computing device.

[0099] During the training phase 1202, the machine learning algorithm(s) 1220 may be trained by providing at least the training data 1210 as training inputs using unsupervised, supervised, semi-supervised, and / or reinforcement learning techniques. Unsupervised learning involves providing some (or all) of the training data 1210 to the machine learning algorithm(s) 1220, and the machine learning algorithm(s) 1220 determining one or more output inferences based on the provided portion (or all) of the training data 1210. Supervised learning involves providing some (or all) of the training data 1210 to the machine learning algorithm(s) 1220, and the machine learning algorithm(s) 1220 determining one or more output inferences based on the provided portion (or all) of the training data 1210, where the output inference(s) are either accepted or corrected based on the correct results associated with the training data 1210. In some examples, the supervised learning of the machine learning algorithm(s) 1220 may be governed by a set of rules and / or a set of labels for the training inputs, which may be used to correct the inferences of the machine learning algorithm(s) 1220.

[0100] Semi-supervised learning involves having correct results for some, but not all, of the training data 1210. During semi-supervised learning, supervised learning is used for portions of the training data 1210 that have correct results, and unsupervised learning is used for portions of the training data 1210 that do not have correct results. Reinforcement learning involves the machine learning algorithm(s) 1220 receiving a reward signal related to a prior inference, where the reward signal may be a numerical value. During reinforcement learning, the machine learning algorithm(s) 1220 can output an inference and receive a reward signal in response, where the machine learning algorithm(s) 1220 are configured to attempt to maximize the numerical value of the reward signal. In some examples, reinforcement learning also utilizes a value function that provides a numerical value representing the expected sum of the numerical values ​​provided by the reward signal over time. In some examples, the machine learning algorithm(s) 1220 and / or the trained machine learning model(s) 1232 may be trained using other machine learning techniques, including, but not limited to, incremental learning and curriculum learning.

[0101] In some examples, the machine learning algorithm(s) 1220 and / or the trained machine learning model(s) 1232 can use transfer learning techniques. For example, transfer learning techniques may include the trained machine learning model(s) 1232 being pre-trained on a set of data and additionally trained using training data 1210. More specifically, the machine learning algorithm(s) 1220 may be pre-trained on data from one or more computing devices and the resulting trained machine learning model provided to computing device CD1, which is intended to execute the trained machine learning model during inference stage 1204. Then, during training stage 1202, the pre-trained machine learning model may be additionally trained using training data 1210, which may be derived from kernel data and non-kernel data of computing device CD1. This further training of the machine learning algorithm(s) 1220 and / or the pre-trained machine learning model using training data 1210 of CD1's data may be performed using either supervised learning or unsupervised learning. The training phase 1202 may be completed once the machine learning algorithm(s) 1220 and / or pre-trained machine learning model(s) have been trained with at least the training data 1210. The resulting trained machine learning model(s) may be utilized as at least one of the trained machine learning model(s) 1232.

[0102] In particular, once the training phase 1202 is complete, the trained machine learning model(s) 1232 may be provided to the computing device if not already on the computing device. The inference phase 1204 may begin after the trained machine learning model(s) 1232 are provided to the computing device CD1.

[0103] During the inference stage 1204, the trained machine learning model(s) 1232 may receive input data 1230 and generate and output one or more corresponding inferences and / or predictions 1250 regarding the input data 1230. Thus, the input data 1230 may be used as input to the trained machine learning model(s) 1232 to provide the corresponding inference(s) and / or prediction(s) 1250 to kernel and non-kernel components. For example, the trained machine learning model(s) 1232 may generate the inference(s) and / or prediction(s) 1250 in response to one or more inference / prediction requests 1240. In some examples, the trained machine learning model(s) 1232 may be executed by another piece of software. For example, the trained machine learning model(s) 1232 may be executed by an inference or prediction daemon so that they are readily available to provide inferences and / or predictions upon request. Input data 1230 may include data from computing device CD1 running trained machine learning model(s) 1232 and / or input data from one or more computing devices other than CD1.

[0104] The training data 1210 may include images (e.g., digital photographs) that include user-drawn bounding boxes containing regions of visual salience (e.g., regions where one or more objects of particular interest to the user may reside).

[0105] The input data 1230 may include one or more captured images or previews of images. Other types of input data are possible as well.

[0106] The inference(s) and / or prediction(s) 1250 may include output images, bounding boxes enclosing regions of interest with the greatest probability of visual saliency, and / or other output data generated by trained machine learning model(s) 1232 operating on the input data 1230 (and training data 1210). In some examples, trained machine learning model(s) 1232 may use the output inference(s) and / or prediction(s) 1250 as input feedback 1260. The trained machine learning model(s) 1232 may also utilize past inferences as input to generate new inferences.

[0107] A convolutional neural network, such as a visual saliency model, may be an example of machine learning algorithm(s) 1220. After training, the trained version of the convolutional neural network may be an example of trained machine learning model(s) 1232. In this approach, an example of estimation / prediction request(s) 1240 may be a request to predict a region of interest in a preview of an image, and a corresponding example of estimation and / or prediction(s) 1250 may be an output image having a bounding box containing the region of visual saliency.

[0108] In some examples, one computing device CD_SOLO may include a trained version of convolutional neural network 100, perhaps after training the convolutional neural network. Computing device CD_SOLO may then receive a request to predict a region of interest in a preview of an image and use the trained version of the convolutional neural network to generate an output image having a bounding box containing the region of visual salience.

[0109] In some examples, two or more computing devices CD_CLI and CD_SRV may be used to provide the output image; for example, a first computing device CD_CLI can generate and send a request to a second computing device CD_SRV to predict a region of interest within a preview of the image. CD_SRV can then, perhaps after training a convolutional neural network, use the trained version of the convolutional neural network to generate an output image having a bounding box containing regions of visual salience and, in response to a request from CD_CLI, predict the region of interest within the preview of the image. Then, upon receiving a response to the request, CD_CLI can provide the requested region of interest using a user interface and / or display.

[0110] Data Network Example 13 illustrates a distributed computing architecture 1300, according to an example embodiment. The distributed computing architecture 1300 includes server devices 1308, 1310 configured to communicate with programmable devices 1304a, 1304b, 1304c, 1304d, and 1304e via a network 1306. The network 1306 may correspond to a local area network (LAN), a wide area network (WAN), a WLAN, a WWAN, a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 1306 may also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.

[0111] While FIG. 13 shows only five programmable devices, the distributed application architecture may service tens, hundreds, or thousands of programmable devices. Furthermore, programmable devices 1304a, 1304b, 1304c, 1304d, and 1304e (or any additional programmable devices) may be any type of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mountable device (HMD), a network terminal, a mobile computing device, etc. In some examples, as shown by programmable devices 1304a, 1304b, 1304c, and 1304e, the programmable devices may be directly connected to the network 1306. In other examples, as shown by programmable device 1304d, the programmable devices may be indirectly connected to the network 1306 through an associated computing device, such as programmable device 1304c. In this example, programmable device 1304c may serve as an associated computing device for passing electronic communications between programmable device 1304d and the network 1306. In other examples, such as shown by programmable device 1304e, the computing device may be part of and / or within a vehicle such as a car, truck, bus, boat or ship, airplane, etc. In other examples not shown in Figure 13, the programmable device may be connected both directly and indirectly to the network 1306.

[0112] The server devices 1308, 1310 may be configured to perform one or more services requested by the programmable devices 1304a-1304e. For example, the server devices 1308 and / or 1310 may provide content to the programmable devices 1304a-1304e. The content may include, but is not limited to, web pages, hypertext, scripts, binary data such as compiled software, images, audio, and / or video. The content may include compressed and / or uncompressed content. The content may be encrypted and / or decrypted. Other types of content are possible as well.

[0113] As another example, server devices 1308 and / or 1310 may provide programmable devices 1304a-1304e with access to software for database, search, computation, graphics, audio, video, World Wide Web / Internet usage, and / or other functions. Many other examples of server devices are possible as well.

[0114] Computing Device Architecture 14 is a block diagram of an exemplary computing device 1400, according to an example embodiment. In particular, the computing device 1400 shown in FIG. 14 may be configured to perform at least one function of and / or related to method 1600.

[0115] The computing device 1400 may include a user interface module 1401, a network communication module 1402, one or more processors 1403, data storage 1404, one or more cameras 1418, one or more sensors 1420, and a power system 1422, all of which may be linked to each other via a system bus, network, or other connection mechanism 1405.

[0116] The user interface module 1401 may be operable to transmit data to and / or receive data from external user input / output devices. For example, the user interface module 1401 may be configured to transmit data to and / or receive data from user input devices such as a touchscreen, a computer mouse, a keyboard, a keypad, a touchpad, a trackball, a joystick, a voice recognition module, and / or other similar devices. The user interface module 1401 may also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices now known or later developed. The user interface module 1401 may also be configured to generate audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, earphones, and / or other similar devices. User interface module 1401 may further comprise one or more haptic devices that may generate tactile output, such as vibration and / or other output detectable by touch and / or physical contact with computing device 1400. In some examples, user interface module 1401 may be used to provide a graphical user interface (GUI) for utilizing computing device 1400.

[0117] The network communication module 1402 may include one or more devices providing one or more wireless interfaces 1407 and / or one or more wired interfaces 1408 configurable to communicate over a network. The wireless interface(s) 1407 may include one or more wireless transmitters, receivers, and / or transceivers, such as a Bluetooth® transceiver, a Zigbee® transceiver, a Wi-Fi™ transceiver, a WiMAX™ transceiver, an LTE™ transceiver, and / or other type of wireless transceiver configurable to communicate over a wireless network. The wired interface(s) 1408 may include one or more wired transmitters, receivers, and / or transceivers, such as an Ethernet® transceiver, a Universal Serial Bus (USB) transceiver, or similar transceiver configurable to communicate over a twisted pair wire, coaxial cable, fiber optic link, or similar physical connection to a wired network.

[0118] In some examples, the network communication module 1402 can be configured to provide reliable, secure, and / or authenticated communications. For each communication described herein, information to facilitate reliable communications (e.g., guaranteed message delivery) may be provided, perhaps as part of the message header and / or footer (e.g., packet / message ordering information, encapsulation header and / or footer, size / time information, and transmission verification information such as a cyclic redundancy check (CRC) and / or parity check value). Communications may be protected (e.g., encoded or encrypted) and / or decrypted / decoded using one or more cryptographic protocols and / or algorithms, such as, but not limited to, the Data Encryption Standard (DES), the Advanced Encryption Standard (AES), the Rivest-Shamir-Adelman (RSA) algorithm, the Diffie-Hellman algorithm, a secure socket protocol such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS), and / or the Digital Signature Algorithm (DSA). Other encryption protocols and / or algorithms may be used similar to or in addition to those listed herein to protect (and subsequently decrypt / decode) communications.

[0119] The one or more processors 1403 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPUs), graphics processing units (GPUs), application-specific integrated circuits, etc.) The one or more processors 1403 may be configured to execute computer-readable instructions 1406 contained in data storage 1404 and / or other instructions described herein.

[0120] Data storage 1404 may include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of the one or more processors 1403. The one or more computer-readable storage media may include volatile and / or non-volatile storage components, such as optical storage, magnetic storage, organic storage, or other memory or disk storage, which may be integrated, in whole or in part, with at least one of the one or more processors 1403. In some examples, data storage 1404 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other examples, data storage 1404 may be implemented using two or more physical devices.

[0121] Data storage 1404 may include computer-readable instructions 1406 and possibly additional data. In some examples, data storage 1404 may include storage necessary to execute at least some of the methods, scenarios, and techniques described herein and / or at least some of the functionality of the devices and networks described herein. In some examples, data storage 1404 may include storage for trained neural network models 1412 (e.g., models of trained convolutional neural networks). In particular, in these examples, computer-readable instructions 1406 may include instructions that, when executed by processor(s) 1403, enable computing device 1400 to provide some or all of the functionality of trained neural network models 1412.

[0122] In some examples, computing device 1400 may include one or more cameras 1418. Camera(s) 1418 may include one or more image capture devices, such as still and / or video cameras, equipped to capture light and record the captured light into one or more images; i.e., camera(s) 1418 may generate image(s) of the captured light. The one or more images may be one or more still images and / or one or more images utilized in video capture. Camera(s) 1418 may capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as light of one or more other frequencies.

[0123] In some examples, computing device 1400 may include one or more sensors 1420. Sensors 1420 may be configured to measure conditions within computing device 1400 and / or conditions in the environment of computing device 1400 and provide data regarding these conditions.For example, sensors 1420 may include: (i) sensors for acquiring data about computing device 1400, such as, but not limited to, a thermometer for measuring the temperature of computing device 1400, a battery sensor for measuring the power of one or more batteries in power system 1422, and / or other sensors for measuring the state of computing device 1400; (ii) identification sensors for identifying other objects and / or devices, such as, but not limited to, a radio frequency identification (RFID) reader, a proximity sensor, a one-dimensional barcode reader, a two-dimensional barcode (e.g., a quick response (QR) code) reader, and a laser tracker, which may be configured to read identifiers such as RFID tags, barcodes, QR codes, and / or other devices and / or objects configured to read and provide at least identification information; and (iii) tilt sensors, gyroscopes, accelerometers. (iv) sensors that measure the position and / or movement of the computing device 1400, such as, but not limited to, a Doppler sensor, a GPS device, a sonar sensor, a radar device, a laser displacement sensor, and a compass; (iv) environmental sensors that acquire data indicative of the environment of the computing device 1400, such as, but not limited to, an infrared sensor, an optical sensor, a light sensor, a biosensor, a capacitive sensor, a touch sensor, a temperature sensor, a wireless sensor, a radio sensor, a movement sensor, a microphone, a sound sensor, an ultrasonic sensor, and / or a smoke sensor; and / or (v) force sensors that measure one or more forces (e.g., inertial forces and / or G-forces) acting about the computing device 1400, such as, but not limited to, one or more sensors that measure force, torque, ground force, friction in one or more dimensions, and / or a zero moment point (ZMP) sensor that identifies the ZMP and / or the location of the ZMP. Many other examples of sensors 1420 are possible as well.

[0124] The power supply system 1422 may include one or more batteries 1424 and / or one or more external power interfaces 1426 for providing power to the computing device 1400. When electrically coupled to the computing device 1400, each battery of the one or more batteries 1424 may serve as a source of stored power for the computing device 1400. The one or more batteries 1424 of the power supply system 1422 may be configured to be portable. Some or all of the one or more batteries 1424 may be easily removable from the computing device 1400. In other examples, some or all of the one or more batteries 1424 may be internal to the computing device 1400 and therefore may not be easily removable from the computing device 1400. Some or all of the one or more batteries 1424 may be rechargeable. For example, the rechargeable batteries may be recharged via a wired connection between the batteries and another power source, such as by one or more power sources external to the computing device 1400 and connected to the computing device 1400 via one or more external power interfaces. In other examples, some or all of the one or more batteries 1424 may be non-rechargeable batteries.

[0125] The one or more external power interfaces 1426 of power system 1422 may include one or more wired power interfaces, such as a USB cable and / or a power cord, that enable a wired power connection to one or more power sources external to computing device 1400. The one or more external power interfaces 1426 may include one or more wireless power interfaces, such as a Qi wireless charger, that enable a wireless power connection to one or more external power sources, such as via a Qi wireless charger. Once a power connection to an external power source is established using one or more external power interfaces 1426, computing device 1400 can draw power from the external power source via the established power connection. In some examples, power system 1422 may include associated sensors, such as battery sensors or other types of power sensors associated with one or more batteries.

[0126] Cloud-based Server FIG. 15 illustrates a cloud-based server system according to an example embodiment. In FIG. 15, the functionality of a convolutional neural network and / or computing devices may be distributed among computing clusters 1509a, 1509b, and 1509c. Computing cluster 1509a may include one or more computing devices 1500a, a cluster storage array 1510a, and a cluster router 1511a connected by a local cluster network 1512a. Similarly, computing cluster 1509b may include one or more computing devices 1500b, a cluster storage array 1510b, and a cluster router 1511b connected by a local cluster network 1512b. Similarly, computing cluster 1509c may include one or more computing devices 1500c, a cluster storage array 1510c, and a cluster router 1511c connected by a local cluster network 1512c.

[0127] In some embodiments, each of computing clusters 1509a, 1509b, and 1509c may have an equal number of computing devices, an equal number of cluster storage arrays, and an equal number of cluster routers. However, in other embodiments, each computing cluster may have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each computing cluster may depend on one or more computing tasks assigned to each computing cluster.

[0128] In computing cluster 1509a, for example, computing device 1500a may be configured to perform various computing tasks, such as convolutional neural networks, confidence training, and / or computing device functions. In one embodiment, the functions of the convolutional neural networks, confidence training, and / or computing devices may be distributed among one or more of computing devices 1500a, 1500b, and 1500c. Computing devices 1500b and 1500c in each computing cluster 1509b and 1509c may be configured similarly to computing device 1500a in computing cluster 1509a. Meanwhile, in some embodiments, computing devices 1500a, 1500b, and 1500c may be configured to perform different functions.

[0129] In some embodiments, computing tasks and stored data associated with the convolutional neural network and / or computing devices may be distributed across computing devices 1500a, 1500b, and 1500c based at least in part on the processing requirements of the convolutional neural network and / or computing devices, the processing capabilities of computing devices 1500a, 1500b, and 1500c, the latency of network links between computing devices within each computing cluster and between the computing clusters themselves, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design goals of the overall system architecture.

[0130] The cluster storage arrays 1510a, 1510b, 1510c of the computing clusters 1509a, 1509b, 1509c may be data storage arrays that include disk array controllers configured to manage read and write access to groups of hard disk drives. The disk array controllers, alone or in combination with their respective computing devices, may also be configured to manage backup or redundant copies of data stored in the cluster storage arrays to protect against disk drive or other cluster storage array failures and / or network failures that prevent one or more computing devices from accessing one or more cluster storage arrays.

[0131] Just as the functionality of convolutional neural networks and / or computing devices may be distributed across computing devices 1500a, 1500b, 1500c of computing clusters 1509a, 1509b, 1509c, various active and / or backup portions of these components may be distributed across cluster storage arrays 1510a, 1510b, 1510c. For example, some cluster storage arrays may be configured to store some portions of data for convolutional neural networks and / or computing devices, while other cluster storage arrays may store other portions(s) of data for convolutional neural networks and / or computing devices. Also, for example, some cluster storage arrays may be configured to store data for a first convolutional neural network, while other cluster storage arrays may store data for a second and / or third convolutional neural network. Additionally, some cluster storage arrays may be configured to store backup versions of data stored in other cluster storage arrays.

[0132] Cluster routers 1511 a, 1511 b, 1511 c in computing clusters 1509 a, 1509 b, 1509 c may include network equipment configured to provide internal and external communications for the computing clusters. For example, cluster router 1511 a in computing cluster 1509 a may include one or more Internet switching and routing devices configured to provide (i) local area network communications between computing device 1500 a and cluster storage array 1510 a via local cluster network 1512 a, and (ii) wide area network communications between computing cluster 1509 a and computing clusters 1509 b and 1509 c via wide area network link 1513 a to network 1306. Cluster routers 1511b and 1511c may include network equipment similar to cluster router 1511a, and cluster routers 1511b and 1511c may perform networking functions for computing clusters 1509b and 1509b similar to that performed by cluster router 1511a for computing cluster 1509a.

[0133] In some embodiments, the configuration of the cluster routers 1511a, 1511b, 1511c may be based at least in part on the data communication requirements of the computing devices and cluster storage arrays, the data communication capabilities of the network equipment within the cluster routers 1511a, 1511b, 1511c, the latency and throughput of the local cluster networks 1512a, 1512b, 1512c, the latency, throughput, and cost of the wide area network links 1513a, 1513b, 1513c, and / or other factors that may contribute to the cost, speed, fault tolerance, resilience, efficiency, and / or other design criteria of the mitigation system architecture.

[0134] Exemplary Methods of Operation 16 illustrates a method 1600 according to an example embodiment. The method 1600 may include various blocks or steps. The blocks or steps may be performed individually or in combination. The blocks or steps may be performed in any order and / or sequentially or in parallel. Additionally, blocks or steps may be omitted or added to the method 1600.

[0135] The blocks of the method 1600 may be performed by various elements of the computing device 1400 as shown and described with reference to FIG.

[0136] Block 1610 includes displaying, via a display screen of the image capture device, a preview of an image representative of the field of view of the image capture device.

[0137] Block 1620 involves determining a region of interest in the preview of the image.

[0138] Block 1630 includes transitioning the image capture device from a normal operating mode to a zoom operating mode, which includes determining a motion trajectory of a region of interest based on sensor data collected by a sensor associated with the image capture device, and generating an adjusted preview representing a zoomed portion of the field of view based on the determined motion trajectory, the adjusted preview displaying the region of interest at or near the center of the zoomed portion.

[0139] Block 1640 includes providing, via a display screen, an adjusted preview of the portion of the field of view.

[0140] Some embodiments include providing, via a display screen, an image overlay that displays a representation of the zoomed portion relative to the field of view.

[0141] Some embodiments include determining a bounding box for the region of interest, and providing the image overlay includes providing the region of interest contained within the bounding box.

[0142] Some embodiments include determining one or more of: (i) a low-resolution version of the displayed image; (ii) coordinates of the adjusted region of interest within the adjusted preview; or (iii) a crop ratio. Such embodiments also include generating an image overlay to enable dynamic visualization of the bounding box.

[0143] Some embodiments include detecting when the region of interest is approaching a boundary of the image overlay, and such embodiments also include providing a notification to a user indicating that the region of interest is approaching a boundary of the image overlay.

[0144] Some embodiments include determining a motion vector associated with the previous image frame and the current image frame, and such embodiments also include determining a size of the region of interest based on the determined motion vector.

[0145] Some embodiments include determining an optical flow corresponding to the region of interest. Such embodiments also include tracking the region of interest within the portion of the field of view based on the determined optical flow, and adjusting the preview of the image based on the tracking of the region of interest.

[0146] Some embodiments include tracking a region of interest within a field of view based on a combination of motion vector processing and optical flow. For example, once it has been determined that the transformation associated with the region of interest is rigid and free of occlusion, a combination of motion vector processing and optical flow may be used to track the region of interest.

[0147] Some embodiments include tracking a region of interest within the field of view based on a hybrid tracker. For example, upon determining that a transformation associated with the region of interest is non-rigid or involves occlusion, a hybrid tracker may be used to track the region of interest. In some embodiments, the hybrid tracker tracks the region of interest after a downsizing operation based on a center cropped frame. In some embodiments, the hybrid tracker includes (a) one or more motion vectors associated with the current image frame and (b) a saliency map indicating the region of interest.

[0148] Some embodiments include generating the saliency map by a neural network.

[0149] In some embodiments, the preview includes a plurality of objects, and the method includes selecting one object from the plurality of objects using a saliency map, and determining the region of interest is based on the selected object.

[0150] In some embodiments, the sensor is one of a gyroscope or an optical image stabilization (OIS) sensor.

[0151] In some embodiments, the trajectory of the movement of the region of interest exhibits a variable rate of movement between successive frames, and in such embodiments, adjusting the preview includes maintaining smooth movement of the region of interest at or near the center of the zoomed portion between successive frames of the preview.

[0152] In some embodiments, the trajectory of the movement of the region of interest exhibits a substantially constant velocity of movement between successive frames, and in such embodiments, adjusting the preview includes locking the position of the region of interest at or near the center of the zoomed portion between successive frames of the preview.

[0153] In some embodiments, determining the region of interest includes receiving a user designation of the region of interest.

[0154] In some embodiments, determining the region of interest includes determining a saliency map indicative of the region of interest based on a neural network.

[0155] Some embodiments include determining that the adjusted preview is at a magnification below a threshold magnification. Such embodiments include transitioning from a zoom mode of operation to a normal mode of operation.

[0156] The particular configuration shown in the figures should not be considered limiting. It should be understood that other embodiments may include more or fewer of each element shown in a given figure. Furthermore, some of the illustrated elements may be combined or omitted. Furthermore, example embodiments may include elements not shown in the figures.

[0157] Steps or blocks representing the processing of information may correspond to circuitry that can be configured to perform specific logical functions of the methods or techniques described herein. Alternatively or additionally, steps or blocks representing the processing of information may correspond to modules, segments, or portions of program code (including associated data). The program code may include one or more instructions executable by a processor to implement specific logical functions or operations in the method or technique. The program code and / or associated data may be stored on any type of computer-readable medium, such as a storage device, including a disk, hard drive, or other storage medium.

[0158] Computer-readable media may also include non-transitory computer-readable media, such as register memory, processor cache, and computer-readable media that store data for a short period of time, such as random access memory (RAM). Computer-readable media may also include non-transitory computer-readable media that store program code and / or data for a longer period of time. Thus, computer-readable media may include, for example, secondary or persistent long-term storage, such as read-only memory (ROM), optical or magnetic disks, compact disk read-only memory (CD-ROM), etc. Computer-readable media may also be any other volatile or non-volatile storage system. Computer-readable media may be considered, for example, to be a computer-readable storage medium or a tangible storage device.

[0159] While various examples and embodiments have been disclosed, other examples and embodiments will be apparent to those skilled in the art. The various examples and embodiments disclosed are for purposes of illustration and are not intended to be limiting, the true scope being indicated by the following claims.

Claims

1. 1. A computer-implemented method comprising: displaying, via a display screen of an image capture device, a preview of an image representative of a field of view of said image capture device; determining a region of interest in the preview of the image; transitioning the image capture device from a normal mode of operation to a zoom mode of operation, the zoom mode of operation comprising: determining a trajectory of movement of the region of interest based on sensor data collected by a sensor associated with the image capture device; generating an adjusted preview representing the zoomed portion of the field of view based on the determined trajectory of movement, the adjusted preview displaying the region of interest at or near the center of the zoomed portion, the computer-implemented method further comprising: providing, by the display screen, the adjusted preview of the portion of the field of view. A computer-implemented method.

2. providing, by the display screen, an image overlay displaying a representation of the zoomed portion relative to the field of view. The method of claim 1.

3. determining a bounding box of the region of interest; providing the image overlay includes providing the region of interest contained within the bounding box. The method of claim 2.

4. Determining one or more of: (i) a low-resolution version of the displayed image; (ii) coordinates of the adjusted region of interest within the adjusted preview; or (iii) a crop ratio; generating the image overlay to enable dynamic visualization of the bounding box. The method of claim 3.

5. Detecting when the region of interest is approaching a boundary of the image overlay; providing a user notification indicating that the region of interest is approaching the boundary of the image overlay. The method of claim 2.

6. determining a motion vector relating to a previous image frame and a current image frame; determining a size of the region of interest based on the determined motion vector. The method of claim 1.

7. determining an optical flow corresponding to the region of interest; tracking the region of interest within the portion of the field of view based on the determined optical flow; the adjusting of the preview of the image is based on the tracking of the region of interest. The method of claim 1.

8. and tracking the region of interest within the field of view based on a combination of motion vector processing and optical flow. The method of claim 1.

9. tracking the region of interest within the field of view based on a hybrid tracker. The method of claim 1.

10. The method of claim 9 , wherein the hybrid tracker tracks the region of interest after a downsizing operation based on a center cropped frame.

11. the hybrid tracker includes: (a) one or more motion vectors associated with a current image frame; and (b) a saliency map indicating the region of interest.

10. The method of claim 9.

12. generating the saliency map by a neural network. The method of claim 11.

13. The preview includes a plurality of objects, and the method further comprises: further comprising using a saliency map to select an object from the plurality of objects; The method of claim 1 , wherein the determining the region of interest is based on the selected object.

14. The method of claim 1 , wherein the sensor is one of a gyroscope or an optical image stabilization (OIS) sensor.

15. The motion trajectory of the region of interest exhibits a variable speed of movement between successive frames, and the adjustment of the preview includes: The method of claim 1 , further comprising maintaining smooth motion of the region of interest at or near the center of the zoomed portion between the successive frames of the preview.

16. The motion trajectory of the region of interest exhibits a substantially constant rate of movement between successive frames, and the adjustment of the preview on the screen comprises: The method of claim 1 , further comprising locking the position of the region of interest at or near the center of the zoomed portion between the successive frames of the preview.

17. determining the region of interest includes: The method of claim 1 , further comprising receiving a user designation of the region of interest.

18. determining the region of interest includes: The method of claim 1 , further comprising determining a saliency map indicative of the region of interest based on a neural network.

19. determining that the adjusted preview is at a magnification below a threshold magnification; transitioning from the zoom operation mode to the normal operation mode. The method of claim 1.

20. The method of claim 1 , wherein the region of interest includes a human face, and wherein the determining the region of interest is based on a face detection algorithm.

21. A mobile device, an image capture device including a display screen; one or more processors; and a data storage storing computer-executable instructions that, when executed by the one or more processors, cause the mobile device to perform functions comprising the computer-implemented method of any one of claims 1 to 20. Mobile devices.

22. A non-transitory computer-readable medium comprising program instructions executable by one or more processors to cause the one or more processors to perform operations comprising the computer-implemented method of any one of claims 1 to 20.

Citation Information

Patent Citations

  • Image processing apparatus, imaging apparatus and program

    JP2015012493A

  • Imaging apparatus, imaging apparatus control method, imaging apparatus control program, and storage medium

    JP2015043558A

  • Determination device, determination method, and determination program

    JP2018022332A

  • Systems and methods for performing automatic zooming

    JP2018538712A

  • Explanatory sentence creation device, object information representation system, and explanatory sentence creation method

    JP2020013427A