Methods, apparatuses, media, and products for processing one or more frames

By identifying the region of interest during video capture and using cropping and scaling techniques, the problem of a fixed size of the target object in a frame sequence is solved, achieving stable display and tracking of the object in the video.

CN116018616BActive Publication Date: 2026-04-14QUALCOMM INC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2021-04-26
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

During video capture, it is difficult to maintain a fixed size of the target object in a frame sequence, especially when the target object moves relative to the camera. Existing technologies often require manual adjustment or cannot effectively maintain the size and position of the object.

Method used

By identifying the region of interest in a frame sequence, subsequent frames are cropped and scaled to maintain a fixed size of the target object. Smoothing functions are used to control the position and size changes of the object in the frame sequence, which can be applied to single-camera or multi-camera systems.

Benefits of technology

It achieves a fixed size and position of the target object during video capture, reducing the tediousness of manual adjustments and improving video quality and the stability of object tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116018616B_ABST
    Figure CN116018616B_ABST
Patent Text Reader

Abstract

Techniques for processing one or more frames are provided. For example, a region of interest can be determined in a first frame of a sequence of frames. The region of interest in the first frame includes an object having a certain size in the first frame. A portion of a second frame (occurring after the first frame in the sequence of frames) can be cropped and scaled so that the object in the second frame has the same size (and in some cases is located at the same location) as the object in the first frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates generally to video analytics, and more specifically to techniques and systems for maintaining a consistent (e.g., fixed or nearly fixed) size of a target object in one or more frames (e.g., in video analytics, for recording video, and for other purposes). Background Technology

[0002] Many devices and systems allow a scene to be captured by generating images (or frames) and / or video data (including multiple frames). For example, a camera or a computing device that includes a camera (e.g., a mobile device such as a mobile phone or smartphone that includes one or more cameras) can capture a sequence of frames of a scene. In another example, an Internet Protocol camera (IP camera) is a digital video camera that can be used for surveillance or other applications. Unlike analog closed-circuit television (CCTV) cameras, IP cameras can send and receive data via computer networks and the Internet.

[0003] Image and / or video data can be captured and processed by such devices and systems (e.g., mobile devices, IP cameras, etc.) and can be output for consumption (e.g., displayed on this device and / or other devices). In some cases, image and / or video data can be captured by such devices and systems and output for processing and / or consumption by other devices. Summary of the Invention

[0004] In some examples, techniques and systems for processing one or more frames of image or video data to maintain a fixed size of a target object (also called an object of interest) in one or more frames are described. According to at least one illustrative example, a method for using one or more frames is provided. The method includes: determining a region of interest in a first frame of a frame sequence, the region of interest in the first frame including an object of a certain size in the first frame; cropping a portion of a second frame of the frame sequence, the second frame appearing after the first frame in the frame sequence; and scaling the portion of the second frame based on the size of the object in the first frame.

[0005] In another example, an apparatus for processing one or more frames is provided, the apparatus comprising: a memory configured to store at least one frame; and one or more processors implemented in a circuit and coupled to the memory. The one or more processors are configured to: determine a region of interest in a first frame of a frame sequence, the region of interest in the first frame including an object of a certain size in the first frame; crop a portion of a second frame of the frame sequence, the second frame appearing after the first frame in the frame sequence; and scale the portion of the second frame based on the size of the object in the first frame.

[0006] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: determine a region of interest in a first frame of a frame sequence, the region of interest in the first frame including an object of a certain size in the first frame; crop a portion of a second frame of the frame sequence that appears after the first frame of the frame sequence; and scale the portion of the second frame based on the size of the object in the first frame.

[0007] In another example, an apparatus for processing one or more frames is provided. The apparatus includes: components for determining a region of interest in a first frame of a frame sequence, the region of interest in the first frame including an object of a certain size in the first frame; components for cropping a portion of a second frame of the frame sequence, the second frame appearing after the first frame in the frame sequence; and components for scaling the portion of the second frame based on the size of the object in the first frame.

[0008] In some aspects, the methods, apparatus, and computer-readable media described above further include: receiving user input corresponding to the selection of an object in the first frame; and determining a region of interest in the first frame based on the received user input. In some aspects, the user input includes touch input provided using a touch interface of the device.

[0009] In some aspects, the methods, apparatus, and computer-readable media described above further include: determining points of an object region defined for an object in the second frame; and cropping and scaling a portion of the second frame if the points of the object region are located at the center of the cropped and scaled portion.

[0010] In some respects, the point of the object region is the center point of the object region. In other cases, the object region is a bounding box (or other bounded region). The center point can be the center point of the bounding box (or other region), or the center point of the object (e.g., the centroid or center point of the object).

[0011] In some respects, portions of the second frame are scaled based on the size of the objects in the first frame so that the objects in the second frame have the same size as the objects in the first frame.

[0012] In some aspects, the methods, apparatus, and computer-readable media described above further include: determining a first length associated with an object in the first frame; determining a second length associated with an object in the second frame; determining a scaling factor based on a comparison between the first length and the second length; and scaling a portion of the second frame based on the scaling factor.

[0013] In some aspects, the first length is the length of a first object region determined for objects in the first frame, and the second length is the length of a second object region determined for objects in the second frame. In some aspects, the first object region is a first bounding box, and the first length is the diagonal length of the first bounding box, and the second object region is a second bounding box, and the second length is the diagonal length of the second bounding box.

[0014] In some respects, scaling a portion of the second frame based on the scaling factor makes the second object region in the cropped and scaled portion have the same size as the first object region in the first frame.

[0015] In some aspects, the methods, apparatus, and computer-readable media described above further include: determining points of a first object region generated for an object in the first frame; determining points of a second object region generated for an object in the second frame; using a smoothing function based on the points of the first object region and the points of the second object region to determine a motion factor for the object, wherein the smoothing function controls changes in the position of the object in multiple frames of the sequence of frames; and cropping a portion of the second frame based on the motion factor.

[0016] In some aspects, the point of the first object region is the center point of the first object region, and the point of the second object region is the center point of the second object region.

[0017] In some aspects, the smoothing function includes a movement function that determines the position of a point of the corresponding object region in each of the multiple frames of the frame sequence based on a statistical measure of object movement.

[0018] In some aspects, the methods, apparatus, and computer-readable media described above further include: determining a first length associated with an object in the first frame; determining a second length associated with an object in the second frame; determining a scaling factor for the object based on a comparison between the first length and the second length and based on a smoothing function using the first length and the second length, wherein the smoothing function controls the variation in size of the object across multiple frames in the frame sequence; and scaling a portion of the second frame based on the scaling factor.

[0019] In some aspects, the smoothing function includes a shift function that is used to determine the length associated with an object in each of the multiple frames of the frame sequence based on a statistical metric of the object size.

[0020] In some aspects, the first length is the length of a first bounding box generated for an object in the first frame, and the second length is the length of a second bounding box generated for an object in the second frame.

[0021] In some aspects, the first length is the diagonal length of the first bounding box, and the second length is the diagonal length of the second bounding box.

[0022] In some respects, scaling a portion of the second frame based on the scaling factor makes the second bounding box in the cropped and scaled portion have the same size as the first bounding box in the first frame.

[0023] In some respects, cropping and scaling of a portion of the second frame keeps the object centered within that second frame.

[0024] In some aspects, the methods, apparatus, and computer-readable media described above also include: detecting and tracking objects in one or more frames of the frame sequence.

[0025] In some aspects, the device includes a camera (e.g., an IP camera), a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, or other devices. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data.

[0026] The content of this invention is neither intended to identify key or essential features of the claimed subject matter nor to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by referring to the appropriate portions of the complete specification of this patent, any or all of the accompanying drawings, and each claim.

[0027] The foregoing, along with other features and embodiments, will become more apparent from the following description, claims, and drawings. Attached Figure Description

[0028] The illustrative embodiments of this application are described in detail below with reference to the accompanying drawings:

[0029] Figure 1 This is a block diagram illustrating an example architecture of an image capture and processing system based on some examples;

[0030] Figure 2 This is a block diagram illustrating an example of a system including a video source and a video analysis system, based on some examples;

[0031] Figure 3 These are examples of video analytics systems that process video frames, based on some examples.

[0032] Figure 4 This is a block diagram illustrating an example of a blob detection system based on some examples;

[0033] Figure 5 This is a block diagram illustrating an example of an object tracking system based on some examples;

[0034] Figure 6A This is another diagram illustrating examples of machine learning-based object detection and tracking systems based on some examples;

[0035] Figure 6B This is a diagram illustrating an example of an upsampling component in a machine learning-based object detection and tracking system, based on some examples.

[0036] Figure 6C This is a diagram illustrating an example of the backbone architecture for a machine learning-based tracking system, based on several examples.

[0037] Figure 7 This is a diagram illustrating an example of a machine learning-based object classification system based on some examples;

[0038] Figure 8A This is a diagram illustrating an example of a system including frame cropping and scaling systems, based on some examples;

[0039] Figure 8B This is a diagram illustrating an example of a frame cropping and scaling system based on some examples;

[0040] Figure 8C This is a diagram illustrating an example of the frame cropping and scaling process based on some examples;

[0041] Figure 9A This is a flowchart illustrating another example of the frame cropping and scaling process based on some examples; Figure 9B This is a flowchart illustrating another example of the frame cropping and scaling process based on some examples;

[0042] Figure 10A This is a diagram illustrating an example of the initial frame of a video based on some examples;

[0043] Figure 10B This is shown based on some examples in Figure 10A A diagram illustrating examples of subsequent frames of the video that appear after the initial frame;

[0044] Figure 11 It is a diagram illustrating various motion models based on some examples;

[0045] Figure 12 This is a flowchart illustrating an example of a process for performing image stabilization, based on some examples;

[0046] Figure 13A This is a diagram illustrating examples of the processes used to perform various aspects of an automatic zoom function, based on some examples;

[0047] Figure 13B This is a diagram illustrating an example of a process for performing an additional aspect of an automatic zoom function, based on some examples;

[0048] Figure 13C This is a diagram illustrating another example of the process for performing various aspects of an automatic zoom function, based on some examples;

[0049] Figure 13D This is a diagram illustrating an example of a process for performing an additional aspect of an automatic zoom function, based on some examples;

[0050] Figure 14 This is a diagram illustrating an example of a Gaussian filter smoothing function based on some examples;

[0051] Figure 15 This is a diagram illustrating an example of a Fibonacci filter smoothing function based on some examples;

[0052] Figure 16 This is a diagram illustrating an example of the zoom process in a camera production line, based on some examples;

[0053] Figure 17 This is a diagram illustrating an example of zoom delay for a camera pipeline, based on some examples;

[0054] Figure 18 This is a flowchart illustrating an example of a process for processing one or more frames, based on some examples;

[0055] Figures 19 to 23 These are simulated images illustrating the use of the cropping and scaling techniques described herein, based on some examples;

[0056] Figure 24 This is a diagram illustrating examples of machine learning-based object detection and tracking systems, based on several examples.

[0057] Figure 25 This is a flowchart illustrating an example of a pipeline that switches camera lenses based on some examples;

[0058] Figure 26 This is a flowchart illustrating an example of a camera lens switching process based on some examples;

[0059] Figures 27 to 36 This is a diagram illustrating examples of the use of the camera lens switching technique described herein, based on some examples;

[0060] Figures 37 to 41 These are simulated images illustrating the use of the camera lens switching techniques described herein, based on some examples;

[0061] Figure 42 This is a block diagram illustrating examples of deep learning networks based on some examples;

[0062] Figure 43 This is a block diagram illustrating an example of a convolutional neural network based on some examples;

[0063] Figure 44 This is a diagram illustrating an example of a Cifar-10 neural network based on some examples;

[0064] Figures 45A to 45C This is a diagram illustrating an example of a single-shot object detector based on some examples;

[0065] Figures 46A to 46C This is a diagram illustrating examples of the YOLO detector, which you only look once;

[0066] Figure 47 This is a diagram illustrating an example of a system used to implement some aspects described herein. Detailed Implementation

[0067] Certain aspects and embodiments of this disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some can be applied in combination, as will be apparent to those skilled in the art. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of embodiments of this application. However, it will be apparent that various embodiments can be practiced without these specific details. The accompanying drawings and description are not intended to be limiting.

[0068] The following description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of exemplary embodiments will provide those skilled in the art with a feasible description for carrying out the exemplary embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0069] An image capture device (e.g., a camera or a device that includes a camera) is a device that uses an image sensor to receive light and capture image frames such as still images or video frames. The terms “image,” “image frame,” and “frame” are used interchangeably herein. A variety of image capture and image processing settings can be used to configure the camera of an image capture device. Different settings result in images with different appearances. Camera settings such as ISO, exposure time, aperture size, f / stop, shutter speed, focus, and gain are determined and applied before or during the capture of one or more image frames. For example, settings or parameters can be applied to an image sensor to capture one or more image frames. Other camera settings can configure post-processing of one or more image frames, such as changes to contrast, brightness, saturation, sharpness, color levels, curves, or colors. For example, settings or parameters can be applied to a processor (e.g., an image signal processor or ISP) to process one or more image frames captured by an image sensor.

[0070] The camera may include or communicate with a processor (such as an ISP) that can receive and process one or more image frames from an image sensor. For example, raw image frames captured by the camera sensor may be processed by the ISP to generate a final image. In some examples, the ISP may process the image frames using multiple filters or processing blocks applied to them, such as demosaicing, gain adjustment, white balance adjustment, color balance or correction, gamma compression, tone mapping or adjustment, noise reduction or noise filtering, edge enhancement, contrast adjustment, intensity adjustment (such as darkening or brightening), and others. In some examples, the ISP may include a machine learning system (e.g., one or more neural networks and / or other machine learning components) capable of processing image frames and outputting processed image frames.

[0071] In various scenarios (e.g., motion imaging, video analytics, and other use cases), it may be desirable to maintain the size of the region of interest and / or the object of interest (or target object) frame-by-frame from the frame sequence, even if the region of interest and / or object moves relative to one or more cameras in the captured frame sequence. For example, when imaging a person playing soccer in a video capture scene, it may be desirable to maintain the person's constant size throughout the video, even if the person moves relative to the camera (e.g., laterally toward and away from the camera). In another example, regarding video analytics, it may be desirable to maintain the size of the tracked object (e.g., a delivery person) throughout an entire video clip captured by one or more Internet Protocol (IP) camera systems.

[0072] Image capture devices have increased effective zoom range. For example, a multi-camera system can be designed to allow a larger zoom range than a single camera's digital zoom range. However, when a user attempts to record video of a moving object (e.g., people playing football) and has tuned the camera zoom to give the object the desired size in a frame, the object's size ratio (the size of the object relative to the frame, referred to as the object size to frame ratio) will dynamically change as the object moves. As the object moves relative to one or more cameras in a captured frame sequence, it can be difficult to maintain the desired object size in the frame sequence (e.g., the size of the object in the initial frame when video capture is first initiated). For example, manually changing the object size to frame ratio during video capture can be cumbersome for the user. Tracking (e.g., automatically tracking) the subject can also be difficult during video recording.

[0073] This document describes systems, apparatus, processes (also referred to as methods), and computer-readable media (collectively, the “systems and techniques”) for maintaining a fixed size (referred to as a “target fixed-size feature”) of a target object in a frame sequence. The frame sequence can be video, a group of sequentially captured images, or other frame sequences. For example, the systems and techniques described herein can determine a region of interest (ROI) in the first frame (or initial frame). In some cases, the user can select the first frame. For example, in some examples, the user can select any frame from a video as a starting point. In some examples, the systems and techniques can determine the ROI based on the user’s selection of the ROI or objects within the ROI. In some cases, the user’s selection can be based on user input provided using a user interface (e.g., a device’s touchscreen, an electronic drawing tool, a gesture-based user interface, a voice-based user interface, or other user interface). In some examples, the systems and techniques can automatically determine the ROI based on object detection and / or recognition techniques. For example, the systems and techniques can detect and / or recognize people in a frame and can define a ROI around the person.

[0074] In some cases, the system and techniques can determine the size of the object and / or region of interest in the first (or initial) frame when identifying the region of interest (e.g., when user input is provided to identify the object or the region of interest including the object is given). In some cases, the user can provide input (e.g., zooming by providing pinch input) to define the desired size of the object or region of interest, or can maintain the size of the object in the first / initial frame. In some cases, the user can provide input that causes the device to adjust the size of the region of interest and / or the object to define a preferred size for the object in the frame sequence. When determining the region of interest (e.g., when the user selects an object), the system and techniques can crop and scale (e.g., upsample) one or more subsequent frames in the sequence (appearing after the first or initial frame) to maintain the size of the object in each subsequent frame to match the size of the object in the first frame. In some cases, the system and techniques can perform cropping and scaling operations to keep the selected object the same size as the object in the first frame, and can also keep the object at a specific location in each frame (e.g., at the center of each frame, at the position in the frame where the object is located in the first frame, or at another location). In some examples, the system and techniques can utilize object detection and tracking techniques to maintain the object's position and / or size in a sequence of frames.

[0075] In some examples, systems and techniques may apply one or more smoothing functions to an object or a bounding box (or other type of bounding region) associated with the region of interest that includes the object. One or more smoothing functions may result in progressively performed clipping and scaling to minimize frame-by-frame movement and resizing of the object within a frame sequence. The application of one or more smoothing functions can prevent the object from appearing to move in an unnatural (e.g., juddering) manner within the frame sequence due to clipping and scaling performed to maintain the object at a specific size and / or location in each frame. In some implementations, smoothing functions may address displacement (movement within a frame) and / or bounding box size changes (object size changes regardless of the center point). In some cases, displacement may be relative to a point on the object (e.g., the center point) or relative to a point within the bounding box associated with the region of interest that includes the object (e.g., the center point). In some cases, bounding box size changes may include distances relative to the object (e.g., the distance between a first part of the object and a second part of the object) or distances associated with the bounding box corresponding to the region of interest that includes the object (e.g., the diagonal distance of the bounding box).

[0076] In some examples, the system and techniques can be applied to video playback. In other examples, the system and techniques can be applied to other use cases. For instance, the system and techniques can generate video results with a constant (e.g., fixed or nearly fixed, so that the user watching the video does not perceive a change in size) target object size at a specific point in a frame of a video sequence (e.g., at the center point). Multiple video resources can be supported.

[0077] In some examples, a device can implement one or more dual-camera mode features. For instance, a dual-camera mode feature can be implemented by simultaneously using two camera lenses of the device (such as a main camera lens (e.g., a telephoto lens) and an auxiliary camera lens (e.g., a zoom lens, such as a wide-angle lens)). An example of a dual-camera mode feature is the "dual-camera video recording" feature, where two camera lenses simultaneously record two videos. These two videos can then be displayed, stored, sent to another device, and / or used in other ways. Using a dual-camera mode feature (e.g., dual-camera video recording), a device can simultaneously display two perspectives of a scene on a display (e.g., split-screen video). The advantages of a dual-camera mode feature can include: allowing the device to capture a wide field of view of a scene (e.g., with more background and surrounding objects in the scene), allowing the device to capture panoramic views of large-scale events or scenes, and other advantages.

[0078] For video captured using a single camera (or for another sequence of frames or images), various issues may arise regarding maintaining a fixed size of a target object within the frame sequence. For example, when a target object moves toward the device's camera, the device may be unable to perform a zoom-out effect due to the limited field of view of the initial video frame. In another example, when a target object moves away from the device's camera, the magnified image generated based on the initial video frame may be blurry, potentially including one or more visual artifacts, and / or may lack sharpness. Devices implementing dual-camera mode features do not incorporate any artificial intelligence technology. Such systems require end users to manually edit the images using video editing tools or software applications.

[0079] This document also describes systems and techniques for switching between lenses or cameras in devices capable of implementing one or more of the aforementioned dual-camera mode features. For example, systems and techniques can utilize camera lens switching algorithms in dual-camera systems to maintain a fixed size of a target object within a frame sequence of video from the dual-camera system. In some cases, systems and techniques can perform dual-camera zoom. In some cases, these systems and techniques can provide more detailed object zoom effects. In some examples, systems and techniques can be applied to systems or devices with two or more cameras for capturing video or other frame sequences.

[0080] Using such systems and techniques, it is possible to generate or record videos with a constant (e.g., fixed or nearly fixed, so that the user viewing the video does not perceive a change in size) target object size at a specific point (e.g., at the center point) within a frame of a video sequence. Zoom-based systems and techniques can be applied to real-time video recording, capturing still images (e.g., photographs), and / or for other use cases. In some cases, the user can select the object of interest, or the system can automatically determine the salient object (object of interest). Multi-camera system support is also provided, as described above.

[0081] The techniques described herein can be applied to any type of image capture device (e.g., mobile devices including one or more cameras, IP cameras, camera devices such as digital cameras, and / or other image capture devices). The systems and techniques can be applied to any type of content comprising sequences of frames or images, such as pre-recorded video content, live video content (e.g., non-pre-recorded video), or other content.

[0082] The various aspects of the system and technology described herein will be discussed below with reference to the accompanying drawings. Figure 1 This is a block diagram illustrating the architecture of an image capture and processing system 100. The image capture and processing system 100 includes various components for capturing and processing images of a scene (e.g., an image of scene 110). The image capture and processing system 100 can capture individual images (or photographs) and / or can capture video comprising multiple images (or video frames) in a specific sequence. A lens 115 of the system 100 faces scene 110 and receives light from scene 110. The lens 115 bends the light toward an image sensor 130. The light received by the lens 115 passes through an aperture controlled by one or more control mechanisms 120 and is received by the image sensor 130.

[0083] One or more control mechanisms 120 may control exposure, focus, and / or zoom based on information from image sensor 130 and / or image processor 150. One or more control mechanisms 120 may include multiple mechanisms and components; for example, control mechanism 120 may include one or more exposure control mechanisms 125A, one or more focus control mechanisms 125B, and / or one or more zoom control mechanisms 125C. In addition to those shown, one or more control mechanisms 120 may also include additional control mechanisms, such as those controlling analog gain, flash, HDR, depth of field, and / or other image capture attributes.

[0084] The focus control mechanism 125B of the control mechanism 120 can obtain the focus settings. In some examples, the focus control mechanism 125B stores the focus settings in a memory register. Based on the focus settings, the focus control mechanism 125B can adjust the position of the lens 115 relative to the image sensor 130. For example, based on the focus settings, the focus control mechanism 125B can move the lens 115 closer to or away from the image sensor 130 by driving a motor or servo, thereby adjusting the focus. In some cases, the system 100 may include additional lenses, such as one or more microlenses on each photodiode of the image sensor 130, which each bend light towards the corresponding photodiode before the light received from the lens 115 reaches the photodiode. The focus settings can be determined via contrast detection autofocus (CDAF), phase detection autofocus (PDAF), or some combination thereof. The focus settings can be determined using the control mechanism 120, the image sensor 130, and / or the image processor 150. The focus settings can be referred to as image capture settings and / or image processing settings.

[0085] The exposure control mechanism 125A of the control mechanism 120 can obtain the exposure settings. In some cases, the exposure control mechanism 125A stores the exposure settings in a memory register. Based on these exposure settings, the exposure control mechanism 125A can control the aperture size (e.g., aperture size or f / stop), the duration of the aperture opening (e.g., exposure time or shutter speed), the sensitivity of the image sensor 130 (e.g., ISO speed or film speed), the analog gain applied by the image sensor 130, or any combination thereof. The exposure settings may be referred to as image capture settings and / or image processing settings.

[0086] The zoom control mechanism 125C of the control mechanism 120 can obtain zoom settings. In some examples, the zoom control mechanism 125C stores the zoom settings in a memory register. Based on the zoom settings, the zoom control mechanism 125C can control the focal length of an assembly of lens elements, including lens 115 and one or more additional lenses (lens assembly). For example, the zoom control mechanism 125C can control the focal length of the lens assembly by driving one or more motors or servos to move one or more lenses relative to each other. The zoom settings may be referred to as image capture settings and / or image processing settings. In some examples, the lens assembly may include a parfocal zoom lens or a varifocal zoom lens. In some examples, the lens assembly may include a focusing lens (which in some cases may be lens 115) that first receives light from scene 110, and then, before the light reaches image sensor 130, the light passes through a focusless zoom system between the focusing lens (e.g., lens 115) and image sensor 130. In some cases, a focusless zoom system may include two positive (e.g., converging, convex) lenses with equal or similar focal lengths (e.g., within a threshold difference), and a negative (e.g., diverging, concave) lens between the two positive lenses. In some cases, the zoom control mechanism 125C moves one or more lenses in the focusless zoom system, such as the negative lens and one or both positive lenses.

[0087] Image sensor 130 includes one or more arrays of photodiodes or other photosensitive elements. Each photodiode measures the amount of light that ultimately corresponds to a specific pixel in the image generated by image sensor 130. In some cases, different photodiodes may be covered by different color filters in a color filter array, so light matching the color of the color filter covering the photodiode can be measured. Various color filter arrays can be used, including Bayer color filter arrays, four-color color filter arrays (also known as four-Bayer color filters), and / or other color filter arrays. For example, a Bayer color filter array includes a red color filter, a blue color filter, and a green color filter, wherein each pixel of the image is generated based on red light data from at least one photodiode covered in the red color filter, blue light data from at least one photodiode covered in the blue color filter, and green light data from at least one photodiode covered in the green color filter. Other types of color filter arrays may use yellow, magenta, and / or turquoise (also known as "emerald green") color filters as alternatives to or supplements to red, blue, and / or green color filters. Some image sensors may lack color filters entirely, instead using different photodiodes (in some cases stacked vertically) throughout the pixel array. These different photodiodes can have different spectral sensitivity profiles, thus responding to different wavelengths of light. Monochrome image sensors may also lack color filters, and therefore lack color depth.

[0088] In some cases, image sensor 130 may alternatively or additionally include an opaque and / or reflective mask that blocks light from reaching certain photodiodes or portions of certain photodiodes at certain times and / or from certain angles, which can be used for phase detection autofocus (PDAF). Image sensor 130 may also include an analog gain amplifier for amplifying the analog signal output from the photodiodes and / or an analog-to-digital converter (ADC) for converting the analog signal output from the photodiodes (and / or amplified by the analog gain amplifier) ​​into a digital signal. In some cases, image sensor 130 may alternatively or additionally include certain components or functions discussed with respect to one or more control mechanisms 120. Image sensor 130 may be a charge-coupled device (CCD) sensor, an electron-multiplying CCD (EMCCD) sensor, an active pixel sensor (APS), a complementary metal-oxide-semiconductor (CMOS), an N-type metal-oxide-semiconductor (NMOS), a hybrid CCD / CMOS sensor (e.g., sCMOS), or some other combination thereof.

[0089] Image processor 150 may include one or more processors, such as one or more image signal processors (ISPs) (including ISP 154), one or more host processors (including host processor 152), and / or one or more other types of processors 4710 discussed with respect to computing system 4700. Host processor 152 may be a digital signal processor (DSP) and / or other types of processors. Image processor 150 may store image frames and / or processed images in random access memory (RAM) 140 / 4720, read-only memory (ROM) 145 / 4725, cache 4712, system memory 4715, another storage device 4730, or some combination thereof.

[0090] In some implementations, the image processor 150 is a single integrated circuit or chip (e.g., referred to as a system-on-a-chip or SoC) that includes a host processor 152 and an ISP 154. In some cases, the chip may also include one or more input / output ports (e.g., input / output (I / O) port 156), a central processing unit (CPU), a graphics processing unit (GPU), a broadband modem (e.g., 3G, 4G, or LTE, 5G, etc.), memory, and connectivity components (e.g., Bluetooth). TM The I / O port 156 may include any suitable input / output port or interface according to one or more protocols or specifications, such as Inter-Integrated Circuit 2 (I2C) interface, Inter-Integrated Circuit 3 (I3C) interface, Serial Peripheral Interface (SPI) interface, Serial General Purpose Input / Output (GPIO) interface, Mobile Industrial Processor Interface (MIPI) (such as MIPI CSI-2 physical (PHY) layer port or interface), Advanced High Performance Bus (AHB) bus, any combination thereof and / or other input / output ports. In an illustrative example, the host processor 152 may use the I2C port to communicate with the image sensor 130, while the ISP 154 may use the MIPI port to communicate with the image sensor 130.

[0091] The host processor 152 of the image processor 150 can configure the image sensor 130 using parameter settings (e.g., via external control interfaces such as I2C, I3C, SPI, GPIO, and / or other interfaces). In an illustrative example, the host processor 152 can update the exposure settings used by the image sensor 130 based on the internal processing results of the exposure control algorithm from past image frames. The host processor 152 can also dynamically configure the parameter settings of the internal pipeline or modules of the ISP 154 to match the settings of one or more input image frames from the image sensor 130, so that the ISP 154 processes the image data correctly. The processing (or pipeline) blocks or modules of the ISP 154 may include modules for lens / sensor noise correction, demosaicing, color conversion, correction or enhancement / suppression of image attributes, noise reduction filters, sharpening filters, etc. For example, the processing blocks or modules of the ISP 154 can perform a variety of tasks, such as demosaicing, color space conversion, image frame downsampling, pixel interpolation, automatic exposure (AE) control, automatic gain control (AGC), CDAF, PDAF, automatic white balance, merging image frames to form HDR images, image recognition, object recognition, feature recognition, input reception, output management, memory management, or some combination thereof. The host processor 152 can configure the settings of different modules of the ISP 154.

[0092] Image processing device 105B may include various input / output (I / O) devices 160 connected to image processor 150. I / O devices 160 may include a display screen, keyboard, keypad, touchscreen, trackpad, touch-sensitive surface, printer, any other output device 4735, any other input device 4745, or some combination thereof. In some cases, text can be input into image processing device 105B via the physical keyboard or keypad of I / O device 160, or via the virtual keyboard or keypad of the touchscreen of I / O device 160. I / O 160 may include one or more ports, jacks, or other connectors enabling wired connections between system 100 and one or more peripheral devices, through which system 100 can receive data from and / or send data to one or more peripheral devices. I / O 160 may include one or more wireless transceivers enabling wireless connections between system 100 and one or more peripheral devices, through which system 100 can receive data from and / or send data to one or more peripheral devices. Peripheral devices may include any type of I / O device 160 previously discussed, and once a peripheral device is coupled to a port, jack, wireless transceiver, or other wired and / or wireless connector, the peripheral device itself may be considered an I / O device 160.

[0093] In some cases, the image capture and processing system 100 may be a single device. In other cases, the image capture and processing system 100 may be two or more separate devices, including an image capture device 105A (e.g., a camera) and an image processing device 105B (e.g., a computing device coupled to the camera). In some implementations, the image capture device 105A and the image processing device 105B may be wirelessly coupled together, for example, via one or more wires, cables, or other electrical connectors and / or via one or more wireless transceivers. In some implementations, the image capture device 105A and the image processing device 105B may be disconnected from each other.

[0094] like Figure 1 As shown, the vertical dashed line will Figure 1 The image capture and processing system 100 is divided into two parts, namely image capture device 105A and image processing device 105B. Image capture device 105A includes a lens 115, a control mechanism 120, and an image sensor 130. Image processing device 105B includes an image processor 150 (including an ISP 154 and a main processor 152), RAM 140, ROM 145, and I / O 160. In some cases, certain components shown in image capture device 105A, such as ISP 154 and / or main processor 152, may be included in image capture device 105A.

[0095] Image capture and processing system 100 may include electronic devices (e.g., mobile or landline phones, smartphones, cellular phones, etc.), Internet Protocol (IP) cameras, desktop computers, laptop computers, or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, or any other suitable electronic devices), or parts thereof. In some examples, image capture and processing system 100 may include one or more wireless transceivers for wireless communication (such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof). In some implementations, image capture device 105A and image processing device 105B may be different devices. For example, image capture device 105A may include a camera device, and image processing device 105B may include a computing device, such as a mobile phone, desktop computer, or other computing device.

[0096] Although the image capture and processing system 100 is shown as including certain components, those skilled in the art will understand that, with Figure 1Compared to the one illustrated, the image capture and processing system 100 may include more components. The components of the image capture and processing system 100 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, the components of the image capture and processing system 100 may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, GPUs, DSPs, CPUs, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors and processing systems 100 of the electronic device implementing image capture.

[0097] In some examples, an image capture and processing system 100 can be implemented as part of a system capable of performing object detection and / or tracking objects from frames of a video. One example of such a system is a video analytics system. Object detection and tracking are important components in a wide range of applications in computer vision, such as surveillance cameras, human-computer interaction, and others. Assuming the initial state (e.g., position and size) of a target object (or object of interest) in a video frame, the goal of tracking is to estimate the state of that object in subsequent frames. Object detection and tracking systems (e.g., video analytics systems) are capable of outputting patches (e.g., bounding boxes) as detection and tracking results for each frame of the video. Based on these patches, blob or object classification techniques (e.g., neural network-based classification) can be applied to determine whether the object should be classified as a certain type of object (e.g., car or person). One task of object detection, recognition, and tracking is analyzing the movement and behavior of objects in a video. An advantage of this task is that video analytics systems can access high-resolution (e.g., 1080p, 4K, or 8K) video frames, making it possible to access more details about the tracked objects.

[0098] Typically, a video analytics system obtains a sequence of video frames from a video source and can process the video sequence to perform various tasks. An example of a video source can include an IP camera or other video capture device. An IP camera is a digital video camera capable of being used for surveillance, residential security, and / or other suitable applications. Unlike analog closed-circuit television (CCTV) cameras, IP cameras can send and receive data via computer networks and the Internet. In some cases, one or more IP cameras can be located within a scene or environment and can remain static while capturing video sequences of that scene or environment.

[0099] In some cases, IP camera systems can be used for two-way communication. For example, IP cameras can use one or more network cables or a wireless network to send data (e.g., audio, video, metadata, etc.) to allow users to communicate with what they are watching. In an illustrative example, a gas station employee can use video data from an IP camera to assist a customer on how to use the payment pump (e.g., by watching the customer's actions at the pump). Commands for pan, tilt, zoom (PTZ) cameras can also be sent via a single or multiple networks. Furthermore, IP camera systems offer flexibility and wireless capabilities. For example, IP cameras can provide easy network connectivity, adjustable camera positions, and remote access to services via the internet. IP camera systems also provide distributed intelligence. For example, video analytics can be integrated into the camera itself. Encryption and authentication can also be easily provided with IP cameras. For example, IP cameras provide secure data transmission using encryption and authentication methods already defined for IP-based applications. Furthermore, IP cameras can improve labor cost efficiency. For example, video analytics can generate alerts for certain events, reducing the labor cost of monitoring all cameras in the system (based on alerts).

[0100] Video analytics offers a wide range of tasks, from instantaneous detection of events of interest to analyzing pre-recorded video for event extraction over long periods, and many others. Various surveys and real-life experiences indicate that, for example, in surveillance systems, even when monitoring images from a single camera, operators typically cannot maintain alertness and focus for more than 20 minutes. When monitoring two or more cameras, or when the time exceeds a certain period (e.g., 20 minutes), the operator's ability to monitor video and respond effectively to events decreases significantly. Video analytics can automatically analyze video sequences from cameras and send alerts about events of interest. This allows operators to monitor one or more scenes in a passive mode. Furthermore, video analytics can analyze large amounts of recorded video and extract specific video clips containing events of interest.

[0101] Video analytics also offers a variety of other features. For example, video analytics can function as an intelligent video motion detector by detecting and tracking moving objects. In some cases, video analytics can generate and display bounding boxes around valid objects. Video analytics can also be used as an intrusion detector, a video counter (e.g., by counting people, objects, vehicles, etc.), a camera tampering detector, an object leaving detector, an object / property removal detector, a property protector, a loitering detector, and / or as a slip and fall detector. Video analytics can also be used to perform various types of recognition functions, such as face detection and recognition, license plate recognition, object recognition (e.g., bags, logos, body markers, etc.) or other recognition functions. In some cases, video analytics can be trained to recognize certain objects. Another function that video analytics can perform includes providing demographic information for customer pointers (e.g., number of customers, gender, age, amount of time spent, and other appropriate metrics). Video analytics can also perform video searches (e.g., extracting basic activities for a given area) and video summaries (e.g., extracting key movements). In some cases, event detection can be performed by video analytics, including detection of fires, smoke, fights, crowds, or any other suitable event, or even by programming or learning video analytics for detection. Detectors can trigger the detection of events of interest and can send warnings or alarms to the central control room to alert users to events of interest.

[0102] In some cases, as described in more detail herein, video analytics systems can generate and detect foreground blobs that can be used to perform various operations, such as object tracking (also known as blob tracking) and / or the other operations described above. An object tracker (also known in some cases as a blob tracker) can be used to track one or more objects (or points representing objects) in a video sequence using one or more boundary regions. Boundary regions can include bounding boxes, boundary circles, boundary ellipses, or any other suitable shape representing an object and / or region of interest. See below for reference. Figures 2 to 5 Describe in detail an exemplary video analytics system with blob detection and object tracking.

[0103] Figure 2This is a block diagram illustrating an example of a video analysis system 200. The video analysis system 200 receives video frames 202 from a video source 230. Video frames 202 may also be referred to herein as a frame sequence. Each frame may also be referred to as a video picture or image. Video frames 202 may be part of one or more video sequences. The video source 230 may include image capture devices (e.g., image capture and processing system 100, camera, camera phone, video phone, or other suitable capture devices), video storage devices, video archives containing stored video, video servers or content providers providing video data, video feed interfaces receiving video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or other video content sources. In one example, the video source 230 may include an IP camera or multiple IP cameras. In the illustrative example, multiple IP cameras may be located throughout the environment and may provide video frames 202 to the video analysis system 200. For example, IP cameras may be placed in various fields of view within the environment to enable surveillance based on the captured video frames 202 of the environment.

[0104] In some embodiments, the video analytics system 200 and the video source 230 may be part of the same computing device. In some embodiments, the video analytics system 200 and the video source 230 may be part of separate computing devices. In some examples, one or more computing devices may include one or more wireless transceivers for wireless communication. One or more computing devices may include electronic devices such as cameras (e.g., IP cameras or other cameras, camera phones, video phones, or other suitable capture devices), mobile or stationary telephones (e.g., smartphones, cellular phones, etc.), desktop computers, laptop computers, or notebook computers, tablet computers, set-top boxes, televisions, display devices, digital media players, video game consoles, video streaming devices, or any other suitable electronic devices.

[0105] Video analytics system 200 includes a blob detection system 204 and an object tracking system 206. Object detection and tracking allow video analytics system 200 to provide various end-to-end features, such as the video analytics features described above. For example, intelligent motion detection and tracking, intrusion detection, and other features can directly use the results from object detection and tracking to generate end-to-end events. Based on the results of object detection and tracking, other features, such as counting and classifying people, vehicles, or other objects, can be significantly simplified. Blob detection system 204 can detect one or more blobs in video frames of a video sequence (e.g., video frame 202), and object tracking system 206 can track one or more blobs across frames of the video sequence. Object tracking system 206 can be based on any type of object tracking algorithm, such as cost-based tracking, machine learning-based tracking, and others.

[0106] As used herein, a blob refers to a foreground pixel of at least a portion (e.g., a part or the entire object) of an object in a video frame. For example, a blob may include a continuous group of pixels that constitute at least a portion of a foreground object in a video frame. In another example, a blob may refer to a continuous group of pixels that constitute at least a portion of a background object in a frame of image data. A blob may also be referred to as an object, a portion of an object, a blob of pixels, a pixel patch, a cluster of pixels, a point of pixels, a block of pixels, a clump of pixels, or any other term referring to a group of pixels of an object or a portion thereof. In some examples, a boundary region may be associated with a blob. In some examples, a tracker may also be represented by a tracker boundary region. The boundary region of a blob or tracker may include a bounding box, a boundary circle, a boundary ellipse, or any other suitable shape representing a tracker and / or a point. Although a bounding box is used herein to describe examples for illustrative purposes, other suitable shapes of boundary regions may be used to apply the techniques and systems described herein. The bounding box associated with a tracker and / or a blob may have a rectangle, a square, or other suitable shape. In the tracking layer, the terms blob and bounding box may be used interchangeably unless it is not necessary to know how to represent the blob in a formula within the bounding box.

[0107] As described in more detail below, a blob tracker can be used to track blobs. A blob tracker can be associated with a tracker bounding box and can be assigned a tracker identifier (ID). In some examples, the bounding box for a blob tracker in the current frame can be the bounding box of a previous blob in the previous frame associated with that blob tracker. For example, when a blob tracker (after being associated with a previous blob in the previous frame) is updated in the previous frame, the update information for the blob tracker can include tracking information for the previous frame as well as a prediction of the blob tracker's position in the next frame (in this case, the current frame). The prediction of the blob tracker's position in the current frame can be based on the blob's position in the previous frame. A history or motion model can be maintained for the blob tracker, including a history of various states, velocities, and positions for the blob tracker across consecutive frames, as described in more detail below.

[0108] In some examples, the motion model used for the blob tracker can determine and maintain two positions for the blob tracker for each frame. For example, the first position of the blob tracker for the current frame can include the predicted position in the current frame. This first position is referred to herein as the predicted position. The predicted position of the blob tracker in the current frame includes the position of the blob associated with the blob tracker in the previous frame. Therefore, the position of the blob associated with the blob tracker in the previous frame can be used as the predicted position of the blob tracker in the current frame. The second position of the blob tracker for the current frame can include the position of the blob associated with the tracker in the current frame. This second position is referred to herein as the actual position. Accordingly, the position of the blob associated with the blob tracker in the current frame is used as the actual position of the blob tracker in the current frame. The actual position of the blob tracker in the current frame can be used as the predicted position of the blob tracker in the next frame. The position of the blob can include the position of the blob's bounding box.

[0109] The velocity of a blob tracker can include the displacement of the blob tracker between consecutive frames. For example, the displacement between the centers (or centroids) of the two bounding boxes of a blob tracker in two consecutive frames can be determined. In an illustrative example, the velocity of a blob tracker can be defined as... , in .item ( , ) represents the center position of the tracker's bounding box in the current frame, where It is the x-coordinate of the bounding box, and This is the y-coordinate of the bounding box. (Item) ( , The bounding box (x and y) represents the center position of the tracker's bounding box in the previous frame. In some implementations, it's possible to use four parameters simultaneously to estimate x, y, width, and height. In some cases, because the timing for video frame data is constant or at least doesn't have significantly different timeouts (depending on the frame rate, e.g., 30 frames per second, 60 frames per second, 120 frames per second, or other suitable frame rates), time variables may not be needed in velocity calculations. In some cases, a time constant (depending on the real-time frame rate) and / or a timestamp can be used.

[0110] Using a blob detection system 204 and an object tracking system 206, the video analysis system 200 can perform blob generation and detection for each frame or image of a video sequence. For example, the blob detection system 204 can perform background subtraction for a frame and then detect foreground pixels in the frame. Foreground blobs are generated from the foreground pixels using morphological operations and spatial analysis. Additionally, the blob tracker from the previous frame needs to be associated with the foreground blobs in the current frame and also needs to be updated. Both the data association between the tracker and the blob, and the tracker update, can rely on cost function calculations. For example, when a blob is detected from the current input video frame, the blob tracker from the previous frame can be associated with the detected blob based on cost calculations. The tracker is then updated based on the data association, including updating the tracker's state and position, enabling object tracking in the current frame. (Reference) Figure 4 and Figure 5 Further details are described regarding the blob detection system 204 and the object tracking system 206.

[0111] Figure 3 This is an example of a video analytics system (e.g., video analytics system 200) that processes video frames at time t. Figure 3 As shown, video frame A 302A is received by blob detection system 304A. Blob detection system 304A generates a foreground blob 308A for the current frame A 302A. After performing blob detection, foreground blob 308A can be used for temporal tracking by object tracking system 306A. Object tracking system 306A can calculate the cost between the blob tracker and the blob (e.g., the cost of distance, weighted distance, or other costs). Object tracking system 306A can use the calculated cost (e.g., using a cost matrix or other suitable association techniques) to associate the execution data to associate or match the blob tracker (e.g., a blob tracker generated or updated based on the previous frame or a newly generated blob tracker) with blob 308A. The blob tracker can be updated based on the data association, including aspects of the tracker's position, to generate an updated blob tracker 310A. For example, the state and position of the blob tracker for video frame A 302A can be calculated and updated. The position of the blob tracker in the next video frame N 302N can also be predicted from the current video frame A 302A. For example, the predicted position of the blob tracker for the next video frame N 302N can include the position of the blob tracker (and its associated points) in the current video frame A 302A. Once the updated blob tracker 310A is generated, tracking of the blobs in the current frame A 302A can be performed.

[0112] Upon receiving the next video frame N 302N, the speckle detection system 304N generates a foreground speckle 308N for frame N 302N. The object tracking system 306N can then perform temporal tracking of the speckle 308N. For example, the object tracking system 306N obtains a speckle tracker 310A updated based on the previous video frame A 302A. The object tracking system 306N can then calculate a cost and use the newly calculated cost to associate the speckle tracker 310A with the speckle 308N. The speckle tracker 310A can be updated based on the data association to generate the updated speckle tracker 310N.

[0113] Figure 4 This is a block diagram illustrating an example of a blob detection system 204. Blob detection is used to segment moving objects from the global background in a scene. The blob detection system 204 includes a background subtraction engine 412 that receives video frames 402. The background subtraction engine 412 can perform background subtraction to detect foreground pixels in one or more video frames 402. For example, background subtraction can be used to segment moving objects from the global background in a video sequence and generate a foreground-background binary mask (referred to herein as a foreground mask). In some examples, background subtraction can be performed between the current frame or image and a background model that includes a background portion of the scene (e.g., a static or predominantly static portion of the scene). Based on the result of background subtraction, a morphological engine 414 and a connected component analysis engine 416 can perform foreground pixel processing to cluster the foreground pixels into foreground blobs for tracking purposes. For example, after subtracting the background, morphological operations can be applied to remove noisy pixels and smooth the foreground mask. Connected component analysis can then be applied to generate points. Then, blob processing can be performed, which may include further filtering out some blobs and merging some blobs together to provide a bounding box as input for tracking.

[0114] Background subtraction engine 412 can use any suitable background subtraction technique (also known as background extraction) to model the background of a scene (e.g., captured in a video sequence). An example of a background subtraction method used by background subtraction engine 412 includes modeling the background of the scene as a statistical model based on relatively static pixels in previous frames that are not considered to belong to any moving area. For example, background subtraction engine 412 can use a Gaussian distribution model for each pixel location and have parameters of mean and variance to model each pixel location in the frames of the video sequence. All values ​​of previous pixels at a particular pixel location are used to calculate the mean and variance of a target Gaussian model for the pixel location. When processing a pixel at a given location in a new video frame, the pixel value is evaluated by the current Gaussian distribution for that pixel location. The pixel can be classified as a foreground pixel or a background pixel by comparing the difference between the pixel value and the mean of the specified Gaussian model. In an illustrative example, if the distance between the pixel value and the Gaussian mean is less than three (3) times the variance, the pixel is classified as a background pixel. Otherwise, in this illustrative example, the pixel is classified as a foreground pixel. At the same time, the Gaussian model for the pixel location will be updated by taking the current pixel value into account.

[0115] Background subtraction engine 412 can also perform background subtraction using Gaussian mixtures (also known as Gaussian mixture models (GMMs)). A GMM models each pixel as a mixture of Gaussian models and uses an online learning algorithm to update the model. Each Gaussian model is represented by a mean, standard deviation (or a covariance matrix if the pixel has multiple channels), and weights. The weights represent the probability that the Gaussian event occurred in the past history.

[0116]

[0117] Equation (1)

[0118] Equation (1) shows the equation for the GMM model, where there are K Gaussian models. Each Gaussian model has a distribution with mean µ and variance ∑, and weights ω. Here, i is the index of the Gaussian model, and t is the time instance. As the equation shows, the parameters of the GMM change over time after processing one frame (at time t). In GMM or any other learning-based background subtraction, the current pixel influences the entire model of the pixel location based on a learning rate, which can be constant or generally at least the same for each pixel location. GMM-based background subtraction (or other learning-based background subtraction) is suited for local changes for each pixel. Therefore, once the moving object stops, for each pixel location of that object, the same pixel value continues to contribute significantly to its associated background model, and the region associated with that object becomes the background.

[0119] The background subtraction techniques mentioned above are based on the assumption that the camera is mounted statically, and that a new background model will need to be calculated if the camera moves or its orientation changes at any time. There are also background subtraction methods that can handle foreground subtraction based on a moving background, including techniques such as keypoint tracking, optical flow, saliency, and other motion estimation-based methods.

[0120] Background subtraction engine 412 can generate a foreground mask with foreground pixels based on the result of background subtraction. For example, the foreground mask can include a binary image containing pixels that constitute a foreground object (e.g., a moving object) in the scene and pixels of the background. In some examples, the background (background pixels) of the foreground mask can be a solid color, such as a pure white background, a pure black background, or other solid colors. In these examples, the foreground pixels of the foreground mask can be a different color than the color used for the background pixels, such as pure black, pure white, or other solid colors. In one illustrative example, the background pixel can be black (e.g., a pixel color value of 0 in 8-bit grayscale or other suitable value), while the foreground pixel can be white (e.g., a pixel color value of 255 in 8-bit grayscale or other suitable value). In another illustrative example, the background pixel can be white, while the foreground pixel can be black.

[0121] Using a foreground mask generated from background subtraction, the morphological engine 414 can execute morphological functions to filter foreground pixels. Morphological functions can include erosion and dilation functions. In one example, an erosion function can be applied, followed by a series of one or more dilation functions. An erosion function can be applied to remove pixels on object boundaries. For example, the morphological engine 414 can apply an erosion function (e.g., FilterErode3×3) to a 3×3 filter window of the currently processed center pixel. A 3×3 window can be applied to each foreground pixel (as the center pixel) in the foreground mask. Those skilled in the art will understand that other window sizes besides 3×3 can be used. The erosion function can include an erosion operation that sets the current foreground pixel as a background pixel if one or more of the neighboring pixels of the current foreground pixel (used as the center pixel) in the foreground mask within the 3×3 window are background pixels. Such an erosion operation can be referred to as a strong erosion operation or a single-neighbor erosion operation. Here, the neighboring pixels of the current center pixel include eight pixels in the 3×3 window, and the ninth pixel is the current center pixel.

[0122] Dilation operations can be used to enhance the boundaries of foreground objects. For example, morphological engine 414 can apply a dilation function (e.g., filter dilation FilterDialte3×3) to a 3×3 filter window for the center pixel. A 3×3 dilation window can be applied to each background pixel (as the center pixel) in the foreground mask. Those skilled in the art will understand that other window sizes besides 3×3 can be used. The dilation function can include a dilation operation that sets the current background pixel as the foreground pixel if one or more of the neighboring pixels of the current background pixel (used as the center pixel) in the foreground mask in the 3×3 window are foreground pixels. The neighboring pixels of the current center pixel include eight pixels in the 3×3 window, and the ninth pixel is the current center pixel. In some examples, multiple dilation functions can be applied after an erosion function. In one illustrative example, three function calls for dilation of a 3×3 window size can be applied to the foreground mask before sending three function calls for dilation of the 3×3 window size to the connected component analysis engine 416. In some examples, an erosion function can be applied first to remove noisy pixels, and then a series of dilation functions can be applied to refine the foreground pixels. In an illustrative example, a erosion function with a 3×3 window size is first called, and these three function calls are applied to the foreground mask before being sent to the connection component analysis engine 416 with three function calls to inflate the 3×3 window size. Details regarding content-adaptive morphological operations are described below.

[0123] After performing morphological operations, the connectivity component analysis engine 416 can apply connectivity component analysis to connect adjacent foreground pixels, so that connectivity components and blobs can be expressed as formulas. In some implementations of connectivity component analysis, a set of bounding boxes is returned such that each bounding box contains one component of the connected pixels. An example of connectivity component analysis performed by the connectivity component analysis engine 416 is implemented as follows:

[0124] For each pixel of the foreground masking {

[0125] - If it is a foreground pixel and has not yet been processed, apply the following steps:

[0126] - Apply the FloodFill function to connect this pixel to other foreground components and generate connection components.

[0127] - Insert the connection component into the list of connection components.

[0128] - Mark pixels in the connected components as being processed.

[0129] The Floodfill function is an algorithm for determining regions connected to seed nodes in a multidimensional array (e.g., a 2-D image in this case). The Floodfill function first obtains the color or intensity value at the seed location (e.g., a foreground pixel) masked by the source foreground, and then finds all neighboring pixels with the same (or similar) values ​​based on 4 or 8 connectivity. For example, in the case of 4 connectivity, the neighbors of the current pixel are defined as those pixels with coordinates (x + d, y) or (x, y + d), where d equals 1 or -1, and (x, y) is the current pixel. Those skilled in the art will understand that other numbers of connections can be used. Some objects are divided into different connectivity components, and some objects are divided into the same connectivity component (e.g., neighboring pixels with the same or similar values). Additional processing can be applied to further process the connectivity components for grouping. Finally, based on the connectivity components, blobs 408 including neighboring foreground pixels are generated. In one example, a blob may consist of a single connectivity component. In another example, a blob may include multiple connectivity components (e.g., when two or more blobs are merged together).

[0130] The blob processing engine 418 can perform additional processing to further process the blobs generated by the connectivity component analysis engine 416. In some examples, the blob processing engine 418 can generate bounding boxes to represent detected blobs and blob trackers. In some cases, point bounding boxes can be output from the blob detection system 204. In some examples, there may be filtering processing for the connectivity components (bounding boxes). For example, the blob processing engine 418 can perform content-based filtering on certain blobs. In some cases, machine learning methods can determine that the current blob contains noise (e.g., leaves in the scene). Using machine learning information, the blob processing engine 418 can determine that the current blob is a noisy blob and can remove it from the result points provided to the object tracking system 206. In some cases, the blob processing engine 418 can filter out one or more small blobs below a certain size threshold (e.g., the area of ​​the bounding box surrounding a blob is below an area threshold). In some examples, there may be a merging process to merge some connectivity components (represented as bounding boxes) into a larger bounding box. For example, the blob processing engine 418 can merge nearby blobs into one large blob to eliminate the risk that too many small blobs may belong to one object. In some cases, even if the foreground pixels of two bounding boxes are completely disconnected, two or more bounding boxes can be merged together based on certain rules. In some embodiments, the blob detection system 204 does not include a blob processing engine 418, or in some cases, the blob processing engine 418 is not used. For example, blobs generated by the connectivity component analysis engine 416 can be input into the object tracking system 206 without further processing to perform blob and / or object tracking.

[0131] In some implementations, density-based speckle area trimming can be performed by the speckle processing engine 418. For example, density-based speckle area trimming can be applied after post-filtering and when all specks have been formulaically represented before being input into the tracking layer. A similar process is applied vertically and horizontally. For example, density-based speckle area trimming can be performed first vertically and then horizontally, or vice versa. The purpose of density-based speckle area trimming is to filter out columns (in the vertical process) and / or rows (in the horizontal process) of the bounding box if a column or row contains only a small number of foreground pixels.

[0132] The vertical process involves calculating the number of foreground pixels in each column of the bounding box and representing this number as a column density. Then, starting with the leftmost column, columns are processed one by one. The column density of each current column (the column currently being processed) is compared to the maximum column density (the column density of all columns). If the column density of the current column is less than a threshold (e.g., a percentage of the maximum column density, such as 10%, 20%, 30%, 50%, or other suitable percentage), that column is removed from the bounding box, and the next column is processed. However, once the current column has a column density that is not less than the threshold, the processing ends, and no further columns are processed. Similar processing can then be applied from the rightmost column. Those skilled in the art will understand that the vertical process can handle columns that begin with a different column than the leftmost column, such as the rightmost column or other suitable columns within the bounding box.

[0133] The horizontal density-based blob area trimming process is similar to the vertical process, except that rows instead of columns of the bounding box are processed. For example, the number of foreground pixels in each row of the bounding box is calculated and represented as a row density. These rows are then processed row by row, starting from the top row. For each current row (the row currently being processed), the row density is compared to the maximum row density (the row density of all rows). If the current row's row density is less than a threshold (e.g., a percentage of the maximum row density, such as 10%, 20%, 30%, 50%, or other suitable percentage), that row is removed from the bounding box, and the next row is processed. However, once the current row has a row density that is not less than the threshold, the processing ends, and no further rows are processed. A similar process can then be applied from the bottom row. Those skilled in the art will understand that horizontal processing can handle rows that begin with a different row than the top row, such as the bottom row or other suitable rows within the bounding box.

[0134] One purpose of density-based spot area trimming is to remove shadows. For example, when a person and their long, thin shadow are detected within a spot (bounding box), density-based spot area trimming can be applied. Because the column density in the shadow area is relatively small, such shadow areas can be removed after applying density-based spot area trimming. Unlike morphology, which alters the thickness of the spot (besides filtering out some isolated foreground pixels from the formulaically represented spot) and roughly preserves the shape of the bounding box, density-based spot area trimming can drastically change the shape of the bounding box.

[0135] Once a blob is detected and processed, object tracking (also known as blob tracking) can be performed to trace the detected blob. In some examples, cost-based techniques can be used to perform the tracking, such as those mentioned above. Figure 5As described. In some examples, one or more machine learning systems (e.g., one or more neural network-based systems) may be used to perform tracing, as further described below.

[0136] Figure 5 This is a block diagram illustrating an example of an object tracking system 206. The input to blob / object tracking is a list of blobs 508 (e.g., blob bounding boxes) generated by the blob detection system 204. In some cases, a unique ID is assigned to the tracker, and the history of the bounding boxes is maintained. Object tracking in video sequences can be used in many applications, including surveillance applications and others. For example, in many security applications, the ability to detect and track multiple objects in the same scene has generated considerable interest. When a blob (constituting at least a portion of an object) is detected from an input video frame, it is necessary, based on cost calculations, to associate the blob tracker from the previous video frame with the blob in the input video frame. The blob tracker can be updated based on the associated foreground blob. In some cases, the steps in object tracking can be implemented in a series of ways.

[0137] The cost determination engine 512 of the object tracking system 206 can obtain the blob 508 of the current video frame from the blob detection system 204. The cost determination engine 512 can also obtain the blob tracker 510A updated from a previous video frame (e.g., video frame A 302A). The cost between the blob tracker 510A and the blob 508 can then be calculated using a cost function. Any suitable cost function can be used to calculate the cost. In some examples, the cost determination engine 512 can measure the cost between the blob tracker and the blob by calculating the Euclidean distance between the centroid of the tracker (e.g., the bounding box for the tracker) and the centroid of the bounding box of the foreground blob. In an illustrative example using a 2-D video sequence, this cost function is calculated as follows:

[0138]

[0139] item and These are the center locations of the blob tracker and the blob bounding box, respectively. As described herein, in some examples, the bounding box of the blob tracker can be the bounding box of a blob associated with a blob tracker in a previous frame. In some examples, other cost function methods can be implemented, using the minimum distance in the x or y direction to calculate the cost. Such techniques may be good for certain controlled situations, such as well-aligned lane transport. In some examples, the cost function can be based on the distance between the blob tracker and the blob, where instead of using the center location of the blob's bounding box and the tracker to calculate the distance, the boundaries of the bounding boxes are considered, such that a negative distance is introduced when the two bounding boxes geometrically overlap. Furthermore, the value of this distance is further adjusted according to the size ratio of the two associated bounding boxes. For example, the cost can be weighted based on the ratio of the area of ​​the blob tracker's bounding box to the area of ​​the blob's bounding box (e.g., by multiplying the determined distance by this ratio).

[0140] In some embodiments, a cost is determined for each tracker-spot pair between each tracker and each spot. For example, if there are three trackers (including tracker A, tracker B, and tracker C) and three spots (including spot A, spot B, and spot C), a separate cost can be determined between tracker A and each of spots A, B, and C, as well as a separate cost between tracker B and each of spots A, B, and C. In some examples, the costs can be arranged as a cost matrix, which can be used for data association. For example, the cost matrix can be a two-dimensional matrix, where one dimension is spot tracker 510A and the second dimension is spot 508. Each tracker-spot pair or combination between tracker 510A and spot 508 includes the cost included in the cost matrix. The best match between tracker 510A and spot 508 can be determined by identifying the tracker-spot pair with the lowest cost in the matrix. For example, the lowest cost between tracker A and spots A, B, and C is used to determine the spot associated with tracker A.

[0141] The data association between tracker 510A and spot 508, as well as the update of tracker 510A, can be based on a determined cost. Data association engine 514 matches trackers (or tracker bounding boxes) with corresponding spots (or spot bounding boxes) or assigns corresponding spots (or spot bounding boxes) to trackers (or tracker bounding boxes) and vice versa. For example, as previously described, data association engine 514 can associate spot tracker 510A with spot 508 using the tracker-spot pair with the lowest cost. Another technique for associating spot trackers with spots includes the Hungarian method, a combinatorial optimization algorithm that can solve such assignment problems in polynomial time and can anticipate subsequent primal-dual methods. For example, the Hungarian method can optimize the global cost across all spot trackers 510A and spot 508 to minimize the global cost. The spot tracker-spot combination that minimizes the global cost in the cost matrix can be determined and used as the association.

[0142] Besides the Hungarian method, other robust approaches can be used to perform data association between blobs and blobs trackers. For example, additional constraints can be used to address the association problem to make the solution more robust to noise while matching as many trackers and blobs as possible. Regardless of the association technique used, the data association engine 514 can rely on the distance between blobs and trackers.

[0143] Once the association between the speckle tracker 510A and speckle 508 is established, the speckle tracker update engine 516 can use the information from the associated speckles, as well as the tracker's temporal state, to update the position status (or multiple states) of tracker 510A for the current frame. After updating tracker 510A, the speckle tracker update engine 516 can use the updated tracker 510N to perform object tracking and can also provide the updated tracker 510N for processing the next frame.

[0144] The status or condition of a blob tracker can include the identified (or actual) location of the tracker in the current frame and its predicted location in the next frame. The location of the foreground blob is identified by the blob detection system 204. However, as described in more detail below, it may be necessary to predict the location of the blob tracker in the current frame based on information from the previous frame (e.g., using the location of a blob associated with a blob tracker in the previous frame). After performing data association for the current frame, the tracker location in the current frame can be identified as the location of one or more of its associated blobs in the current frame. The tracker location can be further used to update the tracker's motion model and predict its location in the next frame. In addition, in some cases, trackers may be temporarily lost (e.g., when a blob being tracked by the tracker is no longer detected), in which case it is also necessary to predict the location of these trackers (e.g., using a Kalman filter). These trackers are not temporarily displayed to the system. Predicting the bounding box location not only helps to maintain a certain level of tracking for lost and / or merged bounding boxes, but also allows for a more accurate estimation of the tracker's initial location, thus making the association between the bounding box and the tracker more precise.

[0145] As described above, the position of a blob tracker in the current frame can be predicted based on information from the previous frame. One method for performing tracker position updates is to use a Kalman filter. A Kalman filter is a framework that includes two operations. The first operation is to predict the tracker's state, and the second operation is to use measurements to correct or update the state. In this case, the tracker predicts (using the blob tracker update engine 516) its position in the current frame from the previous frame, and when the current frame is received, the tracker first corrects its position state using measurements of (one or more) blobs (e.g., (one or more) blob bounding boxes), and then predicts its position in the next frame. For example, a blob tracker can employ a Kalman filter to measure its trajectory and predict its future position(s). The Kalman filter relies on the associated measurements of (one or more) blobs(s) to correct the motion model for the blob tracker and predict the object tracker's position in the next frame. In some examples, if the blob tracker is associated with a blob in the current frame, the position of that blob is directly used to correct the motion model of the blob tracker in the Kalman filter. In some examples, if the blob tracker is not associated with any blob in the current frame, the position of the blob tracker in the current frame is identified as its predicted position from the previous frame. This means that the motion model for the blob tracker is not corrected and the prediction is propagated with the previous model of the blob tracker (from the previous frame).

[0146] Unlike the location of the tracker, the condition or state of the tracker can also, or alternatively, include the temporal condition or state of the tracker. The temporal state of a tracker can include: whether the tracker is a new tracker that did not exist before the current frame; the normal state of a tracker that has survived for a certain duration and is to be output as an identified tracker-blob pair to the video analysis system; the lost state of a tracker that is not associated with or does not match any foreground blob in the current frame; the dead state of a tracker that is not associated with any blob in a certain number of consecutive frames (e.g., two or more frames, a threshold duration, etc.); and / or other suitable temporal states. Another temporal state that can be maintained for a blob tracker is the duration of the tracker. The duration of a blob tracker includes the number of frames (or other time measurements, such as time) in which the tracker has been associated with one or more blobs.

[0147] Additional state or condition information may be needed to update the tracker, which may require a state machine for object tracking. Given information about the associated (one or more) spots and the tracker's own state history table, the condition also needs to be updated. The state machine gathers all the necessary information and updates the condition accordingly. Various tracker conditions can be updated. For example, in addition to the tracker's life status (e.g., new, lost, dead, or other suitable life status), the tracker's association confidence and relationships with other trackers can also be updated. As an example of tracker relationships, when two objects (e.g., a person, a vehicle, or other object of interest) intersect, the two trackers associated with these two objects will merge for certain frames, and the merging or occlusion status needs to be recorded for advanced video analytics.

[0148] Regardless of the tracking method used, a new tracker begins by associating with a blob in a frame and moves forward, potentially connecting with blobs that may move across multiple frames. When a tracker has been continuously associated with a blob and a certain duration (threshold duration) has elapsed, the tracker can be promoted to a normal tracker. For example, the threshold duration is the duration for which a new blob tracker must be continuously associated with one or more blobs before transitioning to a normal tracker (switching to normal state). The normal tracker is output as an identified tracker-blob pair. For example, when a tracker is promoted to a normal tracker, the tracker-blob pair is output as an event at the system level (e.g., presented on a display as a tracked object, as an alarm output, and / or other suitable event). In some implementations, the normal tracker (e.g., including some condition data of the normal tracker, a motion model for the normal tracker, or other information related to the normal tracker) can be output as part of the object metadata. Metadata including the normal tracker can be output from a video analytics system (e.g., an IP camera running the video analytics system) to a server or other system storage device. Then, metadata can be analyzed for event detection (e.g., via a rule interpreter). Trackers that have not been promoted to normal trackers can be removed (or killed), after which they can be considered dead.

[0149] As described above, in some implementations, one or more machine learning systems (e.g., one or more neural networks) can be used to perform blob or object tracking. In some cases, using machine learning systems for blob / object tracking can allow for online operability and speed.

[0150] Figure 6A This is a diagram illustrating an example of a machine learning-based object detection and tracking system 600 that includes a fully convolutional deep neural network. System 600 can perform object detection, object tracking, and object segmentation. Figure 6A As shown, the input to the object detection and tracking system 600 includes one or more reference object images (referred to as "examples"), and Figure 6A The middle is shown as An image refers to a 255×255 image with three color channels such as red, green, and blue, and one or more query image frames (called "search patches"). Figure 6A The middle is shown as (Image). For example, an example and multiple search patches from that example can be input into System 600 to detect, track, and segment one or more objects in that example.

[0151] The object detection and tracking system 600 includes a ResNet-50 neural network (up to the last convolutional layer of the fourth stage) as the backbone of the neural network of the system 600. To achieve higher spatial resolution in deeper layers, the output path reduces the output stride to 8 by using convolutions with a stride of 1. The receptive field is increased by using dilated convolutions. For example, in the 3×3 convolutional layer of conv4_1 (… Figure 6A In the top conv4 layer (of the array), the stride can be set to 1, and the dilation rate can be set to 2. For example... Figure 6A As shown, the feature map size of the top conv4_1 layer is The feature map size of the bottom conv4_2 layer is Unlike the initial ResNet-50 architecture, there is no downsampling in the conv4_1 or conv4_2 layers.

[0152] Place one or more adjustment layers (in) Figure 6A An adjustment layer (marked as "adjustment") is added to the backbone. In some cases, each adjustment layer may include a 1×1 convolutional layer with 256 output channels. Two adjustment layers can perform depth-wise cross-correlation to generate a feature map of a specific size. Figure 6A The size is shown in the middle. For example, the output features of the adjustment layer are cross-correlated in terms of depth, forming a feature map of size 17×17 (with 256 channels). The purpose of the adjustment layer is to extract features from lower-level networks (e.g., in image size). (In the context of system 600) Locating the target object. For example, the adjustment layer can be used to extract feature maps from a reference object image (example) and a query image frame (search patch). The RoW in the last layer of the second row of system 600 represents the response of the candidate window, which is the target object region from the query image frame input to system 600. The example and search patch share the network parameters from conv_1 to conv4_x, but do not share the parameters of the adjustment layer.

[0153] An optimization module U-shaped structure can be used, which combines the feature maps of the backbone with upsampling to achieve better results. For example, the layers in the top row of System 600 perform deconvolution, followed by upsampling (shown as upsampling components U2, U3, and U4), the purpose of which is to restore the target object location to a higher level (e.g., to image size). An example of a U3 component is... Figure 6B As shown in the diagram. Components U2 and U4 have similar structures and operations to component U3. The last convolutional layer before the sigmoid operation (labeled "...") ,1) is used to change the dimension of the feature map from The size is reduced to 127*127*1. The sigmoid function is used to binarize the output of the object mask, which is the result of object segmentation. The object mask can include a binary mask with a value of 0 or 1 for each pixel. The purpose of generating the object mask is to have an accurate object bounding box. The bounding box can be a rectangle in any orientation. In some cases, the object bounding box is close to the center point or centroid of the object (e.g., centered about the center point or centroid of the object). In some cases, a scoring branch can be included in System 600 to generate a scoring matrix based on the object mask. In these cases, the scoring matrix can be used for accurate object localization. As described above, the first four stages of the ResNet-50 network share parameters, and the output is connected to a 1×1 convolution with shared parameters to adjust the channels. To conduct in-depth cross-correlation. Regarding... Figure 6A Other details of the backbone architecture are... Figure 6C As shown in the image.

[0154] In some implementations, the classification system can be used to classify objects that have been detected and tracked in one or more video frames of a video sequence. Different types of object classification applications can be used. In the first example classification application, a lower-resolution input image is used to provide the category and confidence score for the entire input image. In such an application, classification is performed on the entire image. In the second example classification system, a relatively high-resolution input image is used, and the output consists of multiple objects within the image, each with its own bounding box (or ROI) and classified object type. The first example classification application is referred to herein as "image-based classification," while the second example classification application is referred to herein as "blob-based classification." Both applications can achieve high classification accuracy when utilizing neural network-based solutions (e.g., deep learning).

[0155] Figure 7 Figure 700 illustrates an example of a machine learning-based classification system. As shown, machine learning-based classification (also known as region-based classification) first extracts region proposals (e.g., blobs) from an image. The extracted region proposals (which may include blobs) are fed into a deep learning network for classification. Deep learning classification networks typically begin with an input layer (image or blob), followed by a sequence of convolutional and pooling layers (and other layers), and then end with a fully connected layer. Following the convolutional layers may be a layer of rectified linear unit (ReLU) activation function. The convolutional, pooling, and ReLU layers serve as learnable feature extractors, while the fully connected layer acts as a classifier.

[0156] In some cases, when blobs are fed into a deep learning classification network, one or more shallow layers in the network can learn simple geometric objects, such as lines and / or other objects, that represent the objects to be classified. Deeper layers will learn more abstract, detailed features about the objects, such as sets of lines or other detailed features that define the shape, and finally, a set of shapes from the earlier layers that constitutes the shape of the object being classified (e.g., a person, car, animal, or any other object). See below for reference. Figures 42 to 46C To describe more details about the structure and function of neural networks.

[0157] Since point-based classification requires less computational complexity and less memory bandwidth (e.g., the memory needed to maintain the network structure), it can be used directly.

[0158] Various deep learning-based detectors can be used to classify or detect objects in video frames. For example, a detector based on the Cifar-10 network can be used to perform blob-based classification to categorize blobs. In some cases, a Cifar-10 detector can be trained to classify only people and vehicles. A Cifar-10 network-based detector can take blobs as input and classify them into one of many predefined categories with confidence scores. See below for reference. Figure 21 More details describing the Cifar-10 detector.

[0159] Another deep learning-based detector is the Single-Shot Detector (SSD), a fast single-shot object detector capable of being applied to multiple object categories. A key feature of the SSD model is the use of multi-scale convolutional bounding box outputs, which are connected to multiple feature maps at the top of the neural network. This representation allows SSD to efficiently model various box shapes. It has been demonstrated that, given the same VGG-16 infrastructure, SSD outperforms state-of-the-art object detectors in both accuracy and speed. The SSD deep learning detector is described in more detail in K. Simonyan and A. Zisserman's "Very deep convolutional networks for large-scale image recognition," CoRR, abs / 1409.1556, 2014, the entire contents of which are incorporated herein by reference for all purposes. See below for references. Figure 25 A to Figure 25 C describes more details about the SSD detector.

[0160] Another example of a deep learning-based detector capable of detecting or classifying objects in video frames is the "You Only Look Once" (YOLO) detector. Running on a Titan X, the YOLO detector processes images at 40-90 frames per second (fps) with an mAP of 78.6% (based on VOC 2007). An SSD300 model runs at 59 fps on an Nvidia Titan X and typically performs faster than the current YOLO 1, which has recently been superseded by its successor, YOLO 2. The YOLO deep learning detector is described in more detail in J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You Only Look Once: Unified, Real-Time Object Detection," arXiv preprint arXiv:1506.02640, 2015, the entire text of which is incorporated herein by reference for all purposes. References below. Figures 46A to 46C Further details about the YOLO detector are described. Although the SSD and YOLO detectors are described to provide illustrative examples of deep learning-based object detectors, those skilled in the art will understand that any other suitable neural network can be used to perform object classification.

[0161] As described above, in many cases, when a region of interest and / or an object moves relative to one or more cameras capturing a frame sequence, it may be desirable to maintain the size of the region of interest and / or the object of interest frame by frame in the frame sequence. Examples of such scenarios can include when a user provides input to a device to capture video of an event including the object of interest. For example, the device might record video of a person performing a dance routine, where the person moves relative to the camera (in both the depth and lateral directions) while the video is being captured. When the person is moving relative to the camera, the user may expect the person to maintain a constant size throughout the video (and, in some cases, a consistent position within the captured frames). Another example of this is in video analytics when an IP camera is capturing video of a scene. For example, an IP camera might capture video of a user's living room, where it may be desirable to maintain the size of one or more people in the room (and, in some cases, a consistent position within the captured frames), even as one or more people move away from the camera (in the depth direction).

[0162] When a device is capturing a frame sequence of an object (e.g., video of a person performing a dance routine), that object can move relative to one or more cameras capturing the frame sequence. As a result, the device may struggle to maintain the desired object size (e.g., the object's size in the initial frame when video capture first begins) while the object is moving during frame sequence capture. For example, the user may have adjusted the camera zoom to give the object the desired size in the frame. However, the object's size ratio (the size of the object relative to the frame, called the object size to frame ratio) will change dynamically as the object moves. Manually changing the object size to frame ratio during video capture can be cumbersome for the user. Automatically tracking the object during video recording can also be difficult.

[0163] As described above, this paper describes systems and techniques for maintaining a fixed size of a target object within a frame sequence. In an illustrative example, an initial frame or other frame sequence of video can be captured and displayed. In some cases, the user can provide user input in the initial frame to indicate the object of interest (e.g., by drawing a bounding box around the object, selecting the object, zooming in on the object, etc.). In some cases, the object can be detected automatically without user input. In some cases, the size of the object in the initial frame can be determined and used as a reference size for the object in subsequent video frames after the initial frame. In some cases, a bounding box can be set for the object in the initial frame. In some examples, the center point coordinates (or other points associated with the bounding box or object) and the diagonal length of the bounding box (or other lengths associated with the bounding box or object) can be determined and used as a reference for subsequent video frames.

[0164] Object detection and tracking can be initialized and performed to detect and track objects in subsequent video frames. For each subsequent video frame, the coordinates of the object's bounding box center point (or other points associated with the bounding box or object) and the length of the bounding box's diagonal (or other length associated with the bounding box or object) can be determined or recorded. Once the set of bounding box center point (or other point) coordinates and diagonal lengths (or other lengths) is obtained for each frame of the video, a smoothing function can be applied to smooth the amount of variation in the bounding box diagonal length (and therefore size) in each frame of the video. In some cases, a smoothing function can also be applied to smooth the movement trajectory of the bounding box center point in the video frame. As described in this paper, a scaling factor can be calculated for each frame by comparing the length of the bounding box diagonal in the initial video frame (called the reference frame) with the current frame being processed. The scaling factor can be used to scale or resize each frame. Each video frame can be cropped and scaled based on the center point coordinates and the scaling factor. In some cases, video stabilization can be applied after cropping and scaling. An object can then be provided to the output video, which retains a reference size and, in some cases, remains at a common location within the video frames (e.g., at the center of each frame).

[0165] Figure 8A This is a diagram illustrating an example of a system for capturing and processing frames or images. Figure 8A The system includes: an image sensor 801, one or more image processing engines 803, a video processing engine 805, a display processing engine 807, an encoding engine 809, an image analysis engine 811, a sensor image metadata engine 813, and a frame cropping and scaling system 815. See below for reference. Figure 8B Describe an exemplary frame cropping and scaling system 800.

[0166] Figure 8A The system may include electronic devices or parts thereof, such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), IP cameras, desktop computers, laptop or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, or any other suitable electronic devices. In some examples, the system may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, the frame clipping and scaling system 800 may be implemented as... Figure 1 This is part of the image capture and processing system 100 shown.

[0167] although Figure 8AThe system is shown as including certain components, but those skilled in the art will understand that the system may include more than [other components]. Figure 8A More components are shown below. The components of the system may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, components may include and / or be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, GPU, DSP, CPU, and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof, and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. Software and / or firmware may include those stored on a computer-readable storage medium and may be implemented by... Figure 8A One or more instructions executed by one or more processors of an electronic system.

[0168] Image sensor 801 can perform the same functions as the referenced above. Figure 1 The image sensor 130 described operates similarly. For example, image sensor 801 may include one or more arrays of photodiodes or other photosensitive elements. Each photodiode can measure the amount of light corresponding to a specific pixel in an image generated by image sensor 130. In some examples, one or more image processing engines 803 may include a camera serial interface decoder module, an image front end, a Bayer processing section (e.g., for snapshot or preview images), an image processing engine, any combination thereof, and / or other components.

[0169] The video processing engine 805 can perform video encoding and / or video decoding operations. In some cases, the video processing engine 805 includes a combined video encoder-decoder (also known as a "CODEC"). The video processing engine 805 can perform any type of video decoding technology to encode and / or decode encoded video data. Examples of video decoding technologies or standards include Universal Video Decoding (VVC), High-Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), MPEG-2 Part 2 Decoding, VP9, ​​AOMedia Video 1 (AV1), and others. Using video decoding technologies, the video processing engine 805 can perform one or more prediction methods that utilize redundancy present in video images or sequences (e.g., inter-frame prediction, intra-frame prediction, etc.). The goal of video encoding is to compress video data to a form using a lower bit rate while avoiding or minimizing video quality degradation. The goal of video decoding is to decompress video data and obtain any other information in the encoded video bitstream that can be used for decoding and / or playback of the video data. The video output by the video processing engine 805 can be stored in memory 817 (e.g., a decoded picture buffer (DPB), random access memory (RAM), one or more cache memories, any combination thereof, and / or other memories), and / or can be output for display. For example, decoded video data can be stored in memory 817 for decoding other video frames and / or can be displayed on display 819.

[0170] Display processing engine 807 can be used to preview images. For example, display processing engine 807 can process, manipulate, and / or output a preview image that has the same (or similar in some cases) aspect ratio as a camera output image, but with a lower image resolution. The preview image can be displayed (as a “preview”) on a display of the system or a device including the system before the actual output image is generated.

[0171] Image decoding engine 809 can perform image encoding (compression) and / or image decoding (decompression) operations. In some cases, image decoding engine 809 includes a combined image encoder-decoder (or CODEC). Image decoding 809 can perform any type of image decoding technology to encode image data and / or decode compressed image data. Examples of image decoding technologies or standards include Joint Picture Experts Group (JPEG), Tagged Image File Format (TIFF), and others. Using image decoding technologies, image decoding engine 809 can leverage the visual perception and statistical properties of image data to compress images with lower loss of fidelity or quality.

[0172] Frame analysis engine 811 can perform frame or image analysis on preview frames acquired or received from display processing engine 807. For example, frame analysis engine 811 can acquire or receive a copy of the preview image from display processing engine 807 (the copy of the preview image has a lower image resolution compared to the camera output image). Frame analysis engine 811 can perform object detection and / or tracking operations on the preview image to detect and / or track one or more objects (e.g., target objects) in the image. Frame analysis engine 811 can determine and output the size, position, and center point (or other point) information of the bounding boxes for one or more tracked objects (e.g., tracked target objects). The bounding box information for one or more tracked objects can be output to frame cropping and scaling system 815.

[0173] The sensor frame metadata engine 813 generates and outputs the final output image. The sensor frame (or image) metadata represents the output image information and has the same image resolution as the output image.

[0174] Figure 8B This is a diagram illustrating an example of a frame cropping and scaling system 800, which can process one or more frames to maintain an object at a fixed size (and in some cases, at a fixed position) within one or more frames. In some cases, the frame cropping and scaling system 800 is... Figure 8A The example shown is a frame cropping and scaling system 815. In some cases, the frame cropping and scaling system 800 can be used with... Figure 8A The systems shown are separate. The frame cropping and scaling system 800 includes a region of interest (ROI) determination engine 804, an object detection and tracking system 806, a frame cropping engine 808, a frame scaling engine 810, and a smoothing engine 812. References will follow below. Figures 8C to 41 Examples describing the operations of the cropping and scaling system 800. In some examples, operations can be performed based on user-selected actions. Figure 8C Process 820 Figure 9A Process 930 Figure 9B The device may perform process 935 and / or other processes described herein. For example, the device may receive user input from a user (e.g., touch input via the device's touchscreen, voice input via the device's microphone, gesture input using one or more of the device's cameras, etc.) that instructs the device to capture video and maintain the object at a fixed size in the video. Based on the user input, the device may perform process 820, process 930 of Figure 9, and / or one or more other processes described herein.

[0175] The frame cropping and scaling system 800 may include an electronic device or part of an electronic device, such as a mobile or landline phone (e.g., a smartphone, cellular phone, etc.), an IP camera, a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, or any other suitable electronic device. In some cases, the frame cropping and scaling system 800 may be integrated with... Figure 8A The frame clipping and scaling system 800 is part of the same device as the system. In some examples, the frame clipping and scaling system 800 may include one or more wireless transceivers for wireless communication, such as cellular network communication, 802.11 Wi-Fi communication, wireless local area network (WLAN) communication, or some combination thereof. In some implementations, the frame clipping and scaling system 800 may be implemented as... Figure 1 This is part of the image capture and processing system 100 shown.

[0176] Although the frame cropping and scaling system 800 is shown as including certain components, those skilled in the art will understand that the frame cropping and scaling system 800 may include more than [other components]. Figure 8B Further components are shown. Components of the frame cropping and scaling system 800 may include software, hardware, or one or more combinations of software and hardware. For example, in some implementations, components of the frame cropping and scaling system 800 may include and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, GPU, DSP, CPU, and / or other suitable electronic circuitry) and / or may include computer software, firmware, or any combination thereof and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. The software and / or firmware may include one or more instructions stored on a computer-readable storage medium and executable by one or more processors of an electronic device implementing the frame cropping and scaling system 800.

[0177] Frame sequence 802 is input to frame cropping and scaling system 800. Frame 802 can be part of a frame sequence. The frame sequence can be video, a set of consecutively captured images, or other frame sequences. ROI determination engine 804 can determine the initial region of interest (ROI) in a specific frame based on user input and / or automatically. For example, in Figure 8CAt box 822, process 820 can perform object selection in the initial frame of the frame sequence (e.g., in video). The ROI can be represented by a bounding box or other boundary region. In some implementations, the boundary region is visible in the frame when output to a display device. In some implementations, the boundary region may be invisible (e.g., to an observer, such as a user) when outputting the frame to a display device. The frame that determines the initial ROI is called the initial frame (or reference frame) of the frame sequence.

[0178] In some examples, the ROI determination engine 804 (or other components of the frame cropping and scaling system 800) can determine a video frame from the frame sequence 802 to be used as the initial frame. In some cases, the first frame in the frame sequence 802 can be selected as the initial frame. In one illustrative example, such as in live video recording, the initial frame could be the first frame of the video based on input provided by the end user, indicating the desired size of an object in the frame (e.g., a pinch gesture to zoom to the desired camera zoom level), after which video recording can begin. In another illustrative example, such as in video playback (e.g., previously recorded video) or in any post-processing-based auto-zoom functionality, the end user can select any frame of the video and provide input for the frame (e.g., a pinch input to zoom) indicating the desired size of an object, which would result in that frame being set as the initial frame.

[0179] In some examples, the ROI can be determined based on user selection of a portion of the initial frame, such as an object depicted in the initial frame. User input can be received using any input interface, including a frame cropping and scaling system 800 or other devices. For example, the input interface can include a touchscreen, electronic drawing tools, a gesture-based user interface (e.g., one or more image sensors for detecting gesture input), a voice-based user interface (e.g., a speaker and a voice recognition tool for recognizing voice input), and / or other user interfaces. In some examples, object selection can include tapping (e.g., single-click, double-click, etc.) on an object displayed in the initial frame, the user drawing a bounding box around the object, the user providing input on a touchscreen interface (e.g., a pinch, such as bringing two fingers together or apart) causing the interface to zoom in on the object, or other types of object selection. In some cases, guidance can be provided to the end user on how to utilize features that keep the size of the target object constant throughout the video or other frame sequence. For example, a prompt can be displayed to the user indicating how to select an object that remains fixed throughout the video. For video, a user can select an object of interest by (e.g., on a touchscreen) tapping an object in the initial frame of the video or drawing a bounding box around that object. Based on the selected portion of the initial frame, the ROI determination engine 804 can define an ROI around the selected portion (e.g., around a selected object). The ROI indicates the size (e.g., the ideal size) of the object to be maintained throughout the video or other frame sequence. For example, a user can zoom in on an object to indicate the size the user expects the object to maintain throughout the video or other frame sequence, and the ROI determination engine 804 can use the indicated size to define the ROI around the object.

[0180] In some examples, objects in the initial frame can be automatically detected (e.g., using object detection and / or recognition), and the ROI determination engine 804 can define the ROI around the detected objects. Object detection and / or recognition techniques (e.g., face detection and / or recognition algorithms, feature detection and / or recognition algorithms, edge detection algorithms, boundary tracking functions, any combination thereof, and / or other object detection and / or recognition techniques) can be used to detect objects. Any of the above detection and tracking techniques can be used to automatically detect objects in the initial frame. In some cases, feature detection can be used to detect (or locate) the features of objects from the initial frame. Based on these features, object detection and / or recognition can detect objects, and in some cases, can identify and classify the detected objects into a category or type of object. For example, feature recognition can identify numerous edges and corners in a region of a scene. Object detection can detect that the edges and corners detected in that region belong to a single object. If face detection is performed, face detection can identify that the object is a face. Object recognition and / or face recognition can further identify the person corresponding to that face.

[0181] In some implementations, object detection and / or recognition algorithms can be based on machine learning models trained on images of the same type of objects and / or features using machine learning algorithms. These machine learning models can extract features from the images and are trained on the model based on the algorithms to detect and / or classify objects containing those features. For example, the machine learning algorithm can be a neural network (NN), such as a convolutional neural network (CNN), a time-delayed neural network (TDNN), a deep feedforward neural network (DFFNN), a recurrent neural network (RNN), an autoencoder (AE), a variant AE (VAE), a denoising AE (DAE), a sparse AE (SAE), a Markov chain (MC), a perceptron, or some combination thereof. The machine learning algorithm can be a supervised learning algorithm, an unsupervised learning algorithm, a semi-supervised learning algorithm, a generative adversarial network (GAN) based learning algorithm, any combination thereof, or other learning techniques.

[0182] In some implementations, computer vision-based feature detection and / or recognition techniques can be used. Different types of computer vision-based object detection algorithms can be used. In an illustrative example, template matching-based techniques can be used to detect one or more hands in an image. Various types of template matching algorithms can be used. An example of a template matching algorithm can perform Haar or Haar-like feature extraction, integral image generation, Adaboost training, and a cascaded classifier. This object detection technique performs detection by applying a sliding window (e.g., with a rectangle, circle, triangle, or other shape) onto the image. The integral image can be computed as an image representation evaluating features of a specific region (e.g., rectangular or circular features) from the image. For each current window, Haar features of the current window can be computed from the aforementioned integral image, which can be computed before the Haar features are computed.

[0183] Haar features can be computed by summing the image pixels within a specific feature region of an object image (e.g., a specific feature region of an integral image). For example, on a face, the region with the eyes is typically darker than the region with the bridge of the nose or cheeks. Haar features can be selected using a learning algorithm (e.g., the Adaboost learning algorithm), which selects the best features and / or trains a classifier using those best features. This can be used to classify a window as a face (or other object) window or a non-face window with a cascaded classifier. A cascaded classifier consists of multiple classifiers combined in a cascaded manner, allowing background regions of the image to be quickly discarded while performing more computation on the object-class regions. Using the face as an example of a body part for an external observer, the cascaded classifier can classify the current window into either the face category or the non-face category. If a classifier classifies the window as a non-face category, the window is discarded. Otherwise, if a classifier classifies the window as a face category, the next classifier in the cascaded arrangement is used for retesting. The window is labeled as a candidate for a hand (or other object) until all classifiers determine that the current window is a face (or other object). After all windows have been detected, the windows surrounding each face can be grouped using a non-maximum suppression algorithm to generate a final result of one or more detected faces.

[0184] return Figure 8BThe ROI determination engine 804 can define an ROI based on a selected portion of the initial image (e.g., a selected object) or based on objects detected in the initial image. As described above, an ROI can be represented by a bounding box or other type of boundary region. In some cases, the ROI determination engine 804 can generate a bounding box for an ROI that fits the boundaries of objects within the ROI. For example, for an object, the maximum x-coordinate (horizontal), minimum x-coordinate, maximum y-coordinate (vertical), and minimum y-coordinate can be determined, and the ROI can be defined with maximum x-coordinate, minimum x-coordinate, maximum y-coordinate, and minimum y-coordinate. In some cases, the bounding box for the ROI can be defined around the object, not limited to the boundaries of objects within the ROI.

[0185] The ROI determination engine 804 can determine the size of the object and / or the region of interest (ROI) including the object in the initial frame. The object size can be used as a reference size to determine how much cropping and scaling should be performed on subsequent frames of frame sequence 802. In some cases, the user can adjust the size of the ROI and / or the object to define an optimal size for the object in the frame sequence. For example, the first frame can be displayed (e.g., as a preview image), and the user can adjust the zoom amount of the image to make the object larger (by zooming in) or smaller (by zooming out). In such an example, once the user has finished zooming and the final object size has been determined for the initial frame, the size of the object and / or the ROI including the object is determined and used as a reference size. The reference size can then be used to determine how much cropping and scaling should be performed on subsequent frames in frame sequence 802.

[0186] Subsequent frames of frame sequence 802 (captured after the initial frame) can then be input to frame cropping and scaling system 800. The operation of object detection and tracking system 806, frame cropping engine 808, frame scaling engine 810, and smoothing engine 812 will be described for specific subsequent frames after the initial frame (e.g., the first subsequent frame that appears after the initial frame). However, the same or similar operations can be performed on some or all of the subsequent frames that occur after the initial frame in frame sequence 802.

[0187] The object detection and tracking system 806 can detect and track objects in subsequent frames of a frame sequence. For example, in Figure 8C At box 824, process 820 can perform object detection and tracking to detect and track objects in the frame sequence. In some examples, the above reference can be used. Figures 2 to 7 The video analytics system 200 described uses techniques to detect and track objects.

[0188] Frame cropping engine 808 can crop subsequent frames, and frame scaling engine 810 can scale subsequent frames so that the size of the object in subsequent frames remains the same as the size determined in the initial frame. For example, in Figure 8C At box 826, process 820 can perform cropping and scaling of video frames for subsequent frames. In some cases, cropping and scaling can be performed to keep an object having the same size as determined in the initial frame, and it can also keep the object in a specific position in each frame. For example, cropping and scaling can be performed to keep the object at the center of each subsequent frame, at the position in each subsequent frame where the object was initially in the initial frame, at a user-defined position, or at another position within the subsequent frame. As described in more detail below, frame scaling engine 810 can calculate a scaling factor for each subsequent frame of frame sequence 802. In an illustrative example using the diagonal length of the bounding box for explanation purposes, the scaling factor can be determined by comparing the diagonal length of the bounding box in the initial frame with the diagonal length of the bounding box in the current frame being processed. The ratio between the diagonal lengths can be used as the scaling factor. The scaling factor can be used to scale each subsequent frame so that the object in the current frame has the same size as the object in the initial frame. Details of cropping and scaling will be described below.

[0189] The smoothing engine 812 can apply one or more smoothing functions to progressively perform cropping and scaling on subsequent frames, minimizing frame-to-frame movement and resizing of objects in the frame sequence. For example, the initial cropping and scaling output from the frame cropping engine 808 and frame scaling engine 810 can indicate that subsequent frames will be cropped and scaled by a certain amount. The smoothing engine 812 can determine the modified cropping and scaling amounts to reduce the amount that will be modified in subsequent frames. The smoothing functions (one or more) can prevent objects from moving in an unnatural (e.g., skipping) manner in the frame sequence 802 due to the cropping and scaling amounts determined by the frame cropping engine 808 and frame scaling engine 810.

[0190] In some cases, cropping, scaling, and smoothing can be based on a point on the object (e.g., the center point) or a point within the bounding box associated with the ROI that includes the object (e.g., the center point), and / or can be based on a distance associated with the object (e.g., the distance between a first part and a second part of the object) or a distance representing the ROI including the object associated with the bounding box (e.g., the diagonal distance of the bounding box). For example, the amount of cropping performed relative to a point on the object or a point within the bounding box to move or shift an object in a subsequent frame can be performed. In another example, the amount of scaling performed to make an object larger or smaller can be based on a distance associated with the object (e.g., between different parts of the object) or a distance associated with the bounding box (e.g., the diagonal distance of the bounding box).

[0191] Frame cropping and scaling can be performed along with actual changes in the size of the target object. The smoothing engine 812 can output a final output frame 814 (e.g., output video) that will have the effect of objects of fixed size (based on a reference size determined for the objects in the initial frame) and, in some cases, remain in the same position throughout the entire frame sequence. For example, in Figure 8C At box 828, process 820 can generate an output video that includes fixed-size and positional effects for the object based on the aforementioned target fixed-size features.

[0192] Figure 9A and Figure 9B This is a flowchart illustrating other examples of processes 930 and 935 that can be performed by a frame cropping and scaling system 800 for video. In some examples, processes 930 and / or 935 can be performed based on an action selected by the user. For example, the device can receive user input from the user (e.g., touch input via the device's touchscreen, voice input via the device's microphone, gesture input using one or more cameras of the device, etc.) instructing the device to capture video and maintain objects at a fixed size in the video. Based on the user input, the device can perform processes 930 and / or 935.

[0193] Processes 930 and 935 are described as being performed for pre-recorded video (in which case all frames of the video are available for processing). However, in some cases, processes 930 and 935 can be modified to process live video. In some examples, process 930 can be performed before process 935. For example, process 930 can be performed to select an initial video frame from the frame sequence and set the object bounding box center point (or other point) and the object bounding box diagonal length (or other length) as reference points. Process 935 can be performed to crop and scale subsequent video frames to maintain the size and / or position of the object throughout the frame sequence.

[0194] like Figure 9A As shown, at box 931, process 930 includes obtaining a frame sequence. The frame sequence can be video, a set of consecutively captured images, or other frame sequences. At box 932, process 930 includes selecting or determining a video frame from the frame sequence to use as an initial frame (or reference frame). In one example, the first frame in frame sequence 802 can be selected as the initial frame. As described above, the initial frame can be used as a frame for determining an initial ROI.

[0195] At box 933, process 930 includes: selecting a target object having a given size (e.g., an ideal size). As described above, the target object (or ROI) can be selected based on user input, or the target object (or ROI) can be detected automatically. For example, an object or ROI can be determined based on user input indicating the selection of a portion of an initial frame, such as an object depicted in the initial frame. In an illustrative example, a user can pinch to zoom (e.g., using a pinch gesture on a touchscreen interface) or provide another input to cause the display to zoom in on the target object. In some cases, process 930 may include: generating a bounding box for the target object or ROI in the initial frame of a frame sequence (e.g., a video). For example, an ROI can be determined for an object, and a bounding box can be generated to represent the ROI.

[0196] At box 934, process 930 includes setting the bounding box center point and diagonal length as references for subsequent frames in the frame sequence (e.g., for performing process 935 on subsequent frames). While the bounding box center point and diagonal length are used herein for illustrative purposes, other points and lengths can be used to perform the cropping, scaling, smoothing, and / or other operations described herein. In some examples, instead of the bounding box center point, different points on the bounding box can be used as reference points, such as the top-left corner of the bounding box. In another example, points on an object within the bounding box can be used as reference points, such as the object's center point. In some examples, instead of the bounding box diagonal length, the length between two points of the object within the bounding box can be used to determine the size of the object in the current subsequent frame. For example, if the object is a person, the length between the top of the person's head and the bottom of the person's feet can be used as the length.

[0197] like Figure 9B As shown, at box 937, process 935 includes performing object detection and tracking for each subsequent frame (or subset of subsequent frames) in the frame sequence following the initial frame, similar to box 824 of process 820. Object detection and tracking can be performed to track the object across each frame of the video. In some examples, process 935 may perform coordinate transformations to match each subsequent frame with the initial frame. For example, coordinate transformations may be performed to make each subsequent frame have the same size as the initial frame. In one illustrative example, the coordinate transformation could be a zoom-in process. In another illustrative example, the coordinate transformation could be a zoom-out process.

[0198] At box 938, process 935 includes determining the center point of bounding boxes and the diagonal length of the bounding boxes throughout the entire frame sequence of the video. For example, based on object detection and tracking performed by object detection and tracking system 806, bounding box information can be obtained for each frame from the video. The center point position and diagonal length of each bounding box in each video frame can be determined and used as an indication of object movement trajectories and object size changes throughout the video.

[0199] For example, all frames in a video frame sequence can be processed to determine the center point and diagonal length of each bounding box in each frame. Frame cropping engine 808, frame scaling engine 810, and / or smoothing engine 812 can use the center point and diagonal length of the bounding box in each frame to perform cropping, scaling, and smoothing on subsequent frames of the video (respectively). For example, the center point of the bounding box can be used as a reference to determine the position of an object within the frame, and the diagonal length can be used to determine the size of the object in the current subsequent frame relative to the size of the object in the initial frame. Although this document uses the bounding box center point and the diagonal length of the bounding box for illustrative purposes, in some implementations, other points and lengths can be used to perform cropping, scaling, and smoothing. In some examples, instead of the center point of the bounding box, different points on the bounding box can be used as reference points, such as the top-left corner of the bounding box. In another example, points on the object within the bounding box can be used as reference points, such as the center point of the object. In some examples, instead of the diagonal length of the bounding box, the length between two points of the object within the bounding box can be used to determine the size of the object in the current subsequent frame. For example, if the object is a person, the length between the top of the person's head and the bottom of the person's feet can be used as the length.

[0200] Box 939 represents a smoothing operation that can be performed by the smoothing engine 812. At box 940, process 935 includes performing bounding box center point trajectory smoothing. The smoothing engine 812 can perform bounding box center point trajectory smoothing based on any suitable smoothing algorithm. An example of a smoothing algorithm is based on a moving average algorithm. Moving average techniques can be applied to smooth variations in the position of the bounding box center point and the diagonal length in subsequent frames. Typically, moving averages are used to analyze time-series data (such as video) by calculating the average of different subsets of a complete dataset (e.g., different frames of a video). Based on moving averages, data can be smoothed to result in fewer significant variations between consecutive portions of the data.

[0201] Moving averages can be based on a sliding window, which is used to average over a set number of time periods (e.g., multiple video frames). For example, the number of time periods could be based on the time between consecutive frames of a video (e.g., 33ms in a 30 frames per second video). A moving average can also be an equally weighted average of the first n data points. For example, a sequence of n values ​​could be defined as:

[0202]

[0203] Then, the equally weighted rolling average over the n data points will essentially be the mean of the previous M data points, where M is the size of the sliding window:

[0204]

[0205]

[0206] To calculate the subsequent rolling average, new values ​​can be added to the sum, and values ​​from previous time periods can be discarded. Previous time periods can be discarded because their averages are available, eliminating the need for a full summation each time. The calculation of the subsequent rolling average can be expressed by the following formula:

[0207]

[0208] For system 800 processing the current frame of the video being processed by process 935, a moving average formula can be used to process the (x, y) coordinates of the center points of the bounding boxes of a certain number of video frames out of M video frames. For example, at box 940, the rolling average of the coordinates of the center points of the bounding boxes of the M video frames can be determined. Then you can use the rolling average. Used as the center point location of the bounding box in the current video frame.

[0209] At box 941, process 935 includes performing smoothing on the change in the size of the bounding box diagonal length. The smoothing engine 812 can smooth the change in the size of the bounding box diagonal length based on any suitable smoothing algorithm. In some cases, the smoothing engine 812 can use the moving average algorithm described above. For example, for the current frame of the video being processed, the smoothing engine 812 can use a moving average formula to process the diagonal length of the bounding boxes of a certain number of video frames from the M video frames. For example, at box 942, process 935 can determine a rolling average of the diagonal lengths of the bounding boxes of the M video frames. Processing 935 can reduce the rolling average. Used as the diagonal length of the bounding box in the current video frame.

[0210] In some cases, object detection and tracking may be inaccurate for the current frame of the video being processed. For example, the (x, y) coordinates of the center point of the detected object's bounding box (or other points of the bounding box or object) may be incorrect. The calculated average of the movement (or scrolling) of the bounding box center point coordinates for the current frame is crucial. Error alerts can be minimized (by minimizing the object bounding box position for error detection / tracking) and the object's correct movement or trajectory can be largely maintained. For example, the calculated average movement (or scrolling) of the bounding box center point (or other points on the bounding box or object). It may be more accurate than the actual detected center point. Moving averages can also minimize false alarms about object size (e.g., by minimizing the length of the object bounding box diagonal or the length of the error detection / tracking between parts or portions of the object). For example, the moving (or scrolling) average of the bounding box diagonal length calculated in a given frame. It can be more accurate than the actual detected diagonal length.

[0211] At box 942, process 935 includes calculating a frame scaling factor based on the diagonal length of the initial frame in the video and the smoothed diagonals of other frames. For example, instead of using the actual diagonal length of the bounding box in the current frame, a smoothed diagonal length (e.g., average diagonal length) determined by smoothing engine 812 for the current frame can be used to determine the scaling factor for the current frame (not the initial frame) of the video. In some cases, the scaling factor may be a scaling rate. Frame scaling engine 810 may compare the smoothed diagonal length of the bounding box in the current frame with the diagonal length of the bounding box in the initial frame to determine the scaling factor for the current frame.

[0212] In some examples, process 935 may include determining whether a video resource change has occurred. For example, frame cropping and scaling system 815 may support multiple video resources, where an end user can import multiple videos for performing automatic zoom (cropping and scaling) operations. To determine whether a video resource change has occurred, process 935 may determine whether the current video is still playing. If the current video is playing, process 935 may continue to box 943. If it is determined that another video source has been started, the update operation may not be performed, in which case the system can restart from the beginning of the process (e.g., starting at box 931 of process 930).

[0213] At box 943, process 935 includes cropping and scaling each frame in the video based on the smoothed object bounding box center point position at box 939 (e.g., the average center point position determined for each frame) and a frame scaling factor determined for each frame. Based on the cropping and scaling, a cropped and scaled subsequent frame is generated, and the object in the subsequent frame has the same size and the same relative position as the object in the initial frame.

[0214] refer to Figure 10A and Figure 10B An example is described. Figure 10AThis is a diagram illustrating an example of the initial frame 1002 of the video. The user has selected a person as the object of interest. A bounding box 1004 is generated to represent the region of interest for the person. The bounding box 1004 is shown as having a height h and a width w. The position of the center point 1006 of the bounding box 1004 (e.g., the (x, y) coordinate position) and the diagonal length 1008 of the bounding box 1004 are determined and used as a reference for cropping and scaling subsequent frames of the video to maintain a constant size and position for the person in subsequent frames.

[0215] Figure 10B This is a diagram illustrating an example of a subsequent frame 1012 appearing after the initial frame 1002 in a video. Based on object detection and tracking, a bounding box 1014 is generated around the person in the subsequent frame 1012. The bounding box 1014 has a width wn and a height hm. The width wn of the bounding box 1014 is less than the width w of the bounding box 1004 in the initial frame 1002, while the height hm of the bounding box 1014 is less than the height h of the bounding box 1004 in the initial frame 1002. The position of the center point 1016 (e.g., (x, y) coordinate position) and the diagonal length 1008 of the bounding box 1004 are determined.

[0216] In some examples, frame cropping engine 808 can crop subsequent frames 1012 such that the person depicted in subsequent frames 1012 remains centered in frame 1012. For example, frame cropping engine 808 can crop subsequent frames 1012 to generate a cropped region 1022 such that the center point 1016 of bounding box 1014 is located at the center of cropped region 1022. In some examples, frame cropping engine 808 can crop subsequent frames 1012 such that the person shown in frame 1012 remains in the same relative position as the person was in the initial frame 1002. For example, frame cropping engine 808 can determine the position of the center point 1006 of bounding box 1004 relative to a point in the initial frame 1002 that is common across all frames of the video. For example, the common point across all frames could be the top-left corner of a video frame (e.g., the top-left corner 1007 in the initial frame 1002). Figure 10A The relative distance 1009 shown is from the center point 1006 of the bounding box 1004 in the initial frame 1002 to the top-left corner point 1007. The frame cropping engine 808 can crop the subsequent frame 1012 to generate a cropped region 1022 such that the relative position and distance 1029 of the center point 1016 relative to the top-left corner point 1017 of the cropped region 1022 are the same as the relative position and distance of the center point 1006 in the initial frame 1002 relative to the top-left corner point 1007.

[0217] Frame scaling engine 810 determines a scaling factor (e.g., scaling ratio) for scaling the clipped region 1022 by comparing the smoothed diagonal length of bounding box 1014 in subsequent frame 1012 with the diagonal length 1008 of bounding box 1004 in initial frame 1002. As described above, the smoothed diagonal length of bounding box 1014 can be determined by smoothing engine 812. For example, if the actual diagonal length 1018 of bounding box 1014 is 1.5, the smoothed diagonal length of bounding box 1014 can be determined to be a value of 1.2 (based on the scroll average determined as described above). The diagonal length 1008 of bounding box 1004 in initial frame 1002 could be a value of 3. The scaling factor can be determined as the scaling ratio. Using this formula based on diagonal lengths of 1008 and 1018, it is possible to determine... The scaling factor is 2.5. Based on the scaling factor of 2.5, the cropped area 1022 can be increased by 2.5 times (increased by 2.5 times).

[0218] Due to cropping and scaling, a cropped and scaled subsequent frame 1032 is generated. The diagonal length 1038 of bounding box 1034 is the same as the diagonal length 1008 of bounding box 1004, and therefore the size of the person depicted in the cropped and scaled subsequent frame 1032 is the same as the size of the person depicted in the initial frame 1002. In some examples, the center point 1036 of bounding box 1034 is located at the center of the cropped and scaled subsequent frame 1032. In some examples, the position and distance 1039 of the center point 1036 of frame 1032 relative to the top-left corner point 1037 are the same as the position and distance 1009 of the center point 1006 relative to the top-left corner point 1007 in the initial frame 1002, resulting in the person being maintained in the same position in the cropped and scaled subsequent frame 1032 as the person in the initial frame 1032. Therefore, the size of the person depicted in the cropped and scaled subsequent frame 1032 is the same as the size of the person depicted in the initial frame 1002, and is kept in a position consistent with other frames in the entire video.

[0219] Back Figure 9BThe process 935 at box 944 includes performing video stabilization. Any suitable video stabilization technique can be used to stabilize video frames. Typically, video stabilization techniques are used to avoid visual quality loss by reducing unwanted shaking and jitter of the device (e.g., mobile devices, handheld cameras, head-mounted displays, etc.) during video capture. Video stabilization reduces shaking and jitter without affecting moving objects or intentional camera panning. Video stabilization can benefit handheld imaging devices (e.g., mobile phones), which can be significantly affected by shaking due to their small size. Unstable images are often caused by unintended hand shaking and intentional camera panning, where unintended positional fluctuations of the camera lead to unstable image sequences. Using video stabilization techniques ensures high visual quality and stable video footage even under suboptimal conditions.

[0220] One example of a video stabilization technique that can be implemented is a fast and robust two-dimensional motion model based on the Euclidean transform, which can be used to solve video stabilization problems. In the Euclidean motion model, a square in the image can be transformed into any other square with different positions, sizes, and / or rotations for motion stabilization (because camera movement between consecutive frames of a video is typically small). Figure 11 This is a diagram illustrating an example of an applied motion model, which includes an initial square and various transformations applied relative to the initial square. These transformations include transformations, Euclidean transformations, affine transformations, and homography.

[0221] Figure 12 This is a flowchart illustrating an example of a process 1200 for performing image stabilization. The image stabilization process includes tracking one or more feature points between two consecutive frames. The tracked features allow the system to estimate motion between multiple frames and compensate for that motion. An input frame sequence 1202, comprising a frame sequence, is provided as input to process 1200. The input frame sequence 1202 may include an output frame 814. At block 1204, process 1200 includes performing salient point detection using optical flow. Salient detection is performed to determine feature points in the current frame. At block 1204, any suitable type of optical flow technique or algorithm can be used. In some cases, optical flow motion estimation can be performed on a pixel-by-pixel basis. For example, for each pixel in the current frame y, motion estimation f defines the position of the corresponding pixel in the previous frame x. Motion estimation f for each pixel may include an optical flow vector that indicates the movement of the pixel between multiple frames. In some cases, the optical flow vector for a pixel can be a displacement vector (e.g., indicating horizontal and vertical displacements, such as x-displacement and y-displacement), which shows the movement of the pixel from the first frame to the second frame.

[0222] In some examples, optical flow maps (also known as motion vector maps) can be generated based on the computation of optical flow vectors across multiple frames. Each optical flow map can include a 2D vector field, where each vector is a displacement vector indicating the movement of a point from one frame to another (e.g., indicating horizontal and vertical displacements, such as x-displacement and y-displacement). An optical flow map can include an optical flow vector for each pixel in a frame, where each vector indicates the movement of the pixel across multiple frames. For example, dense optical flow can be computed between adjacent frames to generate an optical flow vector for each pixel in the frame, which can be included in the dense optical flow map. In some cases, an optical flow map can include vectors for fewer than all pixels in a frame, such as vectors for pixels belonging only to one or more parts of a tracked external observer (e.g., the external observer's eye, one or more hands, and / or other body parts). In some examples, Lucas-Kanade optical flow can be computed between adjacent frames to generate optical flow vectors for some or all pixels in the frame, which can be included in the optical flow map.

[0223] As described above, this can be done between adjacent frames in a frame sequence (e.g., between adjacent frames). and The optical flow vector or optical flow map is calculated between sets of frames. Two adjacent frames can include two directly adjacent frames, which are frames captured consecutively, or two frames separated by a certain distance in the frame sequence (e.g., within two frames, within three frames, or other suitable distances). From frames to frame Optical flow can be generated by O , = dof( , The optical flow map is given by , where dof is the dense optical flow. Any suitable optical flow procedure can be used to generate the optical flow map. In an illustrative example, the frame... Pixels Can move distance in the next frame Assuming the pixels are the same, and Frame and next frame If the strength between them remains unchanged, then we can assume the following procedure:

[0224]

[0225] By approximating the equation using the Taylor series on the right-hand side, then removing the common terms and dividing by... The optical flow equation can be derived as follows:

[0226] ,

[0227] in:

[0228] ;

[0229] ;

[0230] ;

[0231] ;as well as

[0232] .

[0233] Using the optical flow equation above, the image gradient can be found. and and the gradient along time (denoted as) ) Items u and v are the velocity or optical flow. The x and y components of the optical flow are unknown. In some cases, when it is impossible to solve the optical flow equation with two unknown variables, an estimation technique may be required. Any suitable estimation technique can be used to estimate the optical flow. Examples of such estimation techniques include difference methods (e.g., Lucas-Kanade estimation, Horn-Schunck estimation, Buxton-Buxton estimation, or other suitable difference methods), phase correlation, block-based methods, or other suitable estimation techniques. For example, Lucas-Kanade assumes that the optical flow (the displacement of an image pixel) is small and approximately constant in the local neighborhood of pixel I, and uses the least squares method to solve the basic optical flow equation for all pixels in that neighborhood.

[0234] At box 1206, process 1200 includes selecting correspondences between salient points in consecutive images. At box 1208, process 1200 performs transform estimation based on noise correspondences. At box 1210, process 1200 includes applying transform approximation and smoothing to generate an output frame sequence 1212 including an output frame sequence. For example, key feature points can be detected from previous image frames and the current image frame, and then the feature points with one-to-one correspondences can be used. Based on the location of the feature points used, a region-based transform can be applied to map image content from the previous frame to the current image frame.

[0235] In some examples, Figure 9B The entire process 935 involves video frame extraction and merging before and after. For example, in some cases, the input and output of system 800 may include image frames (instead of video), in which case video frame extraction and merging are required before and after the entire process.

[0236] In some examples, the inherent zoom rate and camera lens switching capabilities of a device (e.g., a mobile phone or smartphone) can be used to perform one or more of the techniques described herein. For example, the system can output a video with the fixed-size effect of the target object described herein. In some cases, this solution can be used as a real-time feature (for live video) and can automatically adjust the camera zoom rate during video recording.

[0237] In some examples, one or more of the techniques described above and / or other techniques may be used to perform automatic zoom operations. Figure 13A This is a diagram illustrating an example of a process 1300 for performing various aspects of autofocus. For example, process 1300 may determine or set as a reference point (e.g., center point) and distance (e.g., diagonal length) of the bounding box of an object and / or region of interest in a first frame (or initial frame). In some examples, process 1300 may begin based on an autofocus operation selected by a user. For example, the device may receive user input (e.g., touch input via the device's touchscreen, voice input via the device's microphone, gesture input using one or more cameras of the device, etc.) instructing the device to enter autofocus mode. Based on the user input, the device may execute process 1300. In some examples, once an autofocus operation is selected, the device may begin using an object detection and tracking system (e.g., ...). Figure 8B An object detection and tracking system 806 is used to perform object detection and tracking on any region or object of interest.

[0238] At box 1302, process 1300 includes obtaining the first frame (or initial frame) of the frame sequence (e.g., the first video frame of the video to which the user-identified object and / or region of interest is targeted). At box 1304, process 1300 includes: determining the target object of interest in the first frame. For example, as described above, Figure 8B The Region of Interest (ROI) determination engine 804 can determine the ROI in the first frame based on user input and / or automatically. The ROI may correspond to a target object (or object of interest). The ROI and / or target object may be represented by a bounding box or other bounded region. In some examples, the bounding box is visible in the frame when outputting to a display device. In some examples, the bounding box may not be visible when outputting the frame to a display device. The frame in which the initial ROI is determined is called the initial frame (or reference frame) of the frame sequence.

[0239] As described above, in some examples, the ROI can be determined based on user selection of a portion of the initial frame, such as an object depicted in the initial frame. For example, the user can select the target object to be used during auto-zoom so that the object maintains a fixed size (e.g., the size of the object in the initial frame) across multiple frames in a frame sequence. User input can be received using any input interface of the device, such as a touchscreen, electronic drawing tools, gesture-based user interfaces (e.g., one or more image sensors for detecting gesture input), voice-based user interfaces (e.g., speakers and voice recognition tools for recognizing voice input), and / or other user interfaces. (Refer to above) Figure 8C Any inputs and / or other inputs described in Figure 9 can be provided by the user. For example, object selection can be performed based on a tap (e.g., single click, double click, etc.) on an object displayed in the initial frame, a bounding box drawn by the user around the object, or other types of object selection. In some cases, guidance can be provided to the end user on how to utilize features that maintain the target object's size throughout the video or other frame sequence. For example, a prompt can be displayed to the user indicating how to select objects that remain fixed throughout the video. For video, a user can select an object of interest by tapping (e.g., on a touchscreen) an object in the initial frame of the video or by drawing a bounding box around the object. Based on the selected portion of the initial frame, the ROI determination engine 804 can define the ROI around the selected portion (e.g., around the selected object).

[0240] In some examples, objects in the initial frame can be automatically detected (e.g., using object detection and / or recognition), and the ROI determination engine 804 can define the ROI around the detected objects. Object detection and / or recognition techniques (e.g., face detection and / or recognition algorithms, feature detection and / or recognition algorithms, edge detection algorithms, boundary tracking functions, any combination thereof, and / or other object detection and / or recognition techniques) can be used to detect objects.

[0241] At box 1306, process 1300 includes determining or setting a point and distance of the object's bounding box as a reference. In one illustrative example, the point may include the center point of the bounding box. In some cases, other points of the bounding box may also be used, such as the top-left corner or a corner of the bounding box. In another example, a point on the object within the bounding box may be used as a reference point, such as the center point of the object. In another illustrative example, the distance may be the diagonal length of the bounding box (e.g., the length from the bottom-left corner to the top-right corner, or the length from the bottom-right corner to the top-left corner). In some examples, the distance may include the length between two points of the object within the bounding box. For example, if the object is a person, the length between the top of the person's head and the bottom of the person's feet may be used as the length.

[0242] By setting the object's center point (or other point) and diagonal length (or other distance) of the bounding box, process 1300 can initialize target object information, including the object's center point coordinates, the object's bounding box diagonal length, and the current zoom level for the object.

[0243] Figure 13B This is a diagram illustrating an example of a process 1310 for performing an additional aspect of autofocus on one or more subsequent frames captured after the initial frame (e.g., occurring after the initial frame in a frame sequence). At box 1312, process 1310 includes: obtaining one or more subsequent frames. In some cases, process 1310 may be performed in a single iteration for one frame from the one or more subsequent frames at a time. In some cases, process 1310 may be performed in a single iteration for multiple frames from the one or more subsequent frames at a time. The subsequent frame processed by process 1310 is referred to as the current subsequent frame.

[0244] At box 1314, process 1310 includes: obtaining a frame from display processing engine 807. This frame may be referred to as an analysis frame or a preview frame. (As mentioned above...) Figure 8A As described, a preview (or analysis) frame can have the same aspect ratio as the output frame, but with a lower resolution (smaller size). For example, a preview frame can be a lower-resolution version of the current subsequent frame compared to the full output version. A frame cropping and scaling system (e.g., frame cropping and scaling system 815 and / or frame cropping and scaling system 800) can use the preview frame for object detection and tracking processing. For example, at box 1316, process 1310 performs object detection and / or tracking to detect and / or track target objects (determined from the initial frame) in the preview frame (the lower-resolution version of the current subsequent frame processed by process 1310). As described above, frame analysis engine 811 can perform object detection and / or tracking on the analysis (preview) frame.

[0245] At box 1318, process 1310 performs a coordinate transformation on the preview (analysis) frame. For example, since the preview frame and the sensor frame metadata (corresponding to the full output frame) have the same image content but different image resolutions, a coordinate transformation can be performed to make the preview frame and the full output frame the same size. In an illustrative example, the coordinate transformation could be a zoom-in process. For example, process 1310 could zoom in on the preview frame to give it the same resolution as the full output frame corresponding to the sensor frame metadata.

[0246] At box 1320, process 1310 determines the points (e.g., center point or other points) and zoom level of the target object in the current subsequent frame based on the tracked target object information. The tracked target object information includes information associated with the target object detected and tracked from the current subsequent frame. The tracked object information may include: the detected object bounding box for the target object, the location of the bounding box, and the center point (or other point) of the bounding box. The points determined for the target object may include the same points determined for the target object in the initial frame. For example, if the point determined for the target object in the initial frame at box 1306 is the center point of an object or ROI, the point determined for the target object in the current subsequent frame at box 1320 may also include the center point of the object or ROI.

[0247] At box 1322, process 1310 includes determining or calculating a step value for the object point (e.g., center point) determined at box 1320 and a step value for the scaling factor determined at box 1320. In an illustrative example, the step value for the x-coordinate of the point can be determined as diff_x = (curr_x - prev_x) / frame_count, which is a linear step function. The term frame_count can be a constant integer and can be defined as any suitable numerical value (e.g., the value 1, 2, 3, or other suitable integer). Using a linear step function, the step count can be determined as the difference between the x-coordinate of the center point of the target object in the current subsequent frame and the x-coordinate of the center point of the target object in a previous frame (e.g., the frame before the current subsequent frame of the video) divided by the frame count. In another illustrative example, the step value for the y-coordinate of the point can be determined as diff_y = (curr_y - prev_y) / frame_count, similar to the step value for the x-coordinate. In another illustrative example, the step size for the scaling ratio can be determined as diff_zoom = (curr_ratio - prev_ratio) / frame_count. For example, the step size for the scaling ratio can be determined as the difference between the scaling ratio of the target object in the current subsequent frame and the scaling ratio of the target object in the previous frame (e.g., the frame before the current subsequent frame of the video) divided by the number of frames.

[0248] At box 1324, process 1310 includes obtaining sensor frame metadata from sensor frame metadata engine 813. As described above, sensor frame metadata can represent output image information and has the same image resolution as the output image. The image metadata frame has the same aspect ratio as the preview frame but with a higher resolution.

[0249] At box 1326, process 1310 includes determining an updated scaling factor and an updated point (e.g., center point) based on a step size value (e.g., determined using the linear step function described above). The step size value is calculated from a linear step size, where the parameter is progressively increased from a starting value to a stopping value using the number of steps in a linear interval sequence. The number of steps run will always be the parameter input to the step size field.

[0250] At box 1328, process 1310 includes performing scaling smoothing and / or bounding box point trajectory smoothing operations based on the output from box 1320 (object scaling and points determined for objects in the current subsequent frame) and the output from box 1326 (updated scaling and points for objects). For example, smoothing engine 812 can determine a smoothing value for the scaling by performing scaling smoothing. In another example, smoothing engine 812 can determine a smoothing value for the center point of an ROI or object by performing bounding box center point trajectory smoothing. As described above, scaling smoothing smooths changes in the size of the bounding box (e.g., changes in the diagonal length), allowing the size of a target object in the image to change gradually frame by frame. Bounding box point trajectory smoothing allows objects (e.g., based on the object's center point) to move gradually from frame to frame.

[0251] In some examples, the smoothing engine 812 can use the moving average algorithm described above to perform scaling smoothing and / or bounding box point trajectory smoothing operations. In some examples, the smoothing engine 812 can use a Gaussian filter function for scaling smoothing. Figure 14 Figure 1400 shows an example of a Gaussian filter smoothing function. For example, a Gaussian filter function with a window size N can be used, where N represents an empirical threshold that can be set to any suitable value (e.g., N = 31 or other values). An illustrative example of a Gaussian filter smoothing function is shown below (and the window size N is shown as window_size):

[0252] function f = gaussian(window_size)

[0253] sigma = double(window_size) / 5;

[0254] h = exp(-((1:window_size) - ceil(window _size / 2)).^2 / (2* sigma ^2));

[0255] f = h(:) / sum(h)

[0256] end

[0257] In some examples, the smoothing engine 812 can use a median filter function with a window size of M or other values ​​for scaling smoothing. In some examples, the smoothing engine 812 can use a Fibonacci sequence filter function with a window size of M for scaling smoothing. M represents an empirical threshold that can be set to any suitable value, such as M = 31 or other values. Figure 15This is graph 1500 illustrating the Fibonacci filter smoothing function. An illustrative example of the Fibonacci filter smoothing function is shown below:

[0258]

[0259]

[0260]

[0261] At box 1330, process 1310 includes updating the region of the current subsequent frame for zooming. For example, process 1310 may send this region as zoom information (e.g., a zoom rectangle for magnification or upsampling as the final output frame) to a camera pipeline, such as image sensor 801, image capture device 105A including image sensor 130, and one or more zoom control mechanisms 125C. In one example, one or more zoom control mechanisms 125C of image capture device 105A may use the zoom information (the region for zooming) to crop and scale the captured frame so that the object has the desired zoom level. Illustrative examples of the information are provided below:

[0262] curr_ratio += diff_zoom

[0263] curr_x += diff_x

[0264] curr_y += diff_y

[0265] Where curr_ratio is the zoom ratio value of the previous frame, and curr_x and curr_y are the x and y coordinates of the center position of the previous frame, respectively. The signs of diff_zoom, diff_x, and diff_y represent the zoom step size of the camera and the center position of the current frame, respectively.

[0266] At box 1332, process 1300 outputs a frame that has been cropped and scaled from the initially captured frame, thereby keeping the target object at the size it was in the initial frame.

[0267] In some examples, automatic zoom can be performed based on audio analysis, as a supplement to or alternative to using one or more of the techniques described above. For example, by analyzing audio data associated with video, the system can automatically focus on a salient or target object emitting sound. In some cases, the audio source is automatically magnified and focused on the salient object as the camera zooms in. In some examples, background noise can be removed. For example, if a user is recording a video of a person during a performance, the person's voice can be enhanced (e.g., made clearer, such as by increasing volume, eliminating background noise, etc.) as the user zooms in on the person. Such techniques can be used to generate or record video with a consistent target object size at a specific point (e.g., at the center point) within one or more video frames. Such techniques can be applied to real-time video recording and / or other use cases.

[0268] Figure 13C and Figure 13D These are diagrams illustrating examples of processes 1340 and 1350 used to perform automatic zoom based on analyzed audio. (Reference) Figure 13C At box 1342, process 1340 includes: obtaining a first (or initial) audio-video source. The first audio-video source may include a video frame and audio information associated with the video frame.

[0269] At box 1344, process 1340 performs visual processing to process video data from the first audio-video source to detect one or more candidate target objects. For example, one or more target objects may be detected in a given frame. Visual processing may include detecting one or more salient objects (e.g., candidate objects of interest) from the video frame. At box 1346, process 1340 performs audio processing to process audio data from the first audio-video source to detect sounds associated with target objects. Audio processing may include audio recognition and / or classification to identify and / or classify audio associated with the video frame. In an illustrative example, a deep learning neural network (e.g., Figure 42 Deep learning network 4200, Figure 43 The neural network uses a convolutional neural network 4300, or other deep neural networks, to perform visual and audio processing. In such an example, the input is video (audio-video source), and the output of the neural network is an image with at least one object highlighted as the source of the sound.

[0270] At box 1347, process 1340 includes: determining, based on audio processing, whether the detected candidate object (detected based on visual processing) is emitting sound. If it is determined that the candidate object is emitting sound, then it is similar to... Figure 13AIn process 1300, box 1306, and in process 1340 at box 1348, process 1340 may include: determining or setting object bounding box points and distances as references. In an illustrative example, a point may include the center point of the bounding box, and a distance may be the diagonal length of the bounding box (e.g., the length from the bottom-left corner to the top-right corner, or from the bottom-right corner to the top-left corner). Other points and / or distances may be used as described above. If it is determined at box 1347 that a candidate target object is not emitting sound, the object can be analyzed to determine whether the next candidate target object is emitting sound. Similar to a reference. Figure 13A As described, by setting the points of the bounding box (e.g., the center point of the object) and the distances (e.g., the diagonal length), process 1340 can initialize target object information, including the coordinates of the object's center point, the diagonal length of the object's bounding box, and the current zoom level for that object.

[0271] refer to Figure 13D Process 1350 is similar to Figure 13B The process 1320 also includes audio processing operations (at boxes 1352, 1354, 1356, 1358 and 1360). Figure 13D The ones with Figure 13B The boxes containing similar numbers are referenced above. Figure 13B The following description is provided. At box 1352, process 1350 includes performing audio three-dimensional (3D) localization. 3D sound localization refers to an acoustic technique used to locate the source of sound in 3D space. The source location can be determined by the direction of the incident sound wave (e.g., horizontal and vertical angles) and the distance between the source and the sensor. Once the audio 3D reset position is performed, process 1350 proceeds to box 1332 to output a cropped and scaled frame, as referenced above. Figure 13B As described.

[0272] At box 1354, process 1300 includes obtaining subsequent audio. The subsequent audio may be audio associated with one or more subsequent frames obtained at box 1312. At box 1356, process 1300 includes updating by amplifying the audio source and increasing its volume.

[0273] At box 1358, process 1300 includes performing background noise reduction. For example, audio background noise (such as paper creaking, keyboard typing, fan noise, dog barking, and other noise) reduces the audible perception of an audio signal. Audio background noise cancellation can help eliminate interfering noise and filter out distracting noise to create a better audio experience. At box 1360, process 1300 includes outputting the audio associated with the frame output at box 1332.

[0274] Figure 16 This is a diagram illustrating the zoom process in a camera pipeline (e.g., image sensor 801, image capture device 105A including image sensor 130, and one or more zoom control mechanisms 125C, etc.). As shown, the image capture device can stream (e.g., by default) output frames with a 1.0x zoom ratio (meaning zero or no zoom). The region of interest (ROI) 1604 (also called a crop rectangle or zoom rectangle) is shown in the frame 1602 with a 1.0x zoom ratio. For example, as described above, Figure 8B The ROI determination engine 804 can determine the initial region of interest (ROI) in a specific frame based on user input and / or automatically. In an illustrative example, the user can provide user input to define the zoom ROI 1604, including the rectangle's position and size. As shown, the zoom ROI 1604 is cropped from frame 1602. Once cropped from frame 1602, the zoom ROI 1604 is magnified or upsampled for the output stream (shown as magnified frame 1606).

[0275] Figure 17 This is a diagram showing the zoom delay of a camera pipeline with a seven (7) frame delay for zoom requests. Figure 17 The example shown represents a frame delay of seven (7) frames, where a zoom request will be applied after seven frames in a given frame. For example, for a zoom request based on... Figure 17 A request for 1.1x zoom (1702) made in frame 1 will be applied with the corresponding zoom adjustment at frame 8. As shown, frame 8 has a zoom amount of 1.1. The zoom increment can be adjusted every frame. For example, a request for 1.2x zoom (1704) can be made based on frame 2, and the corresponding zoom adjustment (shown as a zoom amount of 1.2) will be applied at frame 9, seven frames later. A request for 1.8x zoom (1706) can be made based on frame 4, and the corresponding zoom adjustment (shown as a zoom amount of 1.8) will be applied at frame 11, seven frames later.

[0276] Several advantages are achieved by using the frame cropping and scaling techniques described above. For example, cropping and scaling techniques enable the provision of fixed-size features to target objects in video systems (e.g., mobile devices, video analytics systems, etc.). Systems implementing cropping and scaling techniques can achieve good performance and can be deployed in any type of device, such as mobile devices, IP cameras, etc.

[0277] Figure 18This is a flowchart illustrating an example of a process 1800 for processing one or more frames using the techniques described herein. At block 1802, process 1800 includes: determining a region of interest (ROI) in a first frame of a frame sequence. The ROI in the first frame includes an object of a certain size in the first frame. As described above, the ROI may be determined based on user input or may be determined automatically. In some examples, process 1800 includes: receiving user input corresponding to a selection of an object in the first frame, and determining the ROI in the first frame based on the received user input. In some aspects, the user input includes: touch input provided using a device's touch interface (e.g., selecting an object, drawing a shape around an object, etc.). As described herein, user input may include other types of user input.

[0278] At box 1804, process 1800 includes: cropping a portion of a second frame in the frame sequence, the second frame appearing after the first frame in the frame sequence. At box 1806, process 1800 includes: scaling that portion of the second frame based on the size of an object in the first frame. For example, scaling that portion of the second frame based on the size of an object in the first frame results in the object in the second frame having the same size as the object in the first frame. In some examples, cropping and scaling the portion of the second frame keeps the object centered in the second frame. In some cases, process 1800 includes: detecting and tracking objects in one or more frames of the frame sequence.

[0279] In some examples, process 1800 includes: determining points of an object region defined for an object in the second frame, and cropping and scaling that portion of the second frame if the points of the object region are located at the center of the cropped and scaled portion. In some cases, the points of the object region are the center points of the object region. In some cases, the object region is a bounding box (or other bounded region). In some cases, the center point is the center point of the bounding box (or other region). In some cases, the center point is the center point of the object (e.g., the centroid or center point of the object). This can be achieved by performing object segmentation (e.g., using...). Figure 6A The system 600 shown in the figure is used to find the center point.

[0280] In some aspects, process 1800 includes: determining a first length associated with an object in a first frame, and determining a second length associated with an object in a second frame. Process 1800 may include: determining a scaling factor based on a comparison between the first length and the second length, and scaling the portion of the second frame based on the scaling factor. In some cases, scaling the portion of the second frame based on the scaling factor results in a second object region in the cropped and scaled portion having the same size as the first object region in the first frame. In some examples, the first length is the length of the first object region determined for the object in the first frame, and the second length is the length of the second object region determined for the object in the second frame. In some cases, the first object region is a first bounding box (or other bounding region), and the second object region is a second bounding box (or other bounding region). The first length may be the diagonal length (or other length) of the first bounding box, and the second length may be the diagonal length (or other length) of the second bounding box. In some cases, the first length may be the length between points of the object in the first frame, and the second length may be the length between multiple points of the object in the second frame.

[0281] In some aspects, process 1800 includes: determining points of a first object region generated for an object in a first frame, and determining points of a second object region generated for an object in a second frame. In some implementations, the points of the first object region are the center points of the first object region (e.g., the center point of the object in the first frame, or the center point of a first bounding box), and the points of the second object region are the center points of the second object region (e.g., the center point of the object in the second frame, or the center point of a second bounding box). Process 1800 may include: determining a motion factor for the object based on a smoothing function using the points of the first and second object regions. The smoothing function can control the positional variation of the object across multiple frames in a frame sequence. For example, the smoothing function can control the positional variation of the object so that the position of the object changes gradually across multiple frames in the frame sequence (e.g., such that the change does not exceed a threshold positional variation, such as 5 pixels, 10 pixels, or other threshold positional variations). In some examples, the smoothing function includes a moving function (e.g., a moving average function or other moving function) used to determine the location of a point of the corresponding object region in each of multiple frames of a frame sequence based on a statistical measure of object movement (e.g., mean, standard deviation, variance, or other statistical measure). In one illustrative example, the smoothing function includes a moving average function used to determine the average location of a point of the corresponding object region in each of multiple frames. For example, as described above, a moving average can reduce or eliminate false alarms (e.g., by minimizing the location of the object bounding box in error detection / tracking). Process 1800 may include cropping this portion of a second frame based on a movement factor.

[0282] In some examples, process 1800 includes: determining a first length associated with an object in a first frame, and determining a second length associated with an object in a second frame. In some examples, the first length is the length of a first bounding box generated for the object in the first frame, and the second length is the length of a second bounding box generated for the object in the second frame. In some cases, the first length is the diagonal length of the first bounding box, and the second length is the diagonal length of the second bounding box. Process 1800 may include: determining a scaling factor for the object based on a comparison between the first and second lengths and based on a smoothing function using the first and second lengths. The smoothing function can control the size variation of the object across multiple frames in a frame sequence. For example, the smoothing function can control the size variation of the object so that the size of the object varies gradually across multiple frames in the frame sequence (e.g., such that the variation does not exceed a threshold size variation, e.g., according to a threshold size variation not greater than 5%, 10%, 20%, or other threshold size variations). In some cases, the smoothing function includes a moving function (e.g., a moving average function, or other moving function) used to determine the length associated with an object in each of multiple frames of a frame sequence based on a statistical measure of object size (e.g., mean, standard deviation, variance, or other statistical measure). In an illustrative example, the smoothing function includes a moving average function used to determine the average length associated with an object in each of multiple frames. For example, as described above, a moving average can reduce or eliminate false alarms (e.g., by minimizing the diagonal length of the object bounding box in error detection / tracking or the length of error detection / tracking between multiple parts of an object). Process 1800 may include scaling the portion of the second frame based on a scaling factor. In some aspects, scaling the portion of the second frame based on a scaling factor results in the second bounding box in the cropped and scaled portion having the same size as the first bounding box in the first frame.

[0283] Figure 19 , Figure 20 , Figure 21 , Figure 22 , Figure 23 The simulation was shown on four video clips. The video clips all contained resolutions of 720p and 1080p, and were all captured at a rate of 30 frames per second (fps). Figures 19 to 23 Each example in the examples is an illustrative example of a magnification effect (where frames are cropped and upsampled or magnified, similar to...). Figure 10A and Figure 10B (Example). Figure 19As shown, "Initial Frame 0" is the first frame from the video, and "Initial Frame X" is the current frame during video recording. To achieve the zoom effect, the frame cropping and scaling system 800 will crop the area from Initial Frame X and then upsample that area to the size of the initial frame.

[0284] As described above, a device may include multiple cameras and / or lenses (e.g., two cameras in a dual-camera lens system) for performing one or more dual-camera mode features. For example, the device's dual-camera lenses (e.g., a mobile phone or smartphone including a rear dual-camera lens or other dual-camera lenses) can be used to record multiple videos simultaneously (e.g., two videos), which may be referred to as the "dual-camera video recording" feature. In some cases, the primary camera lens (e.g., a telephoto lens) of the device's dual-camera lenses can capture (and / or record) a first video, while the secondary camera lens (e.g., a zoom lens, such as a wide-angle lens) can capture (and / or record) a second video. In some cases, the second video can be used to perform the frame cropping and scaling techniques described above to maintain a fixed size of the target object during video recording. In some cases (e.g., processing images output by an ISP and before display or storage), this solution can be used as a video post-processing feature.

[0285] In some cases, dual-camera mode features can be achieved by using two camera lenses of a device simultaneously, such as the device's main camera lens (e.g., a telephoto lens) and a secondary camera lens (e.g., a zoom lens). The dual-camera video recording feature mentioned above allows two camera lenses to record two videos simultaneously. For example, a device can use a wide-angle lens and a standard lens to record separate videos. In some cases, a device can use three, four, or even more camera lenses to record video simultaneously. The videos can then be displayed (e.g., simultaneously), stored, sent to another device, and / or used in other ways. For example, using dual-camera mode features (e.g., dual-camera video recording), a device can display two perspectives of a scene on a monitor at the same time (e.g., split-screen video). Dual-camera mode features offer various advantages, such as allowing the device to capture a wide field of view of a scene (e.g., with more background and surrounding objects in the scene), allowing the device to capture panoramic views of large-scale events or scenes, etc.

[0286] In some cases, multiple camera modules and lenses can be used to perform zoom functionality. For example, a secondary camera lens can be set to a higher zoom level (e.g., 2.0x camera zoom) compared to the main camera and / or lens (e.g., which may have a 1.0x camera zoom rate).

[0287] Various issues can arise regarding maintaining a fixed size of a target object within a frame sequence. In one example, the device may fail to perform a zoom-out effect when the target object moves toward the device's camera. This problem could be due to limitations in the field of view from the initial video frame. For instance, there might not be enough space in the frame to zoom out and maintain the size of the target object (e.g., resulting in black spaces around the zoomed-out frame). In another example, when the target object moves away from the device's camera, the magnified image generated based on the initial video frame may suffer from poor quality, such as blurring, including one or more visual artifacts, lack of sharpness, etc. Furthermore, devices implementing dual-camera mode features do not incorporate any artificial intelligence technology. Such systems require end users to manually edit the images using video editing tools or software applications.

[0288] As described above, this paper describes systems and techniques for switching between multiple cameras or lenses in a device capable of implementing one or more of the dual-camera mode features described above. Although the systems and techniques are described herein with reference to a dual-camera system or two-camera system, they can be applied to systems with more than two cameras (e.g., when using three, four, or other numbers of cameras to capture images or videos). In some cases, the systems and techniques described herein can use camera lens switching algorithms in dual-camera systems to maintain a fixed size of the target object within a sequence of video frames captured using the dual-camera system. In some examples, the systems and techniques can perform dual-camera zoom, which can be used to provide more detailed object zoom effects.

[0289] As described above, an object detection and tracking system (e.g., object detection and tracking system 806) can detect and / or track objects in one or more frames. The object detection and tracking system can utilize any suitable object detection and tracking technique for the multi-camera (e.g., dual-camera) implementations described herein, such as those mentioned above. In some cases, as described above, regions of interest (ROIs) or target objects can be identified based on user input or automatically.

[0290] In some examples, object detection and tracking systems can detect and / or track objects in frames by performing object matching for dual-camera video analysis using machine learning object detection and tracking systems. For example, an object detection and tracking system can extract points of interest (POIs) from one or more input frames. POIs can include stable and repeatable two-dimensional (2D) locations in a frame that are stable across different lighting conditions and viewpoints. POIs can also be referred to as keypoints or landmarks (e.g., facial landmarks on a face). An example of a machine learning system is a convolutional neural network (CNN). In some cases, CNNs may outperform hand-designed representations for various tasks that use frames or images as input. For example, CNNs can be used to predict 2D keypoints or landmarks for various tasks such as object detection and / or tracking.

[0291] Figure 24 This is a diagram illustrating an example of a machine learning-based object detection and tracking system 2400. In some cases, system 2400 uses self-training for self-supervision (rather than using human supervision to define points of interest in the actual training images). Object tracking is performed through point correspondences with point feature matching. In some cases, a large dataset of pseudo-ground truth locations of points of interest in real images or frames is used; this dataset can be pre-configured or pre-set by system 2400 itself, rather than requiring large-scale manual annotation work.

[0292] System 2400 includes a fully convolutional neural network architecture. In some examples, System 2400 receives one or more full-size images as input and operates on them. For example, as Figure 24 As shown, image pairs including images 2402 and 2404 can be input into the system (e.g., during training and / or during inference). In some cases, system 2400 (using a full-size image as input) generates point of interest (POI) detections accompanied by fixed-length descriptors in a single feedthrough. The neural network model of system 2400 includes a single shared encoder 2406 (shown as having four convolutional layers, but may include more or fewer layers) to process and reduce the dimensionality of the input image. Following the encoder, the neural network architecture is split into two decoder heads that learn task-specific weights. For example, a first decoder head 2408 is trained for POI detection, and a second decoder head 2410 is trained for generating POI descriptions (referred to as descriptors). The task of finding POIs may include detection and description (e.g., performed by decoder heads 2408 and 2410, respectively). Detection is the localization of POIs in an image or frame, and description describes each detected point (e.g., with a vector). The overall goal of system 2400 is to efficiently and effectively find features and stable visual characteristics.

[0293] In some examples, the system 2400 can bend each region of a pixel in the input image (e.g., each 8×8 pixel region). This region can be considered as a pixel after bending, in which case each region of the pixel can be represented by a specific pixel in a feature map with 64 channels, followed by a trash can channel. If no points of interest (e.g., keypoints) are detected in a particular 8×8 region, the trash can can have high activation. If keypoints are detected in the 8×8 region, the other 64 channels can be processed using a softmax architecture to find the keypoints within the 8×8 region. In some cases, the system 2400 can compute 2D point of interest locations and descriptors in a single feedthrough and can run on a 480×640 image at 70 frames per second (fps) using a Titan X graphics processing unit (GPU).

[0294] Figure 25 This is a flowchart illustrating an example of a camera lens switching pipeline 2500. Pipeline 2500 is an example of dual-camera lens switching logic. Figure 25 The first lens referred to here is the lens and / or camera used by the device (e.g., based on user input) as the main lens for capturing video. In some cases, the first lens may be a telephoto lens, while... Figure 25 The second lens referred to here can be a wide-angle lens. In some cases, the first lens can be a wide-angle lens, while the second lens can be a telephoto lens. Any other type of lens can be used for both the first and second lenses.

[0295] At box 2502, the target fixation feature can begin with a first lens (e.g., a telephoto lens if the user selects it as the primary lens for recording video). When certain conditions are met (as described below), boxes 2504 and 2508 of pipeline 2500 can switch the primary lens from the first lens (e.g., a telephoto lens) to a second lens (e.g., a wide-angle lens) to perform the target fixation function. In this case, the second lens can be used to capture one or more primary video frames. When certain conditions are met (as described below), boxes 2506 and 2510 of pipeline 2500 can switch the primary lens back from the second lens (e.g., a wide-angle lens) to the first lens (e.g., a telephoto lens) to perform the target fixation feature. In this case, the first lens can be used to capture any primary video frame.

[0296] An example of an algorithm (referred to as Algorithm 1A) that can be used to perform camera lens switching is as follows (using a telephoto lens as the first lens and a wide-angle lens as the second lens):

[0297] The `disp_xy` function is initialized based on the x and y displacements of the target object's bounding box center from the center point of the first or initial frame.

[0298] Initialize done_by_tele to true and tele_lens_ratio to 1.0.

[0299] Initialize the camera zooming_ratio values ​​for telephoto and wide-angle lenses.

[0300] When done_by_tele is true (e.g., assigned a value of 1), the telephoto lens is used to perform target fixed-size features. zooming_ratio is the zoom (or zoom ratio) mentioned above and is used to determine how much the ROI or object from the input frame is magnified.

[0301] In some cases, the camera lens switching algorithm described above can continue with the following operation (referred to as Algorithm 1B):

[0302] For each iteration of the input video frame

[0303] #Option 1) First, the first video frame, starting from the telephoto lens.

[0304] #Option 2) Keep using a telephoto lens

[0305] #Option 3) Switch from wide-angle lens to telephoto lens

[0306] If tele_zooming_ratio is 1.0

[0307] Readjust the width and height of the object's bounding box based on tele_zooming_ratio.

[0308] Reposition the object

[0309] If the object's bbox width or height is outside the image

[0310] # Switch to wide-angle lens

[0311] done_by_tele == fake

[0312] jump over

[0313] Processing video frame cropping and resizing

[0314] Update disp_xy offset

[0315] done_by_tele = true

[0316] Set wide_lens_tirnes_ratio = 1.0

[0317] otherwise

[0318] done_by_tele = fake

[0319] #Option 1) Keep using a wide-angle lens

[0320] #Option 2) Switch from telephoto lens to wide-angle lens

[0321] If done_by_tele == false

[0322] If previous iterations were accomplished using a telephoto lens

[0323] Update wide-angle lens magnification

[0324] if

[0325] If disp_xy! = 0 and

[0326] Update disp_xy offset

[0327] Processing video frame cropping and resizing

[0328] otherwise

[0329] Preserve the original video frames without cropping or resizing.

[0330] Figure 26 This is a flowchart illustrating an example of a camera lens switching process 2600. At box 2602, for a main video including frames captured using a first camera lens, process 2600 can perform video frame selection from the video captured using the first lens (e.g., a telephoto camera lens). For example, a user can select a video frame as a first frame, which will be used as the starting point for performing target fixed-size features. For example, as described above, the first frame can be used to define the ROI and / or target object size, points (e.g., the center point of a bounding box), and distances (e.g., the diagonal length of the bounding box). At box 2603, process 2600 can use a second camera lens (e.g., a wide-angle camera lens) to determine or locate the corresponding video frame from the captured or recorded video. In an illustrative example, video frames from the first and second camera lenses can have reference numerals that correspond to the output time of those frames. Process 2600 (at box 2603) can use the reference numerals of the frames to determine the corresponding video frame. The first camera lens in... Figure 26 The image shows a telephoto lens (also referred to as a "telephoto lens" in this article), while the second lens is... Figure 26 The image shown is of a wide-angle camera lens. However, those skilled in the art will understand that the first and second lenses can be any other suitable lenses.

[0331] At box 2604, process 2600 includes selecting and / or drawing a bounding box (or other boundary region) for a target object in the first video frame. For example, a user can select a target object (e.g., in some cases a single object or multiple objects) from the first video frame by providing user input (e.g., clicking on a target object in a frame displayed on a touchscreen display, drawing a bounding box around the target object, or providing any other suitable type of input). In another example, the techniques described above can be used to automatically determine the target object (e.g., via object detection and tracking system 806).

[0332] At box 2605, process 2600 identifies or finds the same target object in the corresponding video frame of the video captured using the second lens as determined in box 2603. In some cases, in order to find the same target from the video captured using the second lens, process 2600 can (e.g., using object detection and tracking system 806) determine the approximate location of the target object from the video captured using the first lens. Process 2600 can then apply an object matching algorithm (e.g., using data from...) Figure 24 The system 2400) is used to locate target objects in video captured using a second lens, which can be associated with bounding boxes and information.

[0333] At box 2606, process 2600 can perform object detection and tracing. In some cases, object detection and tracing can be similar to the above reference. Figures 8B to 13B The description describes object detection and tracking. For example, object detection and tracking system 806 can automatically detect and track objects in two videos (video captured by a first shot and video captured by a second shot) in parallel. At box 2608, process 2600 (e.g., by object detection and tracking system 806) determines or captures the coordinates and distances (e.g., diagonal lengths) of points (e.g., center points) of bounding boxes determined for target objects across frames in both videos. In some cases, the points (e.g., center points) and distances (e.g., diagonal lengths) can be stored for later use by process 2600.

[0334] At box 2610, process 2600 applies a smoothing function. For example, as described above, smoothing engine 812 may apply a smoothing function to smooth the frame scaling (or resizing) rate. The scaling rate or zoom rate can be calculated by comparing the diagonal length (or other distance) of the bounding box of the target object in the selected first video frame with the diagonal length (or other distance) of the bounding box of the target object in the current frame. As described above, in some cases, the smoothing function may include a moving average function. For example, the smoothing function may be used to determine the average length associated with an object in each of multiple frames in a frame sequence.

[0335] At box 2612, process 2600 can determine whether a camera lens switch should be performed. For example, process 2600 can use the camera lens switch algorithms provided above (e.g., Algorithm 1A and / or Algorithm 1B) to determine whether to use video frames from a first lens (e.g., a telephoto lens) or a second lens (e.g., a wide-angle lens). At box 2614, process 2600 can perform frame cropping and scaling (or zooming). For example, frame scaling engine 810 can upsample (or enlarge) the ROI (e.g., bounding box) of the target object based on the object bounding box point coordinates (e.g., center point coordinates) and the scaling rate or resizing rate. At box 2616, process 2600 performs video stabilization, such as using the references above. Figure 12 The image stabilization technique described. At box 2618, process 2600 outputs a frame that has been cropped and scaled from the initial capture frame, thereby maintaining the target object at the size it was in the initial or first frame.

[0336] In some cases, as described above, the camera lens switching system and techniques described herein can be applied to or extended to other multi-camera systems (e.g., camera systems including three, four, or five cameras) that record multiple images and / or videos at once.

[0337] In some examples, the step-size algorithm can be used to achieve a smoothing effect. In some cases, the techniques described above that use step-size values ​​can be used (e.g., as referenced). Figure 13B (As described). An illustrative example of the moving step algorithm is provided below:

[0338] (1) Operation 1 In the output frame, initialize the target object's coordinates center_xy to (w / 2, h / 2), where w is the width of the output frame and h is the height of the output frame.

[0339] (2) Operation 2 When the lens is held as a telephoto lens (e.g., as described below) Figure 32 and Figure 33 (as shown) or the lens is kept as a wide-angle lens (e.g., as described) Figure 34 , Figure 35 and Figure 36 When (as shown), update center_xy

[0340] (3) Operation 3When the situation of operation 2 changes (switching from telephoto lens to wide-angle lens, or from wide-angle lens to telephoto lens), the target object coordinates are updated to (center_xy [1] ± moving_step, center_xy [2] ± moving_step), and the target object coordinates are applied to the output frame.

[0341] (4) Operation 4 Starting from operation 3, update center_xy using move_step to get closer to (w / 2, h / 2).

[0342] (5) Operation 5 Repeat steps 3 and 4 until center_xy = (w / 2, h / 2).

[0343] Figures 27 to 36 This is a diagram illustrating an example of using the camera lens switching technology described above. It uses an example of a telephoto camera lens (shown as "telephoto frame") as the main lens (e.g., selected by the user) and a wide-angle lens (shown as "wide-angle frame") as the secondary lens to describe... Figures 27 to 36 Examples.

[0344] Figure 27 This is a diagram illustrating an example of lens selection. For instance, when the size of the target object in the current telephoto frame 2704 (shown as telephoto frame N) is smaller than the size of the target object in the reference telephoto frame 2702, the device or system can determine (e.g., in...) Figure 26 (At frame 2612) use telephoto lens frame 2704 to generate the output frame result. Figure 28 This is another diagram illustrating an example of lens selection. For instance, when the size of the target object in the current telephoto frame 2804 (shown as telephoto frame M) is larger than the size of the object in the reference telephoto frame 2802, the device or system can determine to use the wide-angle lens frame 2806 (shown as wide-angle frame M to indicate that the wide-angle frame M and the telephoto frame M are captured simultaneously from the same angle relative to the camera) to generate the output frame result. Figure 29 This is another diagram illustrating an example of lens selection. For instance, if the size of a target object in the current wide-angle frame 2904 (shown as wide-angle frame P) is larger than the size of an object in the reference wide-angle frame 2902, the device or system can determine to use the current wide-angle frame 2904 to generate the output frame result.

[0345] Figure 30This is a diagram illustrating an example of switching from a telephoto lens to a wide-angle lens. For example, if output frame N is generated from telephoto frame N, and the size of the target object in the current telephoto frame 3004 (shown as telephoto frame N + 1) is larger than the size of the object in the reference telephoto frame 3002, the device or system can switch to wide-angle frame 3006 (shown as wide-angle frame N + 1) to generate output frame 3008.

[0346] Figure 31 This is another diagram illustrating an example of switching from a telephoto lens to a wide-angle lens. For example, if output frame N is generated from telephoto frame N, and the target's position in the current telephoto frame 3104 (shown as telephoto frame N + 1) is near the frame boundary (e.g., in this case, the object is not centered in the frame after zooming or scaling), the device or system can switch from the telephoto frame to the wide-angle frame to generate output frame 3108. In some cases, the device or system can determine whether an object is near a frame boundary by determining whether a point of the target object (e.g., the center point of the object's bounding box) is within a threshold distance of the boundary, such as within 10 pixels, within 20 pixels, or at other distances. Even if the size of the target object in the current telephoto frame 3104 is smaller than the size of the target object in the reference telephoto frame 3102, a switch from the current telephoto frame 3104 (captured using a telephoto lens) to the wide-angle frame 3106 (captured using a wide-angle lens) can still be performed when the object is near the boundary.

[0347] Figure 32 This is a diagram illustrating an example of switching from a wide-angle lens to a telephoto lens. For example, if output frame N is generated from wide-angle frame N, and if the size of the target object in the current wide-angle frame 3206 (shown as wide-angle frame N+1) is smaller than the size of that object in the reference telephoto frame 3202, and the position of the target object is within the image boundary after magnification, then the device or system can switch from the current wide-angle frame 3206 to the current telephoto frame 3204 (shown as telephoto frame N+1) to generate output frame 3208.

[0348] Refer again Figure 32 This provides an example of continuing to use a telephoto lens. For instance, if output frame N is generated from telephoto frame N, and if the size of the target object in the current telephoto frame 3204 (telephoto frame N+1) is smaller than the size of the target object in the reference telephoto frame 3202, and the position of the target object is within the image boundary after magnification, then the device or system can continue to use telephoto frame 3204 to generate output frame 3208.

[0349] Figure 33This is a diagram illustrating another example of using a telephoto lens. For example, if the size of a target object in the current telephoto frame 3304 (telephoto frame N) is smaller than the size of that target object in the reference telephoto frame 3302, and starting from the current telephoto frame 3304 (shown as telephoto frame N) 3302, the object's position is near the frame boundary (e.g., the center point of the target object or the bounding box of the target object is within a threshold distance of the boundary), and no camera lens switch has occurred within a threshold time period (e.g., no camera lens switch has occurred within a certain number of frames, a certain amount of time, and / or other time periods), then the device or system can continue to use telephoto frame 3304 to generate output frame 3308.

[0350] Figure 34 This is a diagram illustrating an example of using a wide-angle lens. For example, if the size of the target object in the current wide-angle frame 3408 (shown as wide-angle frame N) is smaller than the size of that target object in the reference wide-angle frame 3404 and the size of the object in the current telephoto frame 3406 (shown as telephoto frame N) is larger than the size of the object in the reference telephoto frame 3402, then the device or system can continue to use the current wide-angle frame 3408 to generate the output frame 3410.

[0351] Figure 35 This is a diagram illustrating another example of using a wide-angle lens. For example, if the size of the target object in the current wide-angle frame 3506 (shown as wide-angle frame M) is larger than the size of the target object in the reference wide-angle frame 3502, the device or system can continue to use the current wide-angle frame 3506 to generate the output frame 3510.

[0352] Figure 36 This diagram illustrates another example of using a wide-angle lens. For instance, if output frame N is generated from wide-angle frame N, and the target object in the current telephoto frame 3604 (shown as telephoto frame N+1) is located near the frame boundary (e.g., the center point of the target object or the bounding box of the target object is within a threshold distance of the boundary), in this case, the frame may not be able to be scaled or zoomed to obtain the output frame. The device or system can continue to use the current wide-angle frame 3606 (shown as wide-angle frame N+1) to generate output frame 3608. Even if the size of the target object in the current telephoto frame 3604 is smaller than the size of the target object in the reference telephoto frame 3602, the device or system can continue to use the current wide-angle frame 3606 to generate output frame 3608 when the object is near the boundary.

[0353] Figures 37 to 41These are simulated images illustrating the use of the camera lens switching system and techniques described herein. For example, to simulate simultaneous dual-camera video recording, two rear cameras from a mobile phone (e.g., a smartphone), including a telephoto lens and a wide-angle lens, are used. The start and end points of the dual-recorded videos are manually aligned from the dual camera lenses. The test sample video used in the simulation results has a 1080p resolution, with 30 frames per second (fps). As described above, the end user can select a target object from the telephoto lens video frames (e.g., using a touchscreen displaying the video frames).

[0354] Figure 37 The starting or initial video frame from the telephoto lens is shown. Figure 37 (left side) and the starting or initial video frame from the wide-angle lens ( Figure 37 (on the right side). Figure 38 The final video frame from the telephoto lens is shown. Figure 38 (left side) and the ending video frame from the wide-angle lens ( Figure 38 (on the right side). Figure 39 This shows the target fixed-size feature applied to a telephoto lens video frame at time point n. Figure 39 (on the left), and after switching from a telephoto lens to a wide-angle lens, at time point n+1, the target fixed-size feature applied to the wide-angle lens video frame ( Figure 39 (on the right side). Figure 40 It shows that at time point m ( Figure 40 The target fixed size feature applied to the wide-angle lens video frame (on the left) and the target fixed size feature applied to the telephoto lens video frame at time point m+1 after switching from the wide-angle lens to the telephoto lens (on the left) Figure 40 (on the right side). Figure 41 This shows the fixed-size feature of the target applied to a telephoto lens video frame at time point p. Figure 41 (on the left side), and after switching from the telephoto lens to the wide-angle lens, the target fixed-size feature applied to the wide-angle lens video frame at time point p+1 ( Figure 41 (on the right side).

[0355] The lens switching system and techniques described in this paper offer various advantages. For example, the lens switching system and techniques enable the aforementioned target localization size features to be used in multi-recording video scenarios (e.g., dual-recording video using two camera lenses) while obtaining high-quality results.

[0356] In some examples, the processes described herein (e.g., processes 820, 930, 1200, 1300, 1310, 1800, 2500, 2600, and / or other processes described herein) can be executed by a computing device or apparatus. In one example, one or more processes can be executed by... Figure 1 The image capture and processing system 100 performs this process. In another example, one or more processes can be performed by... Figure 8B The frame cropping and scaling system 800 performs this. In another example, one or more processes can be performed by... Figure 47 The computing system 4700 shown is used to perform this. For example, it has... Figure 47 The computing device of the computing system 4700 shown may include components of the frame cropping and scaling system 800, and can implement... Figure 8C Process 820 Figure 9A Process 930 Figure 9B Process 935 Figure 13A Process 1300 Figure 13B Process 1310 Figure 18 The operation of process 1800 and / or other processes described herein.

[0357] Computing devices may include any suitable device, such as mobile devices (e.g., mobile phones), desktop computing devices, tablet computing devices, wearable devices (e.g., VR headsets, AR headsets, AR glasses, network-connected watches or smartwatches, or other wearable devices), server computers, autonomous vehicles, or computing devices of autonomous vehicles, robotic devices, televisions, and / or any other computing device with the resource capability to perform the processes described herein (including processes 820, 930, 935, 1800, and / or other processes described herein). In some cases, a computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or one or more other components configured to perform the steps of the processes described herein. In some examples, a computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or one or more other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0358] Components of a computing device can be implemented in a circuit system. For example, components may include and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof, and / or may be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein.

[0359] Processes 820, 930, 1200, 1300, 1310, 1800, 2500, and 2600 are shown as logic flowcharts, representing a series of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operation represents a computer-executable instruction stored on one or more computer-readable storage media, which performs the described operation when one or more processors execute the computer-executable instruction. Typically, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of operations can be combined in any order and / or in parallel to implement the process.

[0360] Additionally, processes 820, 930, 1200, 1300, 1310, 1800, 2500, 2600, and / or other processes described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as program code (e.g., executable instructions, one or more computer programs, or one or more applications) that can be jointly executed on one or more processors via hardware or a combination thereof. As described above, the program code may be stored, for example, on a computer-readable or machine-readable storage medium in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0361] As described above, various aspects of this disclosure can utilize machine learning systems, such as object tracking and object classification. Figure 42This is an illustrative example of a deep learning neural network 4200 capable of implementing the aforementioned machine learning-based object tracking and / or classification. Input layer 4220 includes input data. In one illustrative example, input layer 4220 may contain data representing pixels of an input video frame. Neural network 4200 includes multiple hidden layers 4222a, 4222b to 4222n. Hidden layers 4222a, 4222b to 4222n comprise "n" hidden layers, where "n" is an integer greater than or equal to 1. The number of hidden layers can be as many as necessary for a given application. Neural network 4200 also includes an output layer 4224, which provides the output produced by the process performed by hidden layers 4222a, 4222b to 4222n. In one illustrative example, output layer 4224 can provide a classification of objects in the input video frame. This classification may include categories that identify the type of object (e.g., person, dog, cat, or other object).

[0362] Neural network 4200 is a multi-layer neural network with interconnected nodes. Each node can represent a piece of information. The information associated with a node is shared between different layers, and each layer retains the information as it is processed. In some cases, neural network 4200 may include a feedforward network, in which case there are no feedback connections, and the network's output is fed back into itself. In some cases, neural network 4200 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.

[0363] Information can be exchanged between multiple nodes through node-to-node interconnections between layers. Nodes in input layer 4220 can activate the set of nodes in the first hidden layer 4222a. For example, as shown, each input node of input layer 4220 is connected to each node of the first hidden layer 4222a. Nodes in the first hidden layer 4222a can transform the information of each input node by applying an activation function to the input node information. The information derived from the transformation can then be passed to and can activate nodes in the next hidden layer 4222b, which can perform its own specified function. Exemplary functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 4222b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 4222n can activate one or more nodes in output layer 4224, providing the output at those nodes. In some cases, although a node in neural network 4200 (e.g., node 4226) is shown as having multiple output lines, a node has a single output, and all lines shown as outputs from a single node represent the same output value.

[0364] In some cases, each node or the interconnections between nodes can have weights, which are a set of parameters derived from the training of the neural network 4200. Once the neural network 4200 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, the interconnections between nodes can represent learned information about the interconnected nodes. The interconnections can have tunable numerical weights (e.g., based on the training data set), allowing the neural network 4200 to adapt to the input and learn as it processes more and more data.

[0365] The neural network 4200 is pre-trained to process features from the data in the input layer 4220 using different hidden layers 4222a, 4222b to 4222n, in order to provide an output through the output layer 4224. In an example where the neural network 4200 is used to identify objects in an image, the neural network 4200 can be trained using training data that includes both images and labels. For example, training images can be input into the network, where each training image has a label that indicates the category of one or more objects in each image (basically, indicating to the network what the objects are and what features they have). In an illustrative example, the training images could include images of the number 2, in which case the label for that image could be [0 0 1 0 0 0 0 0 0 0].

[0366] In some cases, the neural network 4200 can use a training process called backpropagation to adjust the weights of its nodes. Backpropagation can include forward path, loss function, back path, and weight updates. For each training iteration, forward path, loss function, back path, and parameter updates are performed. This process can be repeated for a certain number of iterations for each training image set until the neural network 4200 is trained well enough to precisely adjust the weights of each layer.

[0367] For an example of recognizing objects in an image, the forward path may include passing a training image through a neural network 4200. The weights are initially randomized before training the neural network 4200. The image may include, for example, a numerical array representing the pixels of the image. Each number in the array may include a value from 0 to 255, describing the pixel intensity at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and three color components (such as red, green, and blue, or a luminance component and two chrominance components).

[0368] For the first training iteration of the neural network 4200, the output will likely include values ​​that do not give preference to any particular category due to the weights being randomly selected during initialization. For example, if the output is a vector with probabilities that an object includes different categories, the probability values ​​for each of the different categories can be equal or at least very similar (e.g., for 10 possible categories, each category could have a probability value of 0.1). With the initial weights, the neural network 4200 cannot determine low-level features and therefore cannot accurately determine what the object's classification could be. A loss function can be used to analyze errors in the output. Any suitable loss function can be defined. An example of a loss function is mean squared error (MSE). MSE is defined as... It calculates the sum of the actual answer minus half the square of the predicted (output) answer. The loss can be set to equal... The value of .

[0369] For the first training image, the loss (or error) will be high because the actual value will differ significantly from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training label. The Neural Network 4200 can perform a backward path by determining which inputs (weights) have the greatest impact on the network's loss and can adjust the weights to reduce and ultimately minimize the loss.

[0370] The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that have the greatest impact on the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated to change in the opposite direction of the gradient. A weight update can be expressed as... Where w represents the weight, w i η represents the initial weights, and η represents the learning rate. The learning rate can be set to any suitable value, where a high learning rate involves larger weight updates, while a lower value indicates smaller weight updates.

[0371] Neural Network 4200 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which consists of an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN consist of a series of convolutional, non-linear, pooling (for downsampling), and fully connected layers. Neural Network 4200 can also include any other deep network different from a CNN, such as an autoencoder, a Deep Belief Network (DBN), a Recurrent Neural Network (RNN), and so on.

[0372] Figure 43This is an illustrative example of a Convolutional Neural Network (CNN) 4300. The input layer 4320 of the CNN 4300 includes data representing an image. For example, the data could include a numerical array representing the pixels of an image, where each number in the array includes a value from 0 to 255 describing the pixel intensity at its location in the array. Using the previous example above, the array could include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or a luminance component and two chrominance components, etc.). The image can be passed through a convolutional hidden layer 4322a, an optional non-linear activation layer, a pooling hidden layer 4322b, and a fully connected hidden layer 4322c to obtain an output at the output layer 4324. Although in Figure 43 The diagram shows only one hidden layer among the various hidden layers, but those skilled in the art will understand that a CNN 4300 may include multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers. As previously described, the output may indicate a single category of object, or may include the probability that best describes the category of an object in the image.

[0373] The first layer of CNN 4300 is a convolutional hidden layer 4322a. Convolutional hidden layer 4322a analyzes the image data input to layer 4320. Each node in convolutional hidden layer 4322a is connected to a node (pixel) region of the input image called a receptive field. Convolutional hidden layer 4322a can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in convolutional hidden layer 4322a. For example, the region of the input image covered by a filter at each convolutional iteration will be the receptive field for that filter. In an illustrative example, if the input image comprises a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in convolutional hidden layer 4322a. Weights are learned for each connection between a node and its receptive field, and in some cases, an overall bias is learned, such that each node learns to analyze its specific local receptive field in the input image. Each node in hidden layer 4322a will have the same weights and biases (called shared weights and shared biases). For example, the filter has a weight (digital) array and the same depth as the input. For the video frame example, the filter would have a depth of 3 (based on the 3 color components of the input image). An illustrative example of the filter array size is 5×5×3, corresponding to the size of the receptive field of a node.

[0374] The convolutional property of the convolutional hidden layer 4322a is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 4322a can start at the top left corner of the input image array and can be convolved around the input image. As described above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 4322a. At each convolutional iteration, the value of the filter is multiplied by the corresponding number of initial pixel values ​​of the image (e.g., a 5×5 filter array multiplied by 5×5 input pixel values ​​at the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain a sum for that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 4322a. For example, the filter can move to the next receptive field by a stride amount (called stride). The stride can be set to 1 or other suitable amounts. For example, if the stride is set to 1, the filter will move 1 pixel to the right at each convolutional iteration. Processing the filter at each unique location of the input produces a number that represents the filter result for that location, thus determining the sum value for each node of the convolutional hidden layer 4322a.

[0375] The mapping from the input layer to the convolutional hidden layer 4322a is called an activation map (or feature map). An activation map includes values ​​for each node, representing the filter result at each location in the input quantity. An activation map can include an array containing various summed values ​​produced by each iteration of the filter on the input quantity. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 4322a can include multiple activation maps to identify multiple features in the image. Figure 43 The example shown includes three activation maps. Using the three activation maps, the convolutional hidden layer 4322a can detect three different types of features, and each feature can be detected across the entire image.

[0376] In some examples, nonlinear hidden layers can be applied after convolutional hidden layer 4322a. Nonlinear layers can be used to introduce nonlinearity into a system that has already computed linear operations. An illustrative example of a nonlinear layer is a tuned linear unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0, x) to all values ​​in the input, which turns all negative activations to 0. Therefore, ReLU can increase the nonlinearity of network 4300 without affecting the receptive field of convolutional hidden layer 4322a.

[0377] Pooling hidden layer 4322b can be applied after convolutional hidden layer 4322a (and, when used, after nonlinear hidden layers). Pooling hidden layer 4322b is used to simplify the information in the output of convolutional hidden layer 4322a. For example, pooling hidden layer 4322b can obtain each activation map from the output of convolutional hidden layer 4322a and use pooling functions to generate a condensed activation map (or feature map). Max pooling is an example of a function performed by a pooling hidden layer. Pooling hidden layer 4322a uses other forms of pooling functions, such as average pooling, L2-norm pooling, or other suitable pooling functions. Pooling functions (e.g., max pooling filters, L2-norm filters, or other suitable pooling filters) are applied to each activation map included in convolutional hidden layer 4322a. Figure 43 In the example shown, three pooling filters are used to convolve the three activation maps in the hidden layer 4322a.

[0378] In some examples, max pooling can be used by applying a max pooling filter (e.g., of size 2×2) with a stride (e.g., equal to the size of the filter, such as a stride of 2) to the activation map output from convolutional hidden layer 4322a. The output from the max pooling filter includes the maximum value in each sub-region around which the filter is located. Using a 2×2 filter as an example, each unit in the pooling layer is able to summarize a region of 2×2 nodes in the previous layer (where each node is a value in the activation map). For example, 4 values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter at each iteration of the filter, with the maximum value from the 4 values ​​output as the "maximum value". If such a max pooling filter is applied to the activation filter from convolutional hidden layer 4322a with a dimension of 24×24 nodes, the output from pooling hidden layer 4322b will be an array of 12×12 nodes.

[0379] In some examples, L2-norm pooling filters can also be used. L2-norm pooling filters involve calculating the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of calculating the maximum value as in max pooling), and using the calculated value as the output.

[0380] Intuitively, pooling functions (e.g., max pooling, L2-norm pooling, or other pooling functions) determine whether a given feature is found at any location in an image region and discard precise location information. This can be done without affecting the feature detection results because once a feature has been found, its exact location is less important than its approximate location relative to other features. Max pooling (and other pooling methods) offers the benefit of significantly fewer pooled features, thus reducing the number of parameters required for subsequent layers in the CNN 4300.

[0381] The final connected layer in the network is a fully connected layer, which connects each node from the pooling hidden layer 4322b to each output node in the output layer 4324. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 4322a comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (towards the filter) to 3 activation maps, and the pooling hidden layer 4322b comprises a layer of 3×12×12 hidden feature nodes based on applying a max-pooling filter to a 2×2 region across each of the 3 feature maps. Extending this example, the output layer 4324 can include 10 output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 4322b is connected to each node of the output layer 4324.

[0382] The fully connected layer 4322c takes the output of the previous pooling hidden layer 4322b (which should represent an activation map of high-level features) and determines the features most relevant to a particular class. For example, the fully connected layer 4322c can determine the high-level features most relevant to a particular class and can include weights (nodes) for those high-level features. A product can be computed between the weights of the fully connected layer 4322c and the weights of the pooling hidden layer 4322b to obtain probabilities for different classes. For example, if the CNN 4300 is used to predict that an object in a video frame is a person, there will be high values ​​in the activation map that represent the high-level features of a person (e.g., two legs, a face at the top of the object, two eyes at the upper left and upper right of the face, a nose in the middle of the face, a mouth at the bottom of the face, and / or other features about a person).

[0383] In some examples, the output from output layer 4324 may include an M-dimensional vector (M = 10 in the previous example), where M may include the number of categories from which the program must select when classifying objects in an image. Other exemplary outputs may also be provided. Each number in the N-dimensional vector may represent the probability that an object belongs to a certain category. In an illustrative example, if the 10-dimensional output vector representing objects in 10 different categories is [0 0 0.05 0.8 0 0.15 00 0 0], then the vector indicates that the probability that the image is a third category object (e.g., a dog) is 5%, the probability that the image is a fourth category object (e.g., a person) is 80%, and the probability that the image is a sixth category object (e.g., a kangaroo) is 15%. The probability for a category can be considered as the confidence level that the object belongs to that category.

[0384] Various object detectors can be used to perform object detection and / or classification. One example includes a detector based on the Cifar-10 neural network. Figure 44 This is a diagram illustrating an example of a Cifar-10 neural network 4400. In some cases, a Cifar-10 neural network can be trained to classify only people and cars. As shown, the Cifar-10 neural network 4400 includes various convolutional layers (Conv1 layer 4402, Conv2 / ReLU2 layer 4408, and Conv3 / ReLU3 layer 4414), multiple pooling layers (Pool1 / ReLU1 layer 4404, Pool2 layer 4410, and Pool3 layer 4416), and mixed with adjusted linear unit layers. Normalization layers Norm 1 4406 and Norm 2 4412 are also provided. The final layer is ip1 layer 4418.

[0385] Another deep learning-based detector capable of detecting and / or classifying objects in images is the SSD detector, a fast, single-step object detector applicable to multiple object types or categories. The SSD model uses multi-scale convolutional bounding box outputs from multiple feature maps attached to the top of the neural network. This representation allows SSD to efficiently model a wide variety of box shapes. Figure 45A Including images, and Figure 45B and Figure 45C Includes diagrams illustrating how the SSD detector (based on the VGG deep network model) operates. For example, SSD matches objects to default boxes with different aspect ratios (in... Figure 45B and Figure 45C(Displayed as a dashed rectangle). Each element of the feature map has many default boxes associated with it. Any default box that intersects the union of ground truth boxes with values ​​greater than a threshold (e.g., 0.4, 0.5, 0.6, or other suitable thresholds) is considered a match with respect to the object. For example, two default boxes in an 8×8 box (in...) Figure 45B (shown in blue) matches the cat, while one of the 4x4 boxes (in) Figure 45C (Shown in red) is matched with a dog. SSD has multiple feature maps, and each feature map is responsible for objects at different scales, allowing it to identify objects across a wide range of scales. For example, Figure 45B The bounding box in the 8×8 feature map is smaller than Figure 45C The bounding boxes in the 4×4 feature map. In an illustrative example, the SSD detector is able to have a total of 6 feature maps.

[0386] For each default box in each cell, the SSD neural network outputs a probability vector of length c, where c is the class number representing the probability that the box contains an object of each class. In some cases, a background class is included, indicating that there is no object in the box. The SSD network also outputs (for each default box in each cell) an offset vector with four entries containing the predicted offsets needed to match the default box to the bounding box of the underlying object. The vector is given in the format (cx, cy, w, h), where cx represents the center x, cy represents the center y, w represents the width offset, and h represents the height offset. The vector is only meaningful if the default box actually contains an object. Figure 45A In the image shown, all probability labels, except for the three matching boxes (two for cats and one for dogs), indicate the background category.

[0387] Another deep learning-based detector capable of detecting and / or classifying objects in images includes the "You Only Need to See Once" (YOLO) detector, an alternative to the SSD object detection system. Figure 46A Including images, and Figure 46B and Figure 46C This includes diagrams illustrating how the YOLO detector works. The YOLO detector can apply a single neural network to an entire image. As shown, the YOLO network divides the image into multiple regions and predicts bounding boxes and probabilities for each region. These bounding boxes are weighted by the predicted probabilities. For example, as... Figure 46A As shown, the YOLO detector divides the image into a 13×13 grid. Each cell is responsible for predicting 5 bounding boxes. A confidence score is provided, which indicates how well the predicted bounding boxes actually enclose objects. This score does not include classification of objects that might be inside the boxes, but it indicates whether the shape of the boxes is appropriate. The predicted bounding boxes are... Figure 46B As shown in the diagram, boxes with higher confidence scores have thicker borders.

[0388] Each cell also predicts a category for each bounding box. For example, a probability distribution over all possible categories is provided. Any number of categories can be detected, such as bicycle, dog, cat, person, car, or other suitable object categories. The confidence score and category prediction for the bounding box are combined into a final score, which indicates the probability that the bounding box contains a specific type of object. For example, in Figure 46B The yellow box with a thick border on the upper left of the image is 85% certain that it contains the object category "dog". There are 169 grid cells (13×13), and each cell predicts 5 bounding boxes, forming a total of 4645 bounding boxes. Many bounding boxes will have very low scores, in which case only boxes with a final score above a threshold (e.g., above 30%, 40%, 50%, or other suitable thresholds) are retained. Figure 46C The image shown contains images with final predicted bounding boxes and categories, including dogs, bicycles, and cars. As shown, from a total of 4645 bounding boxes generated, only... Figure 46C The three bounding boxes shown were preserved because they had the best final scores.

[0389] Figure 47 These are diagrams illustrating examples of systems used to implement certain aspects of this technology. Specifically, Figure 47 An example of a computing system 4700 is shown. This computing system can be any computing device, such as an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system communicate with each other using connection 4705. Connection 4705 can be a physical connection using a bus, or a direct connection such as into processor 4710 in a chipset architecture. Connection 4705 can also be a virtual connection, a networking connection, or a logical connection.

[0390] In some embodiments, the computing system 4700 is a distributed system, wherein the functions described herein may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent a plurality of such components, each performing some or all of the functions described therein. In some embodiments, the component may be a physical or virtual device.

[0391] The exemplary system 4700 includes at least one processing unit (CPU or processor) 4710 and a connection 4705 that couples various system components, including system memory 4715 such as read-only memory (ROM) 4720 and random access memory (RAM) 4725, to the processor 4710. The computing system 4700 may include a cache 4712 of high-speed memory that is directly connected to, closely proximate to, or integrated into the processor 4710.

[0392] Processor 4710 may include any general-purpose processor as well as hardware or software services configured to control processor 4710 (such as services 4732, 4734, and 4736 stored in storage device 4730) and dedicated processors that incorporate software instructions into the actual processor design. Processor 4710 may essentially be a completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0393] To enable user interaction, the computing system 4700 includes an input device 4745, which can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphic input, a keyboard, a mouse, motion input, voice input, etc. The computing system 4700 may also include an output device 4735, which can be one or more of a plurality of output mechanisms. In some cases, a multimodal system allows the user to provide multiple types of input / output to communicate with the computing system 4700. The computing system 4700 may include a communication interface 4740, which typically controls and manages user input and system output. The communication interface can use wired and / or wireless transceivers to perform or facilitate the reception and / or transmission of wired or wireless communications, including wired or wireless communications that fully utilize the following: audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Apple® Lightning® ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, proprietary wired ports / plugs, BLUETOOTH® wireless signaling, BLUETOOTH® Low Energy (BLE) wireless signaling, IBEACON® wireless signaling, Radio Frequency Identification (RFID) wireless signaling, Near Field Communication (NFC) wireless signaling, Dedicated Short Range Communication (DSRC) wireless signaling, and 802.11. The communication interface 4740 may include Wi-Fi wireless signal transmission, Wireless Local Area Network (WLAN) signal transmission, Visible Light Communication (VLC), Microwave Access Global Interoperability (WiMAX), Infrared (IR) wireless signal transmission, Public Switched Telephone Network (PSTN) signal transmission, Integrated Services Digital Network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 4740 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 4700 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the U.S. Global Positioning System (GPS), the Russian Global Navigation Satellite System (GLONASS), the Chinese BeiDou Navigation Satellite System (BDS), and the European Galileo Global GNSS. There are no limitations on operation on any particular hardware device; therefore, the basic functionality described herein can be easily replaced with improved hardware or firmware devices during development.

[0394] Storage device 4730 may be a non-volatile and / or non-transitory and / or computer-readable storage device and may be a hard disk or other type of computer-readable medium that can store data accessible by a computer, such as magnetic tape, flash memory cards, solid-state storage devices, digital multifunction disks, tape frames, floppy disks, hard disks, magnetic tape, magnetic stripes / sheets, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, optical disc read-only memory (CD-ROM), rewritable optical disc (CD), digital video disc (DVD), Blu-ray disc (BDD), holographic disc, another optical medium, secure digital storage (SD) cards, micro secure digital storage (microSD) cards, Memory Stick® cards, smart card chips, EMV. Chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase-change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or tape frame, and / or combinations thereof.

[0395] Storage device 4730 may include software services, servers, services, etc., which enable the system to perform functions when the code defining such software is executed by processor 4710. In some embodiments, hardware services that perform specific functions may include a combination of software components stored in a computer-readable medium and necessary hardware components (such as processor 4710, connection 4705, output device 4735, etc.) to perform functions.

[0396] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include, but does not include, non-transitory media on which data can be stored, carrier waves and / or transient electronic signals propagated wirelessly or via a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as compressed optical discs (CDs) or digital universal disks (DVDs), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent programs, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments can be coupled to another code segment or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., can be passed, forwarded, or transmitted using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0397] In some embodiments, computer-readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, by reference, non-transitory computer-readable storage media explicitly excludes media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0398] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that these embodiments can be practiced without these specific details. For clarity, in some cases, the technology may be represented as including individual functional blocks, which include functional blocks comprising devices, device components, steps, or routines in a method embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring the embodiments with unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.

[0399] The various embodiments described above can be presented as processes or methods, depicted as flowcharts, data flow diagrams, structural diagrams, or block diagrams. While a flowchart can describe operations as a sequential process, many operations can be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but may have additional steps not included in the diagram. A process can correspond to a method, function, program, subroutine, subroutine, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.

[0400] The processes and methods according to the examples above can be implemented using computer-executable instructions stored in or available from a computer-readable medium. Such instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a particular function or group of functions. The portion may be accessible via a network of the computer resources used. Computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include hard disks, optical disks, flash memory, USB devices equipped with non-volatile memory, networked storage devices, etc.

[0401] Apparatus implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments for performing necessary tasks (e.g., computer program products) may be stored on a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small-form-factor personal computers, personal digital assistants, rack-mount devices, standalone devices, and so on. The functionality described herein may also be embodied in peripheral devices or expansion cards. By further example, this functionality may also be implemented on a circuit board between different chips or different processes executed in a single chip.

[0402] Instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are exemplary components for providing the functionality described in this disclosure.

[0403] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it should be understood that the inventive concept may be implemented and employed in other ways, and the appended claims are intended to be construed as including such variations other than those limited by the prior art. Various features and aspects of the above-described applications may be used individually or in combination. Furthermore, embodiments may be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings are to be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in a different order than that described.

[0404] Those skilled in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“”) and greater than or equal to (“”) symbols, respectively.

[0405] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform operations, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform operations, or any combination thereof.

[0406] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0407] The claim language or other language that states "at least one of" and / or "one or more of" in a set indicates that one or more members of the set (in any combination) satisfy the claim. For example, the claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. "At least one of" and / or "one or more of" in a set does not limit the set to items listed in that set. For example, the claim language stating "at least one of A and B" or "at least one of A or B" can mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0408] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0409] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication handsets, or multi-purpose integrated circuit devices, including applications in wireless communication handsets and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the technology can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging materials. The computer-readable medium can include memory or data storage media, such as random access memory (RAM), such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively or concurrently, the technology may be implemented at least in part by a computer-readable communication medium (such as a propagating signal or wave) that carries or transmits code in the form of instructions or data structures and can be accessed, read, and / or executed by a computer or other processor.

[0410] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, alternatively, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein.

[0411] The specification of this disclosure includes:

[0412] Aspect 1: A method for processing one or more frames, the method comprising: determining a region of interest in a first frame in a frame sequence, the region of interest in the first frame including an object having a certain size in the first frame; cropping a portion of a second frame in the frame sequence, the second frame appearing after the first frame in the frame sequence; and scaling the portion of the second frame based on the size of the object in the first frame.

[0413] Aspect 2: The method according to aspect 1 further includes: receiving user input corresponding to the selection of an object in the first frame; and determining the region of interest in the first frame based on the received user input.

[0414] Aspect 3: According to the method of aspect 2, the user input includes touch input provided using the touch interface of the device.

[0415] Aspect 4: The method according to any one of Aspects 1 to 3 further includes: determining a point of an object region determined for an object in the second frame; and cropping and scaling the portion of the second frame if the point of the object region is located at the center of the cropped and scaled portion.

[0416] Aspect 5: According to the method described in aspect 4, the point of the object region is the center point of the object region.

[0417] Aspect 6: The method according to any one of Aspects 1 to 5, wherein the portion of the second frame is scaled based on the size of the object in the first frame such that the object in the second frame has the same size as the object in the first frame.

[0418] Aspect 7: The method according to any one of Aspects 1 to 6 further includes: determining a first length associated with an object in the first frame; determining a second length associated with an object in the second frame; determining a scaling factor based on a comparison between the first length and the second length; and scaling the portion of the second frame based on the scaling factor.

[0419] Aspect 8: According to the method of aspect 7, wherein the first length is the length of a first object region determined for objects in the first frame, and wherein the second length is the length of a second object region determined for objects in the second frame.

[0420] Aspect 9: According to the method of aspect 8, wherein the first object region is a first bounding box, and the first length is the diagonal length of the first bounding box, and wherein the second object region is a second bounding box, and the second length is the diagonal length of the second bounding box.

[0421] Aspect 10: The method according to any one of Aspects 8 or 9, wherein the portion of the second frame is scaled based on the scaling factor so that the second object region in the cropped and scaled portion has the same size as the first object region in the first frame.

[0422] Aspect 11: The method according to any one of Aspects 1 to 10 further includes: determining points of a first object region generated for an object in the first frame; determining points of a second object region generated for an object in the second frame; determining a motion factor for the object based on a smoothing function using the points of the first object region and the points of the second object region, wherein the smoothing function controls changes in the position of the object in a plurality of frames of the sequence of frames; and cropping the portion of the second frame based on the motion factor.

[0423] Aspect 12: According to the method of aspect 11, the point of the first object region is the center point of the first object region, and the point of the second object region is the center point of the second object region.

[0424] Aspect 13: The method according to any one of Aspects 11 or 12, wherein the smoothing function includes a movement function for determining the position of a point of the corresponding object region in each of the plurality of frames of the frame sequence based on a statistical metric of object movement.

[0425] Aspect 14: The method according to any one of Aspects 1 to 13 further includes: determining a first length associated with an object in the first frame; determining a second length associated with an object in the second frame; determining a scaling factor for the object based on a comparison between the first length and the second length and based on a smoothing function using the first length and the second length, wherein the smoothing function controls the variation in size of the object in a plurality of frames of the frame sequence; and scaling the portion of the second frame based on the scaling factor.

[0426] Aspect 15: According to the method of aspect 14, wherein the smoothing function includes a moving function for determining the length associated with the object in each of a plurality of frames of the frame sequence based on a statistical metric of the object size.

[0427] Aspect 16: The method according to any one of Aspects 14 or 15, wherein the first length is the length of a first bounding box generated for an object in the first frame, and wherein the second length is the length of a second bounding box generated for an object in the second frame.

[0428] Aspect 17: According to the method of aspect 16, wherein the first length is the diagonal length of the first bounding box, and wherein the second length is the diagonal length of the second bounding box.

[0429] Aspect 18: The method according to any one of Aspects 16 or 17, wherein the portion of the second frame is scaled based on the scaling factor such that the second bounding box in the cropped and scaled portion has the same size as the first bounding box in the first frame.

[0430] Aspect 19: The method according to any one of aspects 1 to 18, wherein the cropping and scaling of the portion of the second frame keeps the object centered in the second frame.

[0431] Aspect 20: The method according to any one of aspects 1 to 19 further includes: detecting and tracking the object in one or more frames of the frame sequence.

[0432] Aspect 21: An apparatus for processing one or more frames, comprising: a memory configured to store at least one frame; and a processor implemented in a circuit and configured to: determine a region of interest in a first frame in a frame sequence, the region of interest in the first frame including an object having a size in the first frame; crop a portion of a second frame in the frame sequence, the second frame appearing after the first frame in the frame sequence; and scale the portion of the second frame to maintain the size of the object in the second frame.

[0433] Aspect 22: The apparatus according to aspect 21, wherein the processor is configured to: receive user input corresponding to the selection of an object in the first frame; and determine the region of interest in the first frame based on the received user input.

[0434] Aspect 23: The apparatus according to aspect 22, wherein the user input includes touch input provided using the device's touch interface.

[0435] Aspect 24: The apparatus according to any one of aspects 21 to 23, wherein the processor is configured to: determine a point of an object region determined for an object in the second frame; and, if the point of the object region is located at the center of the cropped and scaled portion, crop and scale the portion of the second frame.

[0436] Aspect 25: The apparatus according to aspect 24, wherein the point of the object region is the center point of the object region.

[0437] Aspect 26: The apparatus according to any one of aspects 21 to 25, wherein the portion of the second frame is scaled based on the size of the object in the first frame such that the object in the second frame has the same size as the object in the first frame.

[0438] Aspect 27: An apparatus according to any one of aspects 21 to 26, wherein the processor is configured to: determine a first length associated with an object in the first frame; determine a second length associated with an object in the second frame; determine a scaling factor based on a comparison between the first length and the second length; and scale the portion of the second frame based on the scaling factor.

[0439] Aspect 28: The apparatus according to aspect 27, wherein the first length is the length of a first object region determined for objects in the first frame, and wherein the second length is the length of a second object region determined for objects in the second frame.

[0440] Aspect 29: The apparatus according to aspect 28, wherein the first object region is a first bounding box and the first length is the diagonal length of the first bounding box, and wherein the second object region is a second bounding box and the second length is the diagonal length of the second bounding box.

[0441] Aspect 30: The apparatus according to any one of Aspects 28 or 29, wherein the portion of the second frame is scaled based on the scaling factor such that the second object region in the cropped and scaled portion has the same size as the first object region in the first frame.

[0442] Aspect 31: An apparatus according to any one of aspects 21 to 30, wherein the processor is configured to: determine points of a first object region generated for an object in the first frame; determine points of a second object region generated for an object in the second frame; determine a motion factor for the object based on a smoothing function using points of the first object region and points of the second object region, wherein the smoothing function controls changes in the position of the object in a plurality of frames of the sequence of frames; and crop the portion of the second frame based on the motion factor.

[0443] Aspect 32: The apparatus according to aspect 31, wherein the point of the first object region is the center point of the first object region, and wherein the point of the second object region is the center point of the second object region.

[0444] Aspect 33: The apparatus according to any one of Aspects 31 or 32, wherein the smoothing function comprises a moving average function, the moving average function being used for the average position of points of the corresponding object region in each of the plurality of frames of the frame sequence.

[0445] Aspect 34: An apparatus according to any one of aspects 21 to 33, wherein the processor is configured to: determine a first length associated with an object in the first frame; determine a second length associated with an object in the second frame; determine a scaling factor for the object based on a comparison between the first length and the second length and based on a smoothing function using the first length and the second length, wherein the smoothing function controls the size of the object to change progressively across a plurality of frames in the frame sequence; and scale the portion of the second frame based on the scaling factor.

[0446] Aspect 35: The apparatus according to aspect 34, wherein the smoothing function includes a moving average function for determining an average length associated with the object in each of a plurality of frames of the frame sequence.

[0447] Aspect 36: The apparatus according to any one of Aspects 34 or 35, wherein the first length is the length of a first bounding box generated for an object in the first frame, and wherein the second length is the length of a second bounding box generated for an object in the second frame.

[0448] Aspect 37: The apparatus according to aspect 36, wherein the first length is the diagonal length of the first bounding box, and wherein the second length is the diagonal length of the second bounding box.

[0449] Aspect 38: The apparatus according to any one of aspects 34 to 37, wherein the portion of the second frame is scaled based on the scaling factor such that the second bounding box in the cropped and scaled portion has the same size as the first bounding box in the first frame.

[0450] Aspect 39: The apparatus according to any one of aspects 21 to 38, wherein the cropping and scaling performed on the portion of the second frame keeps the object centered in the second frame.

[0451] Aspect 40: The apparatus according to any one of aspects 21 to 39, wherein the processor is configured to detect and track the object in one or more frames of the frame sequence.

[0452] Aspect 41: An apparatus according to any one of aspects 21 to 40, wherein the apparatus comprises a mobile device having a camera for capturing the at least one frame.

[0453] Aspect 42: The apparatus according to any one of aspects 21 to 41 further includes a display for displaying one or more images.

[0454] Aspect 43: A computer-readable medium having instructions stored thereon that, when executed by a processor, perform any of the operations described in aspects 1 to 40.

[0455] Aspect 44: An apparatus comprising components for performing any of the operations described in aspects 1 to 40.

Claims

1. A method for processing one or more frames, the method comprising: Determine the region of interest in the first frame of the frame sequence, wherein the region of interest has a certain size; The corresponding portion of the second frame of the frame sequence is cropped based on the size of the region of interest in the first frame and a smoothing function configured to control the change of the position of the region of interest between the first frame and the second frame, the second frame being captured after the first frame; as well as The corresponding portion of the second frame is scaled based on the size of the region of interest in the first frame.

2. The method according to claim 1, further comprising: Receive user input corresponding to the selection of an object in the first frame; as well as The region of interest in the first frame is determined based on the received user input.

3. The method of claim 2, wherein the user input includes touch input provided using the device's touch interface.

4. The method according to claim 1, further comprising: Determine the points of the region of interest in the second frame; as well as If the point in the region of interest is located at the center of the cropped and scaled portion, the corresponding portion of the second frame is cropped and scaled.

5. The method according to claim 4, wherein the point in the region of interest is the center point of the region of interest.

6. The method of claim 1, wherein the corresponding portion of the second frame is scaled based on the size of the region of interest in the first frame so that the region of interest in the second frame has the same size as the region of interest in the first frame.

7. The method according to claim 1, further comprising: Determine a first length associated with the region of interest in the first frame; Determine a second length associated with the region of interest in the second frame; The scaling factor is determined based on a comparison between the first length and the second length; as well as The corresponding portion of the second frame is scaled based on the scaling factor.

8. The method of claim 7, wherein the first length is the length of a first object region determined for the region of interest in the first frame, and wherein the second length is the length of a second object region determined for the region of interest in the second frame.

9. The method of claim 8, wherein the first object region is a first bounding box, and the first length is the diagonal length of the first bounding box, and wherein the second object region is a second bounding box, and the second length is the diagonal length of the second bounding box.

10. The method of claim 8, wherein the corresponding portion of the second frame is scaled based on the scaling factor so that the second object region in the cropped and scaled portion has the same size as the first object region in the first frame.

11. The method according to claim 1, further comprising: Determine the points of the first object region generated for the region of interest in the first frame; Determine the points of the second object region generated for the region of interest in the second frame; A movement factor for the region of interest is determined based on a smoothing function using the points of the first object region and the points of the second object region. as well as Based on the motion factor, the corresponding portion of the second frame is cropped.

12. The method of claim 11, wherein the point in the first object region is the center point of the first object region, and wherein the point in the second object region is the center point of the second object region.

13. The method of claim 11, wherein the smoothing function includes a moving function for determining the position of the point of the corresponding object region in each frame of the frame sequence based on a statistical measure of movement of the region of interest.

14. The method of claim 1, further comprising: Determine a first length associated with the region of interest in the first frame; Determine a second length associated with the region of interest in the second frame; A scaling factor for the region of interest is determined based on a comparison between the first length and the second length and based on a smoothing function using the first length and the second length, wherein the smoothing function controls the variation in the size of the region of interest in the first frame and the second frame; as well as The corresponding portion of the second frame is scaled based on the scaling factor.

15. The method of claim 14, wherein the smoothing function comprises a moving function for determining the length associated with the region of interest in each frame of the frame sequence based on a statistical metric of the region of interest size.

16. The method of claim 14, wherein the first length is the length of a first bounding box generated for the region of interest in the first frame, and wherein the second length is the length of a second bounding box generated for the region of interest in the second frame.

17. The method of claim 16, wherein the first length is the diagonal length of the first bounding box, and wherein the second length is the diagonal length of the second bounding box.

18. The method of claim 16, wherein the corresponding portion of the second frame is scaled based on the scaling factor such that the second bounding box in the cropped and scaled portion has the same size as the first bounding box in the first frame.

19. The method of claim 1, wherein the cropping and scaling of the corresponding portion of the second frame keeps the region of interest centered in the second frame.

20. The method of claim 1, further comprising: Detect and track the region of interest in one or more frames of the frame sequence.

21. An apparatus for processing one or more frames, comprising: A memory configured to store at least one frame; as well as A processor, which is implemented in a circuit and configured to: Determine the region of interest in the first frame of the frame sequence, wherein the region of interest has a certain size; The corresponding portion of the second frame of the frame sequence is cropped based on the size of the region of interest in the first frame and a smoothing function configured to control the change of the position of the region of interest between the first frame and the second frame, the second frame being captured after the first frame; as well as The corresponding portion of the second frame is scaled based on the size of the region of interest in the first frame.

22. The apparatus of claim 21, wherein the processor is configured to: Receive user input corresponding to the selection of an object in the first frame; and The region of interest in the first frame is determined based on the received user input.

23. The apparatus of claim 22, wherein the user input includes touch input provided using the device's touch interface.

24. The apparatus of claim 21, wherein the processor is configured to: Determine the points in the region of interest in the second frame; and If the point in the region of interest is located at the center of the cropped and scaled portion, the corresponding portion of the second frame is cropped and scaled.

25. The apparatus of claim 24, wherein the point in the region of interest is the center point of the region of interest.

26. The apparatus of claim 21, wherein the corresponding portion of the second frame is scaled based on the size of the region of interest in the first frame such that the region of interest in the second frame has the same size as the region of interest in the first frame.

27. The apparatus of claim 21, wherein the processor is configured to: Determine a first length associated with the region of interest in the first frame; Determine a second length associated with the region of interest in the second frame; The scaling factor is determined based on a comparison between the first length and the second length; and The corresponding portion of the second frame is scaled based on the scaling factor.

28. The apparatus of claim 27, wherein the first length is the length of a first object region determined with respect to the region of interest in the first frame, and wherein the second length is the length of a second object region determined with respect to the region of interest in the second frame.

29. The apparatus of claim 28, wherein the first object region is a first bounding box, and the first length is the diagonal length of the first bounding box, and wherein the second object region is a second bounding box, and the second length is the diagonal length of the second bounding box.

30. The apparatus of claim 28, wherein the corresponding portion of the second frame is scaled based on the scaling factor such that the second object region in the cropped and scaled portion has the same size as the first object region in the first frame.

31. The apparatus of claim 21, wherein the processor is configured to: Determine the points of the first object region generated for the region of interest in the first frame; Determine the points of the second object region generated for the region of interest in the second frame; A motion factor for the region of interest is determined based on a smoothing function using the points in the first object region and the points in the second object region; as well as Based on the motion factor, the corresponding portion of the second frame is cropped.

32. The apparatus of claim 31, wherein the point in the first object region is the center point of the first object region, and wherein the point in the second object region is the center point of the second object region.

33. The apparatus of claim 31, wherein the smoothing function includes a movement function for determining the position of the point of the corresponding object region in each frame of the frame sequence based on a statistical metric of movement of the region of interest.

34. The apparatus of claim 21, wherein the processor is configured to: Determine a first length associated with the region of interest in the first frame; Determine a second length associated with the region of interest in the second frame; A scaling factor for the region of interest is determined based on a comparison between the first length and the second length and based on a smoothing function using the first length and the second length, wherein the smoothing function controls the variation in the size of the region of interest in the first frame and the second frame; and The corresponding portion of the second frame is scaled based on the scaling factor.

35. The apparatus of claim 34, wherein the smoothing function comprises a moving function for determining the length associated with the region of interest in each frame of the frame sequence based on a statistical metric of the region of interest size.

36. The apparatus of claim 34, wherein the first length is the length of a first bounding box generated for the region of interest in the first frame, and wherein the second length is the length of a second bounding box generated for the region of interest in the second frame.

37. The apparatus of claim 36, wherein the first length is the diagonal length of the first bounding box, and wherein the second length is the diagonal length of the second bounding box.

38. The apparatus of claim 36, wherein the corresponding portion of the second frame is scaled based on the scaling factor such that the second bounding box in the cropped and scaled portion has the same size as the first bounding box in the first frame.

39. The apparatus of claim 21, wherein the cropping and scaling of the corresponding portion of the second frame keeps the region of interest centered in the second frame.

40. The apparatus of claim 21, wherein the processor is configured to: Detect and track the region of interest in one or more frames of the frame sequence.

41. The apparatus of claim 21, wherein the apparatus comprises a mobile device having a camera for capturing the at least one frame.

42. The apparatus of claim 21, further comprising: A display for displaying one or more images.

43. An apparatus for processing one or more frames, comprising: Components for performing the method according to any one of claims 1-20.

44. A computer-readable medium having program code recorded thereon, wherein, The program code may be executed by one or more processors to cause the one or more processors to perform the method according to any one of claims 1-20.

45. A computer program product comprising computer-readable instructions, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1-20.

Citation Information

Patent Citations

  • Method and apparatus for automatically rendering dolly zoom effect

    US20140240553A1

  • Zooming control apparatus, image capturing apparatus and control methods thereof

    US20170272661A1

  • Image processing device, control method thereof, and imaging device

    US20190096047A1

  • Photographing control method, apparatus and device and device and storage medium

    WO2020047745A1