Sparse optical flow estimation

By using sparse optical flow estimation techniques, a subset of pixels from a frame is selected and sparse optical flow maps are generated using masks and machine learning algorithms. This solves the problems of excessive computational complexity and resource requirements in existing optical flow estimation techniques, and improves the performance and accuracy of computer vision and extended reality systems.

CN116235209BActive Publication Date: 2026-03-17QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-22
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing optical flow estimation techniques suffer from high computational complexity, high energy consumption, and reduced accuracy when processing all pixels in a frame or image, especially in computer vision and extended reality systems where they have excessive computational and storage requirements.

Method used

A sparse optical flow estimation method is adopted, which selects a subset of pixels in the frame and uses masks and machine learning algorithms to determine key features and generate a sparse optical flow map, thereby reducing the amount of computation and resource requirements.

Benefits of technology

It reduces computational complexity and power consumption, improves the performance of computer vision and extended reality systems, reduces storage requirements, and maintains high estimation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116235209B_ABST
    Figure CN116235209B_ABST
Patent Text Reader

Abstract

Systems and techniques for performing optical flow estimation between one or more frames are described herein. For example, a process can include determining a subset of pixels of at least one of a first frame and a second frame, and generating a mask indicative of the subset of pixels. The process can include determining one or more features associated with the subset of pixels of at least the first frame and the second frame based on the mask. The process can include determining optical flow vectors between the subset of pixels of the first frame and corresponding pixels of the second frame. The process can include generating an optical flow map of the second frame using the optical flow vectors.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] field

[0002] This disclosure generally relates to optical flow estimation. In some examples, aspects of this disclosure relate to performing sparse optical flow estimation.

[0003] background

[0004] Many devices and systems allow a scene to be captured by generating images (or frames) and / or video data (including multiple frames). For example, a camera or a device that includes a camera can capture a sequence of frames of a scene (e.g., video of the scene). In some cases, the frame sequence can be processed to perform one or more functions, can be output for display, can be output for processing and / or consumption by other devices, and for other purposes.

[0005] A common type of processing performed on frame sequences is motion estimation, which involves tracking the motion of objects or points across multiple frames. For example, motion estimation may include determining an optical flow map that describes the displacement of pixels in a frame relative to corresponding pixels in previous frames. Motion estimation can be used in a wide variety of applications, including computer vision systems, extended reality (XR) systems, data compression, image segmentation, autonomous vehicle operation, and others.

[0006] Overview

[0007] Systems and techniques for performing sparse optical flow estimation of frames, for example, based on subsets and / or regions within a frame, are described. In an illustrative example, a method for optical flow estimation between one or more frames is provided. The method includes: determining a subset of pixels in at least one of a first frame and a second frame; generating a mask indicating the subset of pixels; determining one or more features associated with the subset of pixels in at least the first frame and the second frame based on the mask; determining an optical flow vector between the subset of pixels in the first frame and corresponding pixels in the second frame; and using the optical flow vector to generate an optical flow map of the second frame.

[0008] In another example, an apparatus for encoding video data is provided, the apparatus including a memory and one or more processors (e.g., implemented in a circuit system) coupled to the memory. In some examples, more than one processor may be coupled to the memory and may be used to perform one or more operations described herein. The one or more processors are configured to: determine a subset of pixels in at least one of a first frame and a second frame; generate a mask indicating the subset of pixels; determine one or more features associated with the subset of pixels in at least the first frame and the second frame based on the mask; determine an optical flow vector between the subset of pixels in the first frame and corresponding pixels in the second frame; and use the optical flow vector to generate an optical flow map of the second frame.

[0009] In another example, a non-transient computer-readable medium is provided for encoding video data, on which instructions are stored, which, when executed by one or more processors, cause the one or more processors to: determine a subset of pixels in at least one of a first frame and a second frame; generate a mask indicating the subset of pixels; determine one or more features associated with the subset of pixels in at least the first frame and the second frame based on the mask; determine an optical flow vector between the subset of pixels in the first frame and corresponding pixels in the second frame; and use the optical flow vector to generate an optical flow map of the second frame.

[0010] In another example, an apparatus for encoding video data is provided. The apparatus includes: means for determining a subset of pixels in at least one of a first frame and a second frame; means for generating a mask indicating the subset of pixels; means for determining one or more features associated with the subset of pixels in at least the first frame and the second frame based on the mask; means for determining an optical flow vector between the subset of pixels in the first frame and corresponding pixels in the second frame; and means for using the optical flow vector to generate an optical flow map of the second frame.

[0011] In some aspects, the methods, apparatus, and computer-readable media described above include: using a machine learning algorithm to determine at least the subset of pixels in the first frame and the second frame.

[0012] In some aspects, determining the subset of pixels of at least the first frame and the second frame may include: identifying a first pixel of the first frame corresponding to a first region of interest; identifying a second pixel of the first frame corresponding to a second region of interest; and including the first pixel in the subset of pixels and excluding the second pixel from the subset of pixels.

[0013] In some respects, determining the subset of pixels of at least the first frame and the second frame includes identifying at least one pixel that corresponds to the boundary between the two objects.

[0014] In some respects, determining the subset of pixels of at least the first frame and the second frame includes sampling a predetermined number of pixels within the first frame.

[0015] In some aspects, at least the subset of pixels in the first frame and the second frame includes at least one pixel having adjacent pixels not included in the subset of pixels. In some cases, determining the one or more features based on the mask includes: determining feature information corresponding to adjacent pixels of the at least one pixel; and storing the feature information corresponding to the adjacent pixels in association with the at least one pixel. In some examples, the above-described methods, apparatus, and computer-readable media include: using the feature information corresponding to the adjacent pixels to determine an optical flow vector between the at least one pixel and a corresponding pixel in the second frame.

[0016] In some aspects, the methods, apparatus, and computer-readable media described above include: determining an importance value for a frame within a frame sequence including the first frame and the second frame; and selecting a frame within the frame sequence for performing optical flow estimation based on the importance value. In some cases, the methods, apparatus, and computer-readable media described above include: selecting a second frame for performing optical flow estimation based on determining that the importance value of the second frame exceeds a threshold importance value.

[0017] In some aspects, the device is a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), an extended reality (XR) device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a personal computer, a laptop computer, a server computer, a camera, a vehicle or a computing device or component of a vehicle, a wearable device, a television set (e.g., a connected television), or other device, and is part of or includes the aforementioned devices. In some aspects, the device includes one or more cameras for capturing one or more frames or images. In some aspects, the device includes a display for displaying one or more frames or images, video content, notifications, and / or other displayable data. In some aspects, the aforementioned device may include one or more sensors.

[0018] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. This subject matter should be understood in conjunction with the appropriate portions of the entire specification of this patent, any or all drawings, and each claim.

[0019] The foregoing, as well as other features and examples, will become more apparent when referenced to the following description, claims and accompanying drawings. Brief description of the attached diagram

[0021] The illustrative examples of this application are described in detail below with reference to the accompanying drawings.

[0022] Figure 1 This is a block diagram illustrating an example optical flow estimation system based on some examples;

[0023] Figure 2A This is an explanation based on some example masks;

[0024] Figure 2B This is an explanation of optical flow estimation for pixels within a mask, based on some examples;

[0025] Figure 3 This is a block diagram illustrating an example optical flow estimation system based on some examples;

[0026] Figure 4A and Figure 4B It is a table explaining the changes in sparsity based on optical flow estimation from some examples;

[0027] Figure 5 This is a flowchart illustrating an example of a process for performing sparse optical flow estimation, based on several examples;

[0028] Figure 6 These are illustrations illustrating examples of deep learning neural networks based on several examples;

[0029] Figure 7 It is an illustration explaining examples of convolutional neural networks based on some examples; and

[0030] Figure 8 These are illustrations illustrating examples of systems used to implement certain aspects described in this article.

[0031] Detailed description

[0032] Certain aspects and examples of this disclosure are provided below. It will be apparent to those skilled in the art that some of these aspects and examples can be applied independently and some can be applied in combination. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of the subject matter of this application. However, it will be apparent that the various examples can be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0033] The following description is provided as an illustrative example only and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description is intended to provide those skilled in the art with enabling descriptions for implementing the illustrative examples. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0034] Motion estimation (also referred to herein as optical flow estimation) is the task of tracking the movement of one or more regions (e.g., an object or a portion of an object, an instance or a portion of an instance, a background portion of a scene or a portion of that background, etc.) across a sequence of frames. For example, an optical flow estimation system can identify pixels in an initial frame that correspond to a portion of a real-world object. The system can then determine corresponding pixels in subsequent frames (e.g., pixels depicting the same portion of a real-world object). The system can estimate the motion of an object between frames by determining optical flow vectors corresponding to the displacement and / or distance between pixels in the initial frame and their corresponding pixels in subsequent frames. For example, an optical flow vector can indicate the displacement (e.g., corresponding to the direction and distance of movement) between coordinates corresponding to an initial pixel and coordinates corresponding to a subsequent pixel.

[0035] An optical flow graph can include one or more optical flow vectors corresponding to motion between two frames. In some cases, an optical flow estimation system can determine an optical flow graph that includes optical flow vectors for each pixel (or approximately each pixel) within a frame. Such an optical flow graph can be referred to as a dense optical flow graph. In some cases, generating a dense optical flow graph may require significant time and / or computational power, which can be detrimental to many applications of motion estimation. Examples of applications utilizing motion estimation include various computer vision tasks and camera applications involving pixel motion (corresponding to object motion). Examples of such applications include video recognition, autonomous driving, video compression, object and / or scene tracking, visual inertial odometry (VIO), video object segmentation, extended reality (XR) (e.g., virtual reality (VR), augmented reality (AR), and / or mixed reality (MR)), and so on. Higher performance is expected for optical flow estimation performed in chips and / or devices, including higher accuracy, lower computational complexity, lower latency, lower power consumption, smaller memory size requirements, and so on.

[0036] As described, optical flow can comprise a dense correspondence estimation problem between a pair of frames or images. Existing solutions generally compute dense optical flow across the entire frame or image (e.g., all pixels in the frame or image), which leads to various problems. The first problem is that uniform (or homogeneous) processing across all pixels and all frames in a video sequence can result in unnecessarily high computational complexity, high latency, high power consumption, and / or high memory usage / requirements (e.g., potentially requiring large memory). Another problem is that uniform (homogeneous) processing across all pixels in a frame or image can lead to degraded accuracy performance. For example, degraded accuracy performance can result from oversampling in smooth regions and undersampling in boundary or discontinuous regions.

[0037] This document describes systems, apparatuses, methods, and computer-readable media (collectively, the “Systems and Techniques”) for performing sparse optical flow estimation for frames. A frame may also be referred to herein as an image. In some cases, the optical flow estimation system may determine a subset of pixels for performing optical flow estimation. The optical flow estimation system may generate a sparse optical flow map based on the subset of pixels. For example, the optical flow estimation system may perform differential and attention field masking techniques to predict or identify spatial regions or temporal frames (e.g., regions with movement) that are important for optical flow estimation. Pixels identified using field masking techniques may be identified using a mask, which may be referred to in some cases as a field mask or an attention field mask. The optical flow estimation system may perform optical flow estimation for pixels indicated by the mask in regions of one or more frames (e.g., only pixels in these regions). In some examples, the mask may be refined (e.g., in terms of resolution) to indicate regions for further flow refinement.

[0038] Because the systems and techniques described herein perform optical flow estimation on significantly fewer pixels than conventional full-frame optical flow estimation, they can generate optical flow maps with reduced latency and / or fewer computational resources. For example, the systems and techniques described herein can result in optical flow estimation being performed on 50% of the pixels in a frame, 40% of the pixels in a frame, 5% of the pixels in a frame, 0% of the pixels in a frame (no pixels), or any other number based on the techniques described herein.

[0039] As mentioned above, an optical flow estimation system can determine a subset of pixels in a frame for which optical flow estimation should be performed. In one example, the subset of pixels may correspond to relevant and / or significant features within the frame. For example, the subset of pixels may correspond to objects in the foreground of the frame, objects moving within the frame, objects with specific labels, and / or other objects within the frame. In another example, the subset of pixels may correspond to the edges of objects and / or the boundaries between objects. In some cases, the optical flow estimation system may use one or more machine learning systems and / or algorithms to determine relevant and / or significant pixels. For example, as described in more detail below, neural network-based machine learning systems and / or algorithms (e.g., deep neural networks) may be used to determine relevant and / or significant pixels in a frame.

[0040] Further details regarding the system used for sparse optical flow estimation are provided in this paper with reference to the accompanying figures. Figure 1 This is a diagram illustrating an example of an optical flow estimation system 100 capable of performing an optical flow estimation process. The optical flow estimation system 100 includes various components, including a selection engine 102, an optical flow vector engine 104, and an optical flow graph engine 106. Each component of the optical flow estimation system 100 may include electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), or other suitable electronic circuitry), computer software, firmware, or any combination thereof, to perform the various operations described herein. While the optical flow estimation system 100 is shown as including certain components, those skilled in the art will appreciate that the optical flow estimation system 100 may include more than […]. Figure 1 The components shown may have more or fewer components. For example, in some instances, the optical flow estimation system 100 may also include... Figure 1 One or more memories (e.g., RAM, ROM, cache, etc.) and / or processing devices not shown.

[0041] Optical flow estimation system 100 may be a computing device or part of multiple computing devices. In some cases, one or more computing devices including optical flow estimation system 100 may also include one or more wireless transceivers for wireless communication and / or a display for displaying one or more frames or images. In some examples, the computing device including optical flow estimation system 100 may be an electronic device, such as a camera (e.g., a digital camera, IP camera, video camera, camera phone, video phone, or other suitable capture device), a mobile or fixed-line handset (e.g., a smartphone, cellular phone, etc.), a desktop computer, a desktop or laptop computer, a tablet computer, an XR device, a vehicle or a computing device or component of a vehicle, a set-top box, a television, a display device, a digital media player, a video game console, a video streaming device, or any other suitable electronic device.

[0042] Optical flow estimation system 100 may receive frame 103 as input. In some examples, optical flow estimation system 100 may perform an optical flow estimation process in response to one or more frames 103 being captured by a camera or a computing device including a camera (e.g., a mobile device, etc.). Frame 103 may include a single frame or multiple frames. For example, frame 103 may include video frames from a video sequence or still images from a set of consecutively captured still images. In one illustrative example, a set of consecutively captured still images may be captured and displayed to a user as a scene preview within the camera's field of view, which can help the user decide when to provide input that allows an image to be captured for storage. In another illustrative example, a set of consecutively captured still images may be captured using burst mode or other similar modes that capture multiple consecutive images. Frames may be red-green-blue (RGB) frames with red, green, and blue components per pixel; luminance, red chrominance, blue chrominance (YCbCr) frames with a luminance component and two chrominance (color) components (red chrominance and blue chrominance) per pixel; or any other suitable type of color or monochrome image.

[0043] In some examples, the optical flow estimation system 100 can capture frame 103. In some examples, the optical flow estimation system 100 can estimate the optical flow from the frame source ( Figure 1(Not shown in the image) Frame 103 is obtained. In some cases, the frame source may include one or more image capture devices and / or one or more video capture devices (e.g., digital camera, digital video camera, telephone with camera, tablet with camera, or other suitable capture device), image and / or video storage devices, image and / or video archives containing stored images, image and / or video servers or content providers providing image and / or video data, image and / or video feed interfaces receiving images from video servers or content providers, computer graphics systems for generating computer graphics images and / or video images, combinations of these sources, or other sources of image frame content. In some cases, multiple frame sources may provide frames to the optical flow estimation system 100.

[0044] In some implementations, the optical flow estimation system 100 and the frame source may be part of the same computing device. For example, in some cases, a camera, telephone, tablet, extended reality (XR) device (e.g., virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, etc.), and / or other devices that include a frame or image source (e.g., camera, storage, etc.) may include an integrated optical flow estimation system. In some implementations, the optical flow estimation system 100 and the frame source may be part of separate computing devices. In an illustrative example, the frame source may include one or more cameras, and the computing device including the optical flow estimation system 100 may include a mobile or landline handset, a desktop computer, a laptop or notebook computer, a tablet computer, or other computing device.

[0045] In some examples, optical flow estimation performed by optical flow estimation system 100 can be performed using a single-camera system of a computing device. In other examples, optical flow estimation performed by optical flow estimation system 100 can be performed using a dual-camera system of a computing device. In some cases, more than two cameras can be used in the camera system to perform optical flow estimation.

[0046] Optical flow estimation system 100 can process frame 103 to generate an optical flow map (e.g., optical flow map 108) by performing optical flow estimation on the pixels within frame 103. Optical flow map 108 may include one or more optical flow vectors corresponding to the movement of pixels between two frames. In some cases, the two frames may be a series of directly adjacent frames within the frame. In some cases, the two frames may be separated by one or more intermediate frames (which may be referred to as non-adjacent frames). In some cases, optical flow map 108 may include optical flow vectors corresponding to a subset of pixels within the frame. For example, optical flow map 108 may be a sparse optical flow map, which includes optical flow vectors corresponding only to a portion of the pixels within the frame. In an illustrative example, optical flow map 108 may include optical flow vectors for 50% of the pixels within the frame. Those skilled in the art will appreciate that optical flow map 108 may include optical flow vectors for any other subset of pixels within the frame (such as 0%, 10%, 40%, 70%, or other numbers).

[0047] Selection engine 102 can process frames to select a subset of pixels in the frame for which optical flow estimation is to be performed. This subset of pixels can be represented by a mask. For example, optical flow estimation system 100 can perform optical flow estimation on pixels included within the mask, but not on pixels outside the mask. In one example, selection engine 102 can determine the mask by sampling the subset of pixels. For example, selection engine 102 can sample a predetermined number of pixels in the frame (e.g., 10 pixels, 100 pixels, etc.). In other examples, selection engine 102 can determine a mask that includes pixels corresponding to important and / or relevant features of the frame. These features can be referred to as key features. In some cases, such key features are associated with motion (e.g., features typically associated with motion within a scene). The mask corresponding to one or more key features can be referred to as a difference field mask and / or an attention field mask. For example, selection engine 102 can identify pixels corresponding to objects in the foreground of the scene (rather than objects in the background of the scene). Additionally or alternatively, in another example, selection engine 102 may identify pixels corresponding to moving objects (rather than stationary objects). Additionally or alternatively, in another example, selection engine 102 may identify pixels corresponding to the boundary between two objects. Additionally or alternatively, in a further example, selection engine 102 may select pixels corresponding to objects with certain labels and / or classifications. For example, selection engine 102 may classify and / or label objects within a frame (e.g., using any type or form of object recognition technology, such as using one or more classification neural networks). Based on classification and / or labels, selection engine 102 may determine pixels corresponding to significant objects (e.g., faces identified using object recognition technology).

[0048] In some cases, selection engine 102 may use a machine learning system and / or algorithm to select a subset of pixels to be included within the mask. For example, the machine learning system and / or algorithm may be any type or form of deep neural network (DNN). In an interpretive example, the machine learning algorithm may include a visual geometry group (VCG) algorithm. In another interpretive example, the machine learning system and / or algorithm may include a residual neural network (ResNet). Any other machine learning system and / or algorithm may be used. The machine learning system and / or algorithm (e.g., DNNs such as VCG, ResNet, etc.) may be trained to determine features of objects within a frame. These features may include object labels, object classifications, object boundaries, and other features. In some cases, the machine learning system and / or algorithm may be trained by feeding the neural network a number of frames or images with known object features (e.g., using supervised training with backpropagation). After the neural network has been sufficiently trained, it may determine features of a new frame (e.g., frame 103) input to the neural network during inference. Interpretive examples of neural networks and associated training are referenced herein. Figure 6 and Figure 7 It has been described.

[0049] In some examples, selection engine 102 may determine a mask comprising a sufficient number of pixels for accurate optical flow estimation (e.g., a minimum or near-minimum number of pixels). For example, the time and / or computational power required to perform optical flow estimation generally increase with the number of pixels involved in the estimation. Therefore, reducing the number of pixels within the mask may be advantageous. To reduce the number of pixels within the mask without significantly and / or undesirably degrading the quality of the optical flow estimation, selection engine 102 may determine the pixels most relevant and / or useful for the optical flow estimation. In some cases, the selection of a subset of pixels may be based on at least two factors. As an example of the first factor, selection engine 102 may perform selection based on regions of interest (or pixels, whether contiguous or not).

[0050] As an example of the second factor, selection engine 102 can perform selection based on the sufficiency of optical flow estimation, given certain estimation quality constraints.

[0051] Selection engine 102 may consider regions of interest (or pixels), estimation quality constraints, combinations thereof, and / or other factors when selecting a subset of pixels. In some examples, selection techniques may include difference fields between two frames, intensity smoothness (or discontinuity), object regions and / or boundaries, semantic regions (e.g., foreground objects versus scene background), semantic uncertainty, active attention regions, any combination thereof, and / or other properties.

[0052] For example, selection engine 102 can search for pixels corresponding to certain object features (e.g., pixels corresponding to the edges of moving objects). In another example, selection engine 102 can search for pixels within certain regions of a frame. For example, selection engine 102 can search for pixels within a frame region known to be foreground. In another example, selection engine 102 can select a subset of pixels that have an approximately uniform distribution throughout the frame. In some cases, selection engine 102 can search for a predetermined number of pixels. For example, selection engine 102 can determine the importance of all or a portion of the pixels in a frame and select a predetermined number or percentage of the most important pixels. In other cases, selection engine 102 can dynamically select the number of pixels within a mask. For example, selection engine 102 can select a relatively large number of pixels in a frame that includes a relatively large number of moving objects and / or relatively fast motion. In an illustrative example, selection engine 102 can use dynamic programming to determine the appropriate number of pixels for the mask. For example, dynamic programming algorithms can be used to solve one or more constrained optimization problems. In one example, given certain available computing resources (e.g., the number of operations per time interval, a certain allowed waiting time, a certain allowed power consumption, etc.), the selection engine 102 can use a dynamic programming algorithm to find the optimal allocation or scheduling of resources based on constraints to achieve a specific goal (e.g., to determine the appropriate number of pixels or masks).

[0053] The optical flow vector engine 104 of the optical flow estimation system 100 can determine the optical flow vector corresponding to the pixel identified by the mask. In some cases, the optical flow vector can indicate the direction and magnitude of pixel movement. For example, the optical flow vector can describe the displacement between the coordinates corresponding to the pixel position in an initial frame and the coordinates corresponding to the pixel position in a subsequent frame. The optical flow vector engine 104 can use any type or form of optical flow estimation technique to determine the pixel position in the subsequent frame. In an illustrative example, the optical flow vector engine 104 can use a differential motion estimation technique (e.g., Taylor series approximation) to determine the optical flow vector. Additionally or alternatively, the optical flow vector engine 104 can use a machine learning algorithm (e.g., a deep neural network) to determine the optical flow vector. In some cases, the machine learning algorithm used to determine the optical flow vector can be different from the machine learning algorithm used to select a subset of pixels for the mask.

[0054] In some cases, the optical flow vector engine 104 can determine the optical flow vector between each pixel in the frame identified by the mask (referred to as the source frame) and the corresponding pixel in a subsequent frame (referred to as the target frame). The optical flow map engine 106 of the optical flow estimation system 100 can generate the value of the optical flow map 108 of the target frame based on the optical flow vector. The mask value can indicate the displacement of a pixel within the subsequent target frame relative to the source frame. In some examples, the optical flow map engine 106 can use the optical flow vector determined for the pixels in the target frame identified by the mask to determine the value of the optical flow map 108 of the target frame. In some examples, the optical flow map engine 106 can estimate the values ​​of other pixels in the target frame based on the optical flow vector determined for the pixels identified by the mask. For example, a stream upsampling technique can be performed to estimate the values ​​of other pixels in the target frame. This stream upsampling technique is described in more detail below.

[0055] In some examples, the optical flow graph engine 106 can generate an incremental optical flow graph corresponding to motion estimation between two adjacent frames. In other examples, the optical flow graph engine 106 can generate a cumulative optical flow graph (in which case the optical flow graph is adjusted or updated in each frame) corresponding to motion estimation between two frames having one or more intermediate frames in between. For example, the optical flow graph engine 106 can determine an incremental optical flow graph between all or a portion of directly adjacent frames within a series of frames. The optical flow graph engine 106 can use the incremental optical flow graph to update the cumulative optical flow graph between the first frame in the series and the current frame in the series. To update the cumulative optical flow graph, the optical flow graph engine 106 can sum the incremental optical flow vector between the current frame and the previous frame with the corresponding optical flow vector of the cumulative optical flow graph.

[0056] The optical flow graph 108 output by the optical flow graph engine 106 can be used for various purposes and / or tasks. For example, as mentioned above, optical flow graphs can be used in various applications or systems, including computer vision systems, XR systems, data compression, image segmentation, autonomous vehicle operation, and other applications.

[0057] Figure 2A and Figure 2B The explanation can be made by Figure 1 A diagram illustrating an example of the optical flow estimation process performed by the optical flow estimation system 100. Figure 2A An example of the first frame 201 in the frame sequence has been explained. Frame 201 can correspond to... Figure 1 One of frames 103. Frame 201 is shown as having a dimension of w pixels wide multiplied by h pixels high (represented as w×h). Those skilled in the art will understand that frame 201 may include dimensions greater than... Figure 2AThe pixel positions described herein are far more numerous than those in the diagram. For example, frame 201 may include a 4K (or Ultra High Definition (UHD)) frame at a resolution of 3840×2160 pixels, an HD frame at a resolution of 1920×1080 pixels, or any other suitable frame with another resolution. Frame 201 includes pixels P1, P2, P3, P4, P5, P6, and P7. As shown, pixel P1 has position 202A. Pixel position 202A may include a (w,h) pixel position (3,1) relative to the top-left pixel position (0,0).

[0058] In one example, pixels P1-P7 are selected to be included within the mask. For example... Figure 2A As shown, pixels within a mask can be adjacent to each other (e.g., in contact), or pixels can be isolated and / or separated from each other. For example, a moving object may correspond to a pixel region within a frame. However, the disclosed sparse optical flow estimation system can be able to perform optical flow estimation on an object without determining the optical flow vector for each pixel within the region. In some cases, a small number of pixels (e.g., 1 pixel, 3 pixels, etc.) may be sufficient for accurate optical flow estimation of a larger region. In the illustrative example, pixels P1, P2, and P3 may correspond to the tip of a person's nose, and pixel P5 may correspond to the boundary between the face and the frame background.

[0059] Figure 2B This is an illustration of an example of the second frame 203 in an interpretive frame sequence. The second frame 203 has corresponding pixel positions (with dimensions w × h) that are the same as those in the first frame 201. For example, the top-left pixel in the first frame 201 (at pixel location or position (0,0)) corresponds to the top-left pixel in the second frame 203 (at pixel location or position (0,0)). As shown, pixel P1 has moved from pixel location 202A in frame 201 to the updated pixel location 202B in frame 203. The updated pixel location 202B may include a (w,h) pixel location (4,2) relative to the top-leftmost pixel location (0,0). An optical flow vector can be calculated for pixel P1, indicating the velocity or optical flow of pixel P1 from the first frame 201 to the second frame 203. In an interpretive example, the optical flow vector of pixel P1 between frames 201 and 203 is (1,1), indicating that pixel P1 has moved one pixel position to the right and one pixel position down. In some cases, the optical flow estimation system 100 can determine the optical flow vectors of the remaining pixels P2-P7 of the mask in a similar manner (while not determining the optical flow vectors of pixels not included in the mask).

[0060] Figure 3 This is a diagram illustrating an example of an optical flow estimation system 300. In some cases, all or part of the optical flow estimation system 300 may correspond to and / or be included in... Figure 1Within the optical flow estimation system 100. For example, the engines of the optical flow estimation system 300 (e.g., masking engine 302, feature extraction engine 304, feature sampling engine 306, correlation engine 308, optical flow calculation engine 310, and upsampling engine 312) can be configured to perform all or some of the functions and / or any additional functions performed by the engines of the optical flow estimation system 100. As will be explained in more detail below, the optical flow estimation system 300 can perform functions optimized for sparse optical flow estimation.

[0061] like Figure 3 As shown, the optical flow estimation system 300 can receive source frame I S and target frame I T In one example, source frame I S Indicates in target frame I T Previously received frames. For example, source frame I. S Can be used with target frame I within the frame sequence T Directly adjacent. Source frame I S and target frame I T This can be input into the masking engine 302 and / or the feature extraction engine 304. For example... Figure 3 As shown, source frame I S and target frame I T They can be cascaded or otherwise combined before being passed to the masking engine 302 and / or the feature extraction engine 304.

[0062] The masking engine 302 can perform mask generation to facilitate pixel and / or frame selection. For example, the masking engine 302 can generate one or more masks ( Figure 3 The masks are shown as masks A, B, C, D, and E. The optical flow estimation system 300 can apply one or more of these masks along the optical flow path. Figure 3 The optical flow estimation system 300 shown represents one or more blocks of the entire optical flow estimation model. For example, in some cases, masks AE can be used in each block (e.g., mask A is used by feature extraction engine 304, mask B by feature sampling engine 306, mask C by correlation engine 308, mask D by optical flow calculation engine 310, and mask E by upsampling engine 312). In some cases, only a subset of masks AE can be used by one or more blocks (e.g., only mask A is used by feature extraction engine 304).

[0063] In some examples, the one or more masks can spatially indicate one or more regions of discrete or continuous points as originating from source frame I. S and / or target frame I TA subset of all points. In some examples, the one or more masks may temporally indicate multiple levels of frames associated with multiple computational levels. In some cases, spatial examples of masked pixels may include associations with object boundaries, smoothness, discontinuities, confidence, and / or information. Temporal examples of masked frames may include the level of importance of the frame in terms of its motion and / or content, etc. In some cases, masking may also be performed using dynamic programming (as described above), whereby the selection of importance and / or relevance as estimated over the entire stream can be derived or predicted.

[0064] In some scenarios, to support sparse and / or multi-level (in terms of importance) flow estimation, feature extraction engine 304 may perform feature extraction. For each applicable pixel p (e.g., identified by mask A), this feature extraction not only extracts features of p but also information related to p in terms of spatial and / or temporal relevance and (expected) contribution. Relevant information can be obtained from information (e.g., features) associated with neighboring pixels surrounding or adjacent to pixel p, or even non-local pixels relative to pixel p. Relative information may also include some spatial and / or temporal contextual information relative to pixel p. Such information relative to each instance of pixel p can be fused for each pixel p. The goal of feature extraction engine 304 is to extract and collect as much desired information (e.g., features) as possible for each applicable p that will be utilized in subsequent processing stages. For example, for each applicable pixel p, feature extraction engine 304 may extract and collect features and fuse these features as described below.

[0065] For example, feature extraction engine 304 can determine the relationship with source frame I. S and / or target frame I T The contextual features associated with a pixel. In one example, the contextual features associated with a pixel may include feature vectors extracted from the frame using a machine learning system and / or algorithm. An example of a machine learning system and / or algorithm that can be used is a deep neural network trained for feature extraction. An illustrative example of a deep neural network is referenced below. Figure 6 and Figure 7The feature vector can indicate features such as a pixel's label or classification, visual attributes and / or characteristics, semantic features, and other features. In some cases, the feature vector may include information related to the spatial characteristics of the pixel. Spatial characteristics may include the pixel's relationship to object boundaries, the pixel's smoothness, discontinuities associated with the pixel, and other characteristics. In some cases, spatial characteristics may include spatial confidence associated with the pixel's importance and / or relevance to the overall optical flow estimation. For example, a pixel with high spatial confidence may be highly important and / or relevant to the optical flow estimation (e.g., high motion). In some cases, the feature vector may include information related to the pixel's temporal characteristics. In some cases, the temporal characteristics of the pixel may include one or more characteristics associated with the pixel's motion, including motion velocity, motion acceleration, and other characteristics. In one example, temporal characteristics may include confidence associated with the importance and / or relevance of the pixel's motion to the overall optical flow estimation. For example, a pixel with high temporal confidence may be highly important and / or relevant to the optical flow estimation.

[0066] In some cases, feature extraction engine 304 can determine multi-scale contextual features associated with a frame. Multi-scale contextual features can include features associated with the frame at various scales (e.g., resolution). For example, feature extraction engine 304 can determine contextual features associated with a high-scale (e.g., full-resolution) version of the frame. Additionally or alternatively, feature extraction engine 304 can determine contextual features associated with one or more lower-scale (e.g., reduced-resolution) versions of the frame. In some cases, contextual features associated with different scales can be used at different steps in the optical flow estimation process. For example, utilizing low-scale feature vectors can improve the efficiency of some optical flow estimation steps, while utilizing high-scale feature vectors can improve the quality and / or accuracy of other optical flow estimation steps.

[0067] In some cases, the contextual features associated with a pixel may include contextual features associated with pixels surrounding and / or near that pixel, as mentioned above. For example, each pixel in a frame may represent a central pixel surrounded by one or more neighboring pixels. In one example, a neighboring pixel may refer to any pixel directly adjacent to (e.g., horizontally, vertically, and / or diagonally) the central pixel. In other examples, a neighboring pixel may refer to a pixel spaced from the central pixel by no more than a threshold distance or a threshold number (e.g., 2 pixels, 3 pixels, etc.). In a further example, a neighboring pixel may be a pixel with a high spatial and / or temporal association with the pixel. These pixels may be adjacent to the central pixel or not adjacent to the central pixel (e.g., not local). Feature extraction engine 304 may determine the contextual features of any number of neighboring pixels associated with the central pixel. For example, feature extraction engine 304 may extract and collect as many contextual features as possible required for one or more steps of the optical flow estimation process (explained in more detail below). Feature extraction engine 304 may also associate the contextual features of neighboring pixels with the central pixel. For example, feature extraction engine 304 can concatenate, group, and / or otherwise store the context features of neighboring pixels within a data structure associated with the center pixel, in conjunction with the context features of the center pixel. This data structure may include an index corresponding to the coordinates of the center pixel. In one example, feature extraction engine 304 can fuse the context features associated with each relevant neighboring pixel using weighting, summation, concatenation, and / or other techniques. For example, feature extraction engine 304 can use formula f... p,i Let i ∈ {0, 1, ..., C-1}, C ∈ R, to determine the fused contextual features so that feature f can be derived from pixel p. p,i .

[0068] By associating the contextual features of neighboring pixels with the contextual features of the center pixel, the feature extraction engine 304 can improve the accuracy of sparse optical flow estimation. For example, by determining and storing the contextual features of neighboring pixels in conjunction with the center pixel, the feature extraction engine 304 can help the optical flow estimation system 300 accurately identify the pixel corresponding to the center pixel in subsequent frames. Since sparse optical flow estimation may not involve a one-to-one mapping between each pixel in the initial frame and each pixel in subsequent frames, the contextual information associated with neighboring pixels can help the optical flow estimation system 300 accurately select the corresponding pixel from multiple candidate pixels.

[0069] Feature extraction engine 304 can determine source frame I S and target frame I T The contextual features of all or a portion of the pixels. In one example, feature extraction engine 304 can determine the source frame I. S and target frame IT The contextual features of the pixels corresponding to the mask (e.g., mask A provided by masking engine 302). For example, feature extraction engine 304 can apply mask A to determine the source frame I to process. S Which pixels are used to extract contextual features and / or to process data from the target frame I? T Which pixels are selected to extract contextual features. For example, feature extraction engine 304 can determine the contextual features of only the pixels identified by mask A. In some cases, the pixels identified by mask A may be selected, at least in part, based on machine learning algorithms trained to extract important and / or key features.

[0070] In some cases, feature sampling engine 306 may receive features extracted by feature extraction engine 304 (e.g., represented by one or more feature vectors). Feature sampling engine 306 may perform sampling operations and / or regrouping operations on the sampled points of the features. For example, feature sampling engine 306 may retrieve and / or group feature vectors (or feature sampled points in feature vectors) to facilitate subsequent processing stages. In some cases, feature sampling engine 306 may sample feature vectors based on a mask (e.g., mask B). For example, feature sampling engine 306 may apply mask B such that only a subset of pixels labeled by mask B is designated as relevant for further computation. In such examples, regions and / or pixels not labeled by mask B may be ignored. In some cases, mask B may be similar to or the same as mask A. In other cases, mask B may be configured to be different from mask A. For example, mask B may identify a different subset of pixels compared to mask A.

[0071] The correlation engine 308 can receive sampled feature vectors from the feature sampling engine 306. The correlation engine 308 can perform correlation calculations on the sampled feature vectors. For example, using data from two input frames (source frame I...)... S and target frame I T Taking the output of the sampled feature map as input, the correlation engine 308 can compute pairwise correlations over several pairwise combinations (e.g., for all possible pairwise combinations). Each correlation parameter represents two features (one feature per frame (e.g., from source frame I)). S A feature and from target frame I T The correlation between features or similarities in some cases. The correlation quantity determined by the correlation quantity engine 308 can be used (e.g., by the optical flow calculation engine 310) as input for subsequent optical flow estimation. In an illustrative example, based on mask M s and M t The pixel collection (e.g., a tensor including data) can each have H s W s C and Ht W t The dimension or shape of C, where the mask M s It is masking engine 302 targeting source frame I S The generated mask, M t It is the masking engine 302 targeting the target frame I T The generated mask, where H represents height, W represents width, and C represents the number of channels (or depth in some cases) in the neural network used for masking engine 302. In some examples, correlation engine 308 can use the following formula: To calculate the relevant quantities, where f s ,f t ∈R C These are for source frame I respectively S and target frame I T The correlation quantity engine 308 can process feature vectors based on an optical flow mask (e.g., mask C). For example, the correlation quantity engine can select or determine correlations labeled only by mask C (in this case, the remaining correlation quantities not labeled by mask C are discarded). In some cases, mask C can be similar to or the same as mask A and / or mask B. In other cases, mask C can be configured to be different from mask A and / or mask B.

[0072] Optical flow calculation engine 310 can receive correlation calculations from correlation engine 308. Optical flow calculation engine 310 can use features from the correlation calculations to perform point-by-point (e.g., per-pixel) optical flow estimation on pixels identified in a mask (e.g., mask A, mask D, or other masks). In some cases, optical flow calculation engine 310 can use one or more neural network operations (e.g., one or more convolutional layers, one or more residual convolutional blocks, and / or other network operations) to refine and / or adjust the optical flow estimation. For example, optical flow calculation engine 310 can determine the optical flow estimation of a specific feature vector based on one or more masks (e.g., mask D). In one example, optical flow calculation engine 310 can perform optical flow estimation to determine the optical flow vector of a pixel or pixel region identified by mask D. In some cases, mask D can be similar to or the same as masks A, B, and / or mask C. In other cases, mask D can be configured to be different from masks A, B, and / or mask C. In some examples, corresponding to source frame I... S and target frame I T The features can have the same as the source frame I S and target frame I T Same resolution.

[0073] Upsampling engine 312 can receive optical flow estimates from optical flow calculation engine 310. Upsampling engine 312 can perform one or more upsampling operations on the optical flow estimates to increase the resolution of the estimated optical flow field. In some cases, upsampling engine 312 can select specific upsampling operations based on the pixel format and / or distribution within the masks(s) used for optical flow estimation. For example, upsampling engine 312 can perform a first upsampling operation on a mask comprising a region of consecutive pixels, a second upsampling operation on a mask comprising discrete (e.g., non-contiguous) pixels, a third upsampling operation on a mask comprising each (or approximately each) pixel within a frame, and / or perform any other upsampling operation.

[0074] As mentioned above, the feature extraction engine 304 can determine multi-scale contextual features associated with pixels in a frame. In some cases, the various steps of the optical flow estimation process can utilize contextual features at different scales. For example, the optical flow calculation engine 310 can utilize extracted features in the form of multi-scale feature pyramids, cascaded and / or fused features with one or more scales, or other feature combinations. In some examples, the optical flow calculation engine 310 can select a mask for a particular step based on the scale of the contextual features used at that step. For example, a mask with low resolution can be used in conjunction with low-resolution feature vectors. Additionally or alternatively, the optical flow calculation engine 310 can select a mask for a particular step based on the operations(s) performed at that step. For example, feature sampling performed by the feature sampling engine 306 can be optimized using a first mask, while correlation calculation determined by the correlation engine 308 can be optimized using a second mask. Thus, in some cases, Figure 3 The masks A, B, C, D, and / or E described herein may include variations and / or differences. However, two or more masks may be identical or similar. Furthermore, steps within the optical flow estimation process may utilize multiple masks or not utilize masks at all.

[0075] By determining a subset of pixels corresponding to key features within a frame, the disclosed optical flow estimation techniques and systems can optimize sparse optical flow estimation based on the spatial characteristics of the frame. In some cases, the disclosed techniques and systems can optimize sparse optical flow estimation based on the temporal characteristics of a series of frames. For example, as mentioned above, optical flow estimation system 300 can determine optical flow estimates between two directly adjacent frames and / or between two frames separated by one or more intermediate frames. In some cases, optical flow estimation system 300 can selectively perform optical flow estimation on specific frames. For example, optical flow estimation system 300 can determine the importance values ​​of all or a portion of the frames within a series of frames. Frames with importance values ​​that meet or exceed a threshold importance value can be designated as key frames for which optical flow estimation is to be performed. In some cases, optical flow estimation system 300 can determine the importance of a frame based on features within the frame. For example, a frame with a large number of important features (e.g., key features) may have a high importance value. In another example, a frame with relatively low motion relative to previous frames may have a low importance value. The disclosed optical flow estimation system can optimize sparse optical flow estimation for any combination of spatial and / or temporal characteristics of one or more frames.

[0076] The described system and techniques provide masks to identify regions used for sparse optical flow estimation (rather than performing dense homogeneous estimation). For example, regions not involving motion (e.g., the background of an image from a still camera) can be skipped when performing optical flow estimation, resulting in reduced computational cost and higher estimation accuracy. This solution can be beneficial in many different fields, such as adaptive regional video source compression as an example. In some cases, more complex optical flow estimation may be required for selected regions or pixels that can be determined or identified using intelligent masks for flow refinement. For example, intelligent masking may include methods based on deep learning neural networks to determine one or more masks. In an example, an intelligent masking engine (e.g., masking engine 302) may determine one or more masks based on input features (or frames), the content of the features and / or frames, the activity associated with the features and / or frames, the semantics associated with the features and / or frames, consistency, and other factors. As described herein, the mask can be used to assist subsequent stages of optical flow processing (e.g., by one or more blocks or engines of the optical flow estimation system 300). As mentioned above, masking of region or frame selection can be made in spatial and / or temporal dimensions, thereby affecting the percentage of content available for refinement or the resolution level. In some cases, streaming network design strategies can fuse multi-scale and contextual features at the beginning of the neural network, prepare relevant quantities in the middle of the network, and subsequently allocate refinement in later or later parts of the network, such as... Figure 3 As shown in the image.

[0077] Figure 4A and Figure 4B Examples are given on optimizing and / or adjusting the sparsity of optical flow estimation data based on motion quantities associated with one or more frames. For example, Figure 4A Table 402 is explained, indicating the frame regions for which optical flow estimation should be performed based on intra-frame motion. In this example, the frame includes a ball in the foreground and the background behind the ball. When the ball is not moving and there is no motion in the background, the optical flow estimation system 100 does not perform optical flow estimation on the frame. When the ball is moving, the optical flow estimation system 100 can perform optical flow estimation on the pixels corresponding to the ball. When the ball is moving and there is motion in the background, the optical flow estimation system 100 can perform optical flow estimation on the entire frame (or a large portion of the frame, such as 80%, 90%, etc.).

[0078] Figure 4B Table 404 is explained, indicating the percentage of pixels within a frame for which optical flow estimation is to be performed based on motion within that frame. For example... Figure 4A and Figure 4B As shown, the size and / or position of the mask implemented by the disclosed optical flow estimation system can be dynamically adjusted and / or optimized based on the current and / or detected motion level within the received frame. For example, when the ball is not moving and there is no motion in the background (according to...) Figure 4A In the first line, the optical flow estimation system 100 performs optical flow estimation for 0% of the frames. When the ball is in motion, the optical flow estimation system 100 can perform optical flow estimation for 50% of the pixels (corresponding to the ball). In one example, when the ball is in motion and there is motion in the background, the optical flow estimation system 100 can perform optical flow estimation for 100% of the frames.

[0079] Figure 5 This is a flowchart illustrating an example of a process 500 for optical flow estimation between one or more frames using one or more techniques described herein. In block 502, process 500 includes: determining a subset of pixels in at least one of a first frame and a second frame. In block 504, process 500 includes: generating a mask indicating the subset of pixels. In some examples, process 500 may include: using a machine learning algorithm to determine a subset of pixels in at least the first and second frames. For example, optical flow estimation system 300 may include a machine learning system and / or algorithm that can determine the subset of pixels (e.g., using masking engine 302).

[0080] In some examples, process 500 may include: identifying a first pixel in the first frame corresponding to a first region of interest (ROI), and identifying a second pixel in the first frame corresponding to a second ROI. In some examples, the first and / or second ROI may be or may include one or more objects (e.g., a car or other vehicle, a person, a building, etc.), a portion of an object (e.g., a car wheel, a person's arm or hand, a sign on a building, etc.), and / or other portions of the scene depicted in the first frame. In some examples, the first ROI may include a moving ROI (e.g., a moving object), and the second ROI may include a stationary object. In some examples, the first ROI may include an object of interest (e.g., a user's hand, regardless of whether the hand is moving in the first frame), and the second ROI may include portions of the scene of no interest (e.g., the scene background, etc.). In some examples, process 500 may include the first pixel (corresponding to the first ROI) within a subset of pixels and exclude (not include) the second pixel (corresponding to the second ROI) from the subset of pixels.

[0081] In some examples, process 500 may include determining a subset of pixels for at least the first and second frames by identifying at least one pixel corresponding to the boundary between the two objects. For example, pixels along the boundary may be included in the pixel subset.

[0082] In some examples, process 500 may include: determining at least a subset of pixels in the first frame and the second frame by sampling a predetermined number of pixels within the first frame.

[0083] In block 506, process 500 includes: determining one or more features associated with a subset of pixels in at least a first frame and a second frame based on a mask. In some examples, the subset of pixels in at least the first and second frames includes at least one pixel having adjacent pixels not included in the pixel subset. In such examples, process 500 may include: determining the one or more features based on the mask by determining feature information corresponding to adjacent pixels of the at least one pixel and storing the feature information corresponding to adjacent pixels in association with the at least one pixel. In some examples, process 500 may include: using the feature information corresponding to adjacent pixels to determine an optical flow vector between the at least one pixel and a corresponding pixel in the second frame.

[0084] In box 508, process 500 includes: determining optical flow vectors between a subset of pixels in the first frame and corresponding pixels in the second frame. In box 510, process 500 includes: using the optical flow vectors to generate an optical flow map for the second frame.

[0085] In some examples, process 500 includes: determining the importance value of frames within a frame sequence including a first frame and a second frame. Process 500 may include: selecting frames within the frame sequence for performing optical flow estimation based on the importance value. In some cases, process 500 includes: selecting a second frame for performing optical flow estimation based on determining that the importance value of the second frame exceeds (or is greater than) a threshold importance value.

[0086] In some examples, the processes described herein (e.g., process 500 and / or other processes described herein) may be computed by a computing device or apparatus (such as having...) Figure 8 The computing device (as shown in the computing device architecture 800) is used to execute the process. In one example, process 500 can be performed by a computing device with an implementation... Figure 3 The optical flow estimation system 300 shown herein is executed by a computing device architecture 800. In some examples, the computing device may include mobile devices (e.g., mobile phones, tablet computing devices, etc.), wearable devices, extended reality (XR) devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), personal computers, laptop computers, video servers, televisions, vehicles (or computing devices of vehicles), robotic devices, and / or any other computing device with the resource capability to perform the processes described herein (including process 500 and / or other processes described herein).

[0087] In some cases, the computing device or apparatus may include various components such as one or more input devices, one or more output devices, one or more processors (e.g., general-purpose processors, neural processing units (NPUs), digital signal processors (DSPs), graphics processors, and / or other processors), one or more microprocessors, one or more microcomputers, one or more transmitters, receivers, or combined transmitter-receiver units (e.g., referred to as transceivers), one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to transmit and / or receive data, any combination thereof, and / or other components. The network interface may be configured to transmit and / or receive Internet Protocol (IP)-based data or other types of data.

[0088] The components of a computing device can be implemented using a circuit system. For example, each component may include electronic circuitry or other electronic hardware (which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), neural processing unit (NPU), and / or other suitable electronic circuitry)) and / or may be implemented therewith, and / or may include computer software, firmware, or any combination thereof and / or may be implemented therewith to perform the various operations described herein.

[0089] Process 500 is interpreted as a logic flowchart, which represents a sequence of operations that can be implemented by hardware, computer instructions, or a combination thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that performs the described operation when executed by one or more processors. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of described operations can be combined and / or performed in parallel in any order to implement the process.

[0090] Additionally, the processes described herein (including process 500 and / or other processes described herein) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, executed by hardware or a combination thereof. As mentioned above, the code can be stored on a computer-readable or machine-readable storage medium, for example in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transient.

[0091] As described above, the optical flow estimation system and techniques described herein can be implemented using neural network-based machine learning systems. Illustrative examples of neural networks that can be used include one or more convolutional neural networks (CNNs), autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), gated recurrent units (GRUs), any combination thereof, and / or any other suitable neural network.

[0092] Figure 6This is an illustrative example of a deep learning neural network 600 that can be used by an object detector. Input layer 620 includes input data. In one illustrative example, input layer 620 may include data representing pixels of an input video frame. Neural network 600 includes multiple hidden layers 622a, 622b through 622n. Hidden layers 622a, 622b through 622n include "n" hidden layers, where "n" is an integer greater than or equal to one. The number of hidden layers can be as many as needed for a given application. Neural network 600 further includes providing the output obtained from the processing performed by hidden layers 622a, 622b through 622n. In one illustrative example, output layer 624 can provide a classification of objects in the input video frame. This classification may include categories identifying the type of object (e.g., person, dog, cat, or other object).

[0093] Neural network 600 is a multi-layered neural network with interconnected nodes. Each node can represent a piece of information. Information associated with a node is shared between different layers, and each layer retains information as it is processed. In some cases, neural network 600 may include a feedforward network, in which there is no feedback connection where the network's output is fed back to itself. In some cases, neural network 600 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read in.

[0094] Information can be exchanged between nodes via node-to-node interconnects between layers. A node in input layer 620 can activate a set of nodes in the first hidden layer 622a. For example, as shown, each input node in input layer 620 is connected to each node in the first hidden layer 622a. Nodes in hidden layers 622a, 622b, through 622n can transform the information of each input node by applying activation functions to this information. The information derived from this transformation can then be passed on and can activate nodes in the next hidden layer 622b, which can execute their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable function. The output of hidden layer 622b can then activate nodes in the next hidden layer, and so on. Finally, the output of hidden layer 622n can activate one or more nodes in output layer 624, where the output is provided. In some cases, although a node in neural network 600 (e.g., node 626) is shown as having multiple output lines, the node has a single output and all lines are shown as outputs from a node representing the same output value.

[0095] In some cases, each node or the interconnections between nodes can have weights derived from a set of parameters trained on the neural network 600. Once trained, the neural network 600 can be referred to as a trained neural network, which can be used to classify one or more objects. For example, interconnections between nodes can represent learned fragments of information related to the interconnected nodes. The interconnections can have tunable numerical weights (e.g., based on the training dataset), allowing the neural network 600 to adapt to the input and learn as it processes more and more data.

[0096] The neural network 600 is pre-trained to process features from the data in the input layer 620 using different hidden layers 622a, 622b to 622n, in order to provide an output through the output layer 624. In an example where the neural network 600 is used to identify objects in an image, the neural network 600 can be trained using training data that includes both images and labels. For example, training images can be input into the network, where each input image has a label indicating the category of one or more objects in each image (basically indicating to the network what these objects are and what features they have). In an interpretive example, the training images could include images of the number 2, in which case the label for the image could be [0 0 1 0 0 0 00 0 0].

[0097] In some cases, the neural network 600 can use a training process called backpropagation to adjust the weights of its nodes. Backpropagation includes forward pass, loss function, backward pass, and weight update. For each training iteration, forward pass, loss function, backward pass, and parameter update are performed. For each set of training images, this process can be repeated up to a certain number of iterations until the neural network 600 is trained well enough that the layer weights are accurately tuned.

[0098] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 600. The weights are initially randomized before training the neural network 600. The image may include, for example, a numerical array representing image pixels. Each number in the array may include a value from 0 to 255, describing the pixel intensity at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chrominance components, etc.).

[0099] For the first training iteration of a neural network 600, the output may include values ​​that do not give preference to any particular category because the weights are randomly selected during initialization. For example, if the output is a vector with probabilities that an object includes different categories, the probability values ​​for each different category may be equal or at least very similar (e.g., 0.1 probability value for each of ten possible categories). Using the initial weights, the neural network 600 cannot determine low-level features and therefore cannot accurately determine what the object's category might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined. An example of a loss function is Mean Squared Error (MSE). MSE is defined as... It is calculated by multiplying half of the actual answer (target) by the square of the predicted answer (output). The loss can be set to equal E. total The value of .

[0100] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the amount of loss so that the predicted output matches the training label. The Neural Network 600 can perform backpropagation by determining which inputs (weights) contribute most to the network's loss, and the weights can be adjusted to reduce and eventually minimize the loss.

[0101] The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute the most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be represented as... Where w represents the weight, w i Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate involves larger weight updates, while a lower value indicates smaller weight updates.

[0102] Neural networks 600 can include any suitable deep network. An example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and output layers. See below. Figure 6 An example of a CNN is described. The hidden layers of a CNN consist of a series of convolutional layers, non-linear layers, pooling layers (for downsampling), and fully connected layers. Neural networks can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), etc.

[0103] Figure 7This is an illustrative example of a Convolutional Neural Network 700 (CNN 700). The input layer 720 of the CNN 700 includes data representing an image. For example, this data could include a numerical array representing the pixels of the image, where each number in the array includes a value from 0 to 255, describing the pixel intensity at that location in the array. Using the previous example above, the array could include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luminance and two chroma components, etc.). The image can be processed by a convolutional hidden layer 722a, optionally a non-linear activation layer, a pooling hidden layer 722b, and a fully connected hidden layer 722c to obtain an output at the output layer 724. While in Figure 7 Only one of each hidden layer is shown in the diagram, but those skilled in the art will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in a CNN 700. As previously described, the output can indicate a single category of an object, or may include probabilities that best describe the category of an object in the image.

[0104] The first layer of CNN 700 is a convolutional hidden layer 722a. Convolutional hidden layer 722a analyzes the image data input to layer 720. Each node in convolutional hidden layer 722a is connected to a node (pixel) region of the input image called its receptive field. Convolutional hidden layer 722a can be thought of as one or more filters (each filter corresponding to a different activation or feature map), where each convolutional iteration of the filter is a node or neuron in convolutional hidden layer 722a. For example, the region of the input image covered by a filter at each convolutional iteration will be the receptive field of that filter. In an illustrative example, if the input image consists of a 28×28 array and each filter (and its corresponding receptive field) is a 5×5 array, then there will be 24×24 nodes in convolutional hidden layer 722a. Each connection between a node and its receptive field learns weights, and in some cases, learns an overall bias, such that each node learns to analyze its specific local receptive field in the input image. Each node in hidden layer 722a will have the same weights and biases (referred to as shared weights and shared biases). For example, the filter has a weight (numerical) array and the same depth as the input. For the video frame example, the filter depth would be 3 (based on the three color components of the input image). An interpretive example of the filter array size is 5×5×3, which corresponds to the size of the receptive field of the node.

[0105] The convolutional property of the convolutional hidden layer 722a is due to the fact that each node of the convolutional layer is applied to its corresponding receptive field. For example, the filters of the convolutional hidden layer 722a can start from the top left corner of the input image array and can convolve around the input image. As mentioned above, each convolutional iteration of the filter can be considered as a node or neuron of the convolutional hidden layer 722a. In each convolutional iteration, the value of the filter is multiplied by the corresponding number of original pixel values ​​of the image (e.g., a 5×5 filter array is multiplied by a 5×5 array of input pixel values ​​from the top left corner of the input image array). The multiplications from each convolutional iteration can be summed to obtain a sum for that iteration or node. Next, the process continues at the next position in the input image based on the receptive field of the next node in the convolutional hidden layer 722a. For example, the filter can move a stride amount to the next receptive field. The stride amount can be set to 1 or other suitable amounts. For example, if the stride amount is set to 1, the filter will move 1 pixel to the right in each convolutional iteration. Processing the filter at each unique location of the input produces a number representing the filter result for that location, thereby causing a sum value to be determined for each node of the convolutional hidden layer 722a.

[0106] The mapping from the input layer to the convolutional hidden layer 722a is called an activation map (or feature map). An activation map includes values ​​representing the filter results at each location of the input quantity for each node. Activation maps can include arrays containing various sums of values ​​obtained from each iteration of the filter on the input quantity. For example, if a 5×5 filter is applied to each pixel of a 28×28 input image (with a stride of 1), the activation map would consist of a 24×24 array. The convolutional hidden layer 722a can include several activation maps to identify multiple features in the image. Figure 7 The example shown includes three activation maps. Using three activation maps, the convolutional hidden layer 722a can detect three different types of features, each of which is detectable across the entire image.

[0107] In some examples, nonlinear hidden layers can be applied after convolutional hidden layer 722a. Nonlinear layers can be used to introduce nonlinearity into a system that is constantly calculating linear operations. An illustrative example of a nonlinear layer is the Rectified Linear Unit (ReLU) layer. A ReLU layer applies the function f(x) = max(0,x) to all values ​​in the input, which changes all negative activations to 0. ReLU can thus increase the nonlinearity of network 700 without affecting the receptive field of convolutional hidden layer 722a.

[0108] A pooling hidden layer 722b can be applied after the convolutional hidden layer 722a (and, when used, after the non-linear hidden layer). The pooling hidden layer 722b is used to simplify the information in the output of the convolutional hidden layer 722a. For example, the pooling hidden layer 722b can take each activation map output from the convolutional hidden layer 722a and use a pooling function to generate dense activation maps (or feature maps). Max pooling is an example of a function performed by the pooling hidden layer. Other forms of pooling functions (such as average pooling, L2 norm pooling, or other suitable pooling functions) can be used by the pooling hidden layer 722a. Pooling functions (e.g., max pooling filters, L2 norm filters, or other suitable pooling filters) are applied to each activation map included in the convolutional hidden layer 722a. Figure 7 In the example shown, three pooling filters are used for the three activation maps in the convolutional hidden layer 722a.

[0109] In some examples, max pooling can be used by applying a max pooling filter (e.g., of 2×2 size) with a step size (e.g., equal to the filter's dimension, such as a step size of 2) to the activation map output from the convolutional hidden layer 722a. The output from the max pooling filter includes the maximum number in each sub-region around which the filter convolves. For example, with a 2×2 filter, each unit in the pooling layer can summarize a region of 2×2 nodes from the previous layer (where each node is a value in the activation map). For instance, four values ​​(nodes) in the activation map will be analyzed by the 2×2 max pooling filter in each iteration of the filter, with the maximum of these four values ​​being output as the "maximum" value. If such a max pooling filter is applied to the activation filter from a convolutional hidden layer 722a of dimension 24×24 nodes, the output from the pooling hidden layer 722b will be an array of 12×12 nodes.

[0110] In some examples, L2 norm pooling filters can also be used. L2 norm pooling filters involve computing the square root of the sum of squares of the values ​​in a 2×2 region (or other suitable region) of the activation map (instead of computing the maximum value as in max pooling), and using the computed value as the output.

[0111] Intuitively, pooling functions (e.g., max pooling, L2 norm pooling, or other pooling functions) determine whether a given feature is found anywhere in a region of an image, discarding the exact location information. This can be done without affecting the feature detection results, because once a feature is found, its exact location is less important than its approximate location relative to other features. The benefit of max pooling (and other pooling methods) is that it pools far fewer features, thus reducing the number of parameters required in the later layers of the CNN700.

[0112] The final layer in the network is a fully connected layer, which connects each node from the pooling hidden layer 724b to each output node in the output layer 724. Using the example above, the input layer comprises 28×28 nodes encoding the pixel intensity of the input image, the convolutional hidden layer 722a comprises 3×24×24 hidden feature nodes based on applying a 5×5 local receptive field (for filtering) to three activation maps, and the pooling hidden layer 722b comprises a layer based on applying a max-pooling filter to a 2×2 region across each of the three feature maps. Extending this example, the output layer 724 may comprise ten output nodes. In such an example, each node of the 3×12×12 pooling hidden layer 724b is connected to each node of the output layer 724.

[0113] The fully connected layer 722c can take the output of the previous pooling hidden layer 722b (which should represent the activation maps of high-level features) and determine the features most relevant to a particular class. For example, the fully connected layer 722c can determine the high-level features most relevant to a particular class and may include weights (nodes) for the high-level features. The product between the weights of the fully connected layer 722c and the pooling hidden layer 722b can be computed to obtain the probabilities for different classes. For example, if CNN 700 is used to predict that an object in a video frame is a person, higher values ​​will appear in the activation maps representing the high-level features of a person (e.g., the presence of two legs, the face being at the top of the object, the two eyes being at the upper left and upper right of the face, the nose being in the middle of the face, the mouth being at the bottom of the face, and / or other features common to people).

[0114] In some examples, the output from output layer 724 may include an M-dimensional vector (M = 10 in the previous example), where M may include the number of categories from which the program must choose when classifying objects in an image. Other example outputs may also be provided. Each number in the N-dimensional vector can represent the probability that an object belongs to a certain category. In an illustrative example, if the 10-dimensional output vector representing objects in ten different categories is [0 0 0.05 0.8 0 0.15 0 0 00], then the vector indicates a 5% probability that the image is an object in the third category (e.g., a dog), an 80% probability that the image is an object in the fourth category (e.g., a person), and a 15% probability that the image is an object in the sixth category (e.g., a kangaroo). The probability for a category can be thought of as the confidence level that an object is part of that category.

[0115] Figure 8 These are illustrations illustrating examples of systems used to implement certain aspects of the techniques described in this paper. Specifically, Figure 8An example of a computing system 800 is described, which can be, for example, any computing device constituting an internal computing system, a remote computing system, a camera, or any component thereof, wherein the components of the system are in communication with each other using connection 805. Connection 805 can be a physical connection using a bus, or a direct connection to processor 810 (such as in a chipset architecture). Connection 805 can also be a virtual connection, a networking connection, or a logical connection.

[0116] In some examples, the computing system 800 is a distributed system, wherein the functions described herein can be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some examples, one or more system components described herein represent a plurality of such components, each performing some or all of the functions described for that component. In some cases, the components can be physical or virtual devices.

[0117] Example system 800 includes at least one processing unit (CPU or processor) 810 and connections 805 that couple various system components (including system memory 815, such as read-only memory (ROM) 820 and random access memory (RAM) 825) to processor 810. Computing system 800 may include a cache 812 of high-speed memory that is directly connected to, adjacent to, or integrated into processor 810.

[0118] Processor 810 may include any general-purpose processor and hardware or software services, such as services 832, 834, and 836 stored in storage device 830 and configured to control processor 810, as well as dedicated processors in which software instructions are incorporated into the actual processor design. Processor 810 may be a substantially self-contained computing system, containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0119] To enable user interaction, the computing system 800 includes input devices 845 that can represent any number of input mechanisms, such as microphones for voice, touchscreens for gesture or graphic input, keyboards, mice, motion input, voice input, etc. The computing system 800 may also include output devices 835, which can be one or more of several output mechanisms. In some instances, multimodal systems allow users to provide multiple types of input / output to communicate with the computing system 800. The computing system 800 may include a communication interface 840, which generally manages and controls user input and system output. The communication interface can perform or facilitate the reception and / or transmission of wired or wireless communications using wired and / or wireless transceivers, including utilizing audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, etc. Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs Wireless signal transmission Low Energy (BLE) wireless signal transmission Wireless signal transmission, including radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), microwave access global interoperability (WiMAX), infrared (IR) wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. The communication interface 840 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers used to determine the location of the computing system 800 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the US-based Global Positioning System (GPS), the Russian-based Global Navigation Satellite System (GLONASS), the Chinese-based BeiDou Navigation Satellite System (BDS), and the European-based Galileo GNSS. There are no limitations on operation on any particular hardware configuration, and therefore the underlying features can be easily replaced to obtain improved hardware or firmware configurations as they are developed.

[0120] Storage device 830 may be a non-volatile and / or non-transient and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as magnetic tape, flash memory cards, solid-state storage devices, digital versatile discs, cartridges, floppy disks, hard disks, magnetic tapes, magnetic stripes, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, CD-ROM discs, rewritable CD discs, DVD discs, Blu-ray discs (BDD discs), holographic discs, another optical medium, secure digital storage (SD) cards, micro-secure digital storage (microSD) cards, Memory Stick. Cards, smart card chips, EMV chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase-change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cassette, and / or combinations thereof.

[0121] Storage device 830 may include software services, servers, etc., which cause the system to perform functions when the code defining such software is executed by processor 810. In some examples, hardware services that perform specific functions may include software components stored in computer-readable media connected to necessary hardware components, such as processor 810, connection 805, output device 835, etc., to perform functions.

[0122] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transient media in which data can be stored and exclude transient electronic signals propagated via carrier waves and / or wirelessly or via wired connections. Examples of non-transient media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer-readable media may have code and / or machine-executable instructions stored thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted using any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0123] In some examples, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transient computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0124] Specific details are provided in the foregoing description to provide a thorough understanding of the examples provided herein. However, those skilled in the art will understand that these examples can be practiced without these specific details. For clarity, in some instances, the technology of the invention may be presented as comprising various functional blocks, including functional blocks containing devices, device components, steps or routines in methods implemented in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring these examples in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without the need for unnecessary detail to avoid confusing the examples.

[0125] The examples above can be described as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts can describe operations as sequential processes, many operations can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination corresponds to the function returning to the calling function or the main function.

[0126] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise available from a computer-readable medium. These instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. Parts of the computer resources used may be accessible via a network. The computer-executable instructions may be, for example, binary files, intermediate format instructions (such as assembly language), firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used during the methods according to the described examples, and / or information created include disks or optical discs, flash memory, USB devices provided with non-volatile memory, networked storage devices, etc.

[0127] Devices implementing the various processes and methods disclosed herein may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include: laptop devices, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mount devices, self-standing devices, etc. The functionality described herein may also be implemented using peripheral devices or plug-in cards. As a further example, such functionality may also be implemented on a circuit board within different chips or different processes executed on a single device.

[0128] Instructions, media for conveying these instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0129] In the foregoing description, aspects of this application have been described with reference to specific examples thereof, but those skilled in the art will recognize that this application is not limited thereto. Thus, although illustrative examples of this application have been described in detail herein, it is to be understood that the various inventive concepts may be implemented and employed in a variety of other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the foregoing applications may be used individually or in combination. Furthermore, the examples may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, this specification and the accompanying drawings should be considered illustrative rather than limiting. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative examples, the methods may be performed in a different order than described.

[0130] Those skilled in the art will appreciate that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this specification.

[0131] When the components are described as being “configured” to perform certain operations, such configurations can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits), or any combination thereof.

[0132] The phrase “coupled to” means that any component is physically connected directly or indirectly to another component, and / or that any component is in communication with another component directly or indirectly (e.g., connected to that other component via a wired or wireless connection and / or other suitable communication interface).

[0133] The language of the claims or other languages ​​that state "at least one" and / or "one or more" in a set indicate that one or more members of the set (in any combination) satisfy the claim. For example, the claim language stating "at least one of A and B" means A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language of "at least one" and / or "one or more" in a set does not limit the set to the items listed in that set. For example, the claim language stating "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0134] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps are described above in a generalized manner in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0135] The techniques described herein can also be implemented using electronic hardware, computer software, firmware, or any combination thereof. These techniques can be implemented using any of a variety of devices, such as general-purpose computers, wireless communication handsets, or multi-purpose integrated circuit devices, including applications in wireless communication handsets and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product and may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. These technologies may additionally or alternatively be implemented, at least in part, by computer-readable communication media carrying or conveying program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.

[0136] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuit systems. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor, but in alternatives, it may be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration. Accordingly, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video codec (CODEC).

[0137] Explanatory examples of this disclosure include:

[0138] Example 1: An apparatus for optical flow estimation between one or more frames. The apparatus includes: a memory configured to store data corresponding to the one or more frames; and a processor configured to: determine a subset of pixels in at least one of a first frame and a second frame; generate a mask indicating the subset of pixels; determine one or more features associated with the subset of pixels in at least the first frame and the second frame based on the mask; determine an optical flow vector between the subset of pixels in the first frame and corresponding pixels in the second frame; and use the optical flow vector to generate an optical flow map of the second frame.

[0139] Example 2: The apparatus of Example 1, wherein the processor is configured to use a machine learning algorithm to determine at least the subset of pixels in the first frame and the second frame.

[0140] Example 3: An apparatus as in either Example 1 or 2, wherein the processor is configured to determine the subset of pixels of at least the first frame and the second frame by: identifying a first pixel of the first frame corresponding to a first region of interest; identifying a second pixel of the first frame corresponding to a second region of interest; and including the first pixel in the subset of pixels and excluding the second pixel from the subset of pixels.

[0141] Example 4: An apparatus as in any of Examples 1 to 3, wherein the processor is configured to determine the subset of pixels of at least the first frame and the second frame by identifying at least one pixel corresponding to the boundary between the two objects.

[0142] Example 5: An apparatus as in any of Examples 1 to 4, wherein the processor is configured to determine the subset of pixels of at least the first frame and the second frame by sampling a predetermined number of pixels within the first frame.

[0143] Example 6: An apparatus as in any of Examples 1 to 5, wherein at least the subset of pixels in the first frame and the second frame includes at least one pixel having adjacent pixels not included in the subset of pixels.

[0144] Example 7: An apparatus as in Example 6, wherein the processor is configured to determine one or more features based on the mask by: determining feature information corresponding to adjacent pixels of the at least one pixel; and storing the feature information corresponding to the adjacent pixels in association with the at least one pixel.

[0145] Example 8: An apparatus as in Example 7, wherein the processor is configured to: use the feature information corresponding to the adjacent pixel to determine the optical flow vector between the at least one pixel and the corresponding pixel of the second frame.

[0146] Example 9: An apparatus as in any of Examples 1 to 8, wherein the processor is configured to: determine an importance value of a frame within a frame sequence including the first frame and the second frame; and select a frame within the frame sequence for performing optical flow estimation based on the importance value.

[0147] Example 10: The apparatus of Example 9, wherein the processor is configured to select the second frame for optical flow estimation based on determining that the importance value of the second frame exceeds a threshold importance value.

[0148] Example 11: A device as in any of Examples 1 to 10, wherein the processor includes a neural processing unit (NPU).

[0149] Example 12: A device as in any of Examples 1 to 11, wherein the device includes a mobile device.

[0150] Example 13: A device as in any of Examples 1 to 12, wherein the device includes an extended reality device.

[0151] Example 14: A device as described in any of Examples 1 to 13, further comprising a display.

[0152] Example 15: An apparatus as in any of Examples 1 to 14, wherein the apparatus includes a camera configured to capture one or more frames.

[0153] Example 16: A method for optical flow estimation between one or more frames. The method includes: determining a subset of pixels in at least one of a first frame and a second frame; generating a mask indicating the subset of pixels; determining one or more features associated with the subset of pixels in at least the first frame and the second frame based on the mask; determining an optical flow vector between the subset of pixels in the first frame and corresponding pixels in the second frame; and using the optical flow vector to generate an optical flow map of the second frame.

[0154] Example 17: The method of Example 16 further includes: using a machine learning algorithm to determine at least the subset of pixels in the first frame and the second frame.

[0155] Example 18: A method as in any of Examples 16 to 17, wherein determining the subset of pixels of at least the first frame and the second frame includes: identifying a first pixel of the first frame corresponding to a first region of interest; identifying a second pixel of the first frame corresponding to a second region of interest; and including the first pixel in the subset of pixels and excluding the second pixel from the subset of pixels.

[0156] Example 19: The method of any of Examples 16 to 18, wherein determining the subset of pixels of at least the first frame and the second frame includes: identifying at least one pixel corresponding to the boundary between the two objects.

[0157] Example 20: The method of any of Examples 16 to 19, wherein determining the subset of pixels of at least the first frame and the second frame includes: sampling a predetermined number of pixels in the first frame.

[0158] Example 21: The method of any of Examples 16 to 20, wherein at least the subset of pixels of the first frame and the second frame includes at least one pixel having adjacent pixels not included in the subset of pixels.

[0159] Example 22: The method of Example 21, wherein determining the one or more features based on the mask includes: determining feature information corresponding to adjacent pixels of the at least one pixel; and storing the feature information corresponding to the adjacent pixels in association with the at least one pixel.

[0160] Example 23: The method of Example 22 further includes: using the feature information corresponding to the adjacent pixel to determine the optical flow vector between the at least one pixel and the corresponding pixel of the second frame.

[0161] Example 24: The method of any of Examples 16 to 23 further includes: determining an importance value of a frame within a frame sequence including the first frame and the second frame; and selecting a frame within the frame sequence for performing optical flow estimation based on the importance value.

[0162] Example 25: The method of Example 24 further includes: selecting the second frame for optical flow estimation based on determining that the importance value of the second frame exceeds a threshold importance value.

[0163] Example 26: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform any of the operations described in any of Examples 16 through 25.

[0164] Example 27: An apparatus including means for performing any of the operations of Examples 16 to 25.

Claims

1. An apparatus for optical flow estimation between a plurality of frames, comprising: a memory configured to store data corresponding to the plurality of frames; and one or more processors configured to: determine a subset of pixels of at least one of a first frame and a second frame, wherein determining the subset of pixels includes identifying first pixels of the first frame corresponding to a first region of interest, identifying second pixels of the first frame corresponding to a second region of interest, and includes including the first pixels within the subset of pixels and excluding the second pixels from the subset of pixels; generate a mask indicating the subset of pixels; determine one or more features associated with the subset of pixels of at least the first frame and the second frame based on the mask; determine optical flow vectors between the subset of pixels of the first frame and corresponding pixels of the second frame; and generate an optical flow map of the second frame using the optical flow vectors. The one or more processors are configured to determine the subset of pixels of at least the first frame and the second frame using a machine learning algorithm.

2. The apparatus of claim 1, wherein, To determine the subset of pixels of at least the first frame and the second frame, the one or more processors are configured to identify at least one pixel corresponding to a boundary between two objects.

3. The apparatus of claim 1, wherein, To determine the subset of pixels of at least the first frame and the second frame, the one or more processors are configured to sample a predetermined number of pixels within the first frame.

4. The apparatus of claim 1, wherein, The subset of pixels of at least the first frame and the second frame includes at least one pixel having an adjacent pixel that is not included in the subset of pixels.

5. The apparatus of claim 1, wherein, To determine the one or more features based on the mask, the one or more processors are configured to:

6. The apparatus of claim 5, wherein, determine feature information corresponding to the adjacent pixel of the at least one pixel; and store the feature information corresponding to the adjacent pixel in association with the at least one pixel. The one or more processors are configured to determine an optical flow vector between the at least one pixel and a corresponding pixel of the second frame using the feature information corresponding to the adjacent pixel.

7. The apparatus of claim 6, wherein, The one or more processors are configured to:

8. The apparatus of claim 1, wherein, determine importance values for frames within a sequence of frames including the first frame and the second frame; and select a frame within the sequence of frames for performing optical flow estimation based on the importance values. The one or more processors are configured to select the second frame for performing optical flow estimation based on determining that an importance value of the second frame exceeds a threshold importance value.

9. The apparatus of claim 8, wherein, The one or more processors include a neural processing unit (NPU).

10. The apparatus of claim 1, wherein, The apparatus includes a mobile device.

11. The apparatus of claim 1, wherein, The apparatus includes an extended reality device.

12. The apparatus of claim 1, wherein, 13. The apparatus of claim 1, further comprising a display. The apparatus includes a camera configured to capture one or more frames.

14. The apparatus of claim 1, wherein, 15. A method of optical flow estimation between a plurality of frames, the method comprising: ​ determining a subset of pixels of at least one of the first frame and the second frame, wherein determining the subset of pixels includes identifying first pixels of the first frame corresponding to a first region of interest, identifying second pixels of the first frame corresponding to a second region of interest, and including the first pixels within the subset of pixels and excluding the second pixels from the subset of pixels; generating a mask indicating the subset of pixels; determining one or more features associated with the subset of pixels of at least the first frame and the second frame based on the mask; determining an optical flow vector between the subset of pixels of the first frame and a corresponding pixel of the second frame; and generating an optical flow map of the second frame using the optical flow vector. The machine learning algorithm is trained using a plurality of frames of a video sequence.

16. The method of claim 15, further comprising: Determining the subset of pixels of at least the first frame and the second frame includes identifying at least one pixel corresponding to a boundary between two objects.

17. The method of claim 15, wherein, Determining the subset of pixels of at least the first frame and the second frame includes sampling a predetermined number of pixels within the first frame.

18. The method of claim 15, wherein, The subset of pixels of at least the first frame and the second frame includes at least one pixel having an adjacent pixel that is not included in the subset of pixels.

19. The method of claim 15, wherein, Determining the one or more features based on the mask includes:

20. The method of claim 19, wherein, determining feature information corresponding to the adjacent pixel of the at least one pixel; and storing the feature information corresponding to the adjacent pixel in association with the at least one pixel. Determining an optical flow vector between the at least one pixel and a corresponding pixel of the second frame using the feature information corresponding to the adjacent pixel.

21. The method of claim 20, further comprising:

22. The method of claim 15, further comprising: determining an importance value for a frame within a sequence of frames including the first frame and the second frame; and selecting a frame within the sequence of frames for performing optical flow estimation based on the importance value. The second frame is selected for performing optical flow estimation based on a determination that an importance value of the second frame exceeds a threshold importance value.

23. The method of claim 22, further comprising:

24. A computer-readable storage medium having instructions stored therein that, when executed by one or more processors, cause the one or more processors to: determine a subset of pixels of at least one of a first frame and a second frame, wherein determining the subset of pixels includes identifying first pixels of the first frame corresponding to a first region of interest, identifying second pixels of the first frame corresponding to a second region of interest, and including the first pixels within the subset of pixels and excluding the second pixels from the subset of pixels; generate a mask indicating the subset of pixels; determine one or more features associated with the subset of pixels of at least the first frame and the second frame based on the mask; determine an optical flow vector between the subset of pixels of the first frame and a corresponding pixel of the second frame; and generate an optical flow map of the second frame using the optical flow vector. ​ ​ 25. The computer-readable storage medium of claim 24, wherein, The instructions, when executed by the one or more processors, cause the one or more processors to determine the subset of pixels of at least the first frame and the second frame using a machine learning algorithm.

26. The computer-readable storage medium of claim 24, wherein, To determine the subset of pixels of at least the first frame and the second frame, the instructions, when executed by the one or more processors, cause the one or more processors to identify at least one pixel corresponding to a boundary between two objects.

27. The computer-readable storage medium of claim 24, wherein, To determine the subset of pixels of at least the first frame and the second frame, the instructions, when executed by the one or more processors, cause the one or more processors to sample a predetermined number of pixels within the first frame.

28. An apparatus for optical flow estimation between a plurality of frames, comprising: a memory configured to store data corresponding to the plurality of frames; and one or more processors configured to: determine a subset of pixels of at least one of a first frame and a second frame, the subset of pixels corresponding to key features within at least one of the first frame and the second frame, wherein the key features include pixels within a first region of interest of the first frame but do not include pixels within a second region of interest of the first frame; generate a mask indicating the subset of pixels; determine one or more contextual features associated with the subset of pixels of at least the first frame and the second frame based on the mask; determine optical flow vectors between a subset of pixels of the first frame and corresponding pixels of the second frame based on the one or more contextual features; and generate an optical flow map for the second frame using the optical flow vectors.

29. The apparatus of claim 28, wherein the subset of pixels of at least the first frame and the second frame includes at least one pixel whose neighboring pixels are not included in the subset of pixels, and determining the one or more contextual features includes determining contextual feature information corresponding to the neighboring pixels of the at least one pixel and storing the contextual feature information corresponding to the neighboring pixels in association with the at least one pixel.

30. An apparatus for optical flow estimation between a plurality of frames, comprising: a memory configured to store data corresponding to the plurality of frames; and one or more processors configured to: determine a subset of pixels of at least one of a first frame and a second frame, wherein the subset of pixels of at least the first frame and the second frame includes at least one pixel whose neighboring pixels are not included in the subset of pixels; generate a mask indicating the subset of pixels; determine one or more features associated with the subset of pixels of at least the first frame and the second frame based on the mask, wherein determining the one or more features includes determining feature information corresponding to the neighboring pixels of the at least one pixel and storing the feature information corresponding to the neighboring pixels in association with the at least one pixel. determining an optical flow vector between a subset of pixels of the first frame and a corresponding pixel of the second frame; and generating an optical flow map for the second frame using the optical flow vector.

Citation Information

Patent Citations

  • Optical flow estimation for motion compensated prediction in video coding

    CN110741640A