Processing video data using incremental quantization
By using incremental quantization technology in video data processing, the differences between video frames are used to generate incremental convolutional output, which solves the problem of waste of computing resources caused by redundant data processing and achieves more efficient video data processing.
Patent Information
- Application Number
- CN202380076478.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-20
- Filing Date
- 2023-10-13
- Publication Date
- 2025-06-10
AI Technical Summary
When processing video data, the artificial neural network independently processes each frame of video, resulting in repeated processing of redundant data and waste of computing resources.
Using incremental quantization technology, a machine learning model generates a convolutional output based on the first frame, and uses the difference between the frame and subsequent frames to generate incremental convolutional output, combining the two to generate a new convolutional output, thereby reducing redundant data processing.
By reducing redundant data processing, the efficiency of video data processing is improved, processor cycles and memory utilization are reduced, and power consumption of training and inference operations is reduced.
Smart Images

Figure CN120129924A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims priority to U.S. Patent Application No. 18 / 338,184, filed on June 20, 2023, entitled "Processing Video Data Using Delta Quantization", which claims the benefit and priority of U.S. Provisional Patent Application S / N. 63 / 383,156, filed on November 10, 2022, entitled "Processing Video Data Using Delta Quantization" and assigned to the assignee of the present application. The entire contents of both applications are incorporated herein by reference.
[0003] Introduction
[0004] Aspects of the present disclosure relate to using machine learning models to process video content.
[0005] Artificial neural networks can be used to perform various operations on video content or other content that includes spatial and temporal components. For example, an artificial neural network can be used to compress video content into a smaller - sized representation to improve storage and transmission efficiency and match the intended use of the video content (e.g., an appropriate data resolution for the size of a device display). The compression of this content can be performed using lossy techniques such that the decompressed data version is an approximation of the original compressed data, or by using lossless techniques that result in a decompressed data version equivalent to the original data. In another example, an artificial neural network can be used to detect objects in video content. Object detection can include, for example, body pose estimation for identifying moving subjects in video content and predicting how the subject will move in the future, object classification for identifying objects of interest in video content, and so on.
[0006] Generally, the temporal component of video content can be represented by different frames in the video content. An artificial neural network can process the frames in the video content independently through each layer of the artificial neural network. Thereby, the cost of video processing by the artificial neural network can grow at a rate different from (and higher than) the rate of growth of the information in the video content. That is, between successive frames in the video content, there may be small changes between each frame because only a small amount of data may change during the amount of time elapsed between different frames. However, since the neural network generally processes each frame independently, the artificial neural network generally processes duplicate data (e.g., invariant parts of the scene) between the frames, which is very inefficient and results in a waste of computing resources (e.g., processor cycles, memory utilization, etc.) due to the repeated processing of the invariant data between different frames.
[0007] Brief Overview
[0008] The systems, methods, and devices of the present disclosure each have several aspects, and no single aspect alone is responsible for their desirable attributes. Without limiting the scope of the present disclosure as expressed in the appended claims, some features will now be briefly discussed. After considering this discussion, and particularly after reading the section entitled "Detailed Description," it will be understood how the features of the present disclosure provide the advantages as described herein.
[0009] Certain aspects provide a method for generating an inference based on quantization of the increment between different parts of an input stream. An example method generally includes: receiving image data including at least a first frame and a second frame; generating a first convolutional output based on the first frame using a machine learning model; generating a second convolutional output based on the difference between the first frame and the second frame using one or more quantizers of the machine learning model; generating a third convolutional output associated with the second frame according to a combination of the first convolutional output and the second convolutional output; and performing image processing based on the first convolutional output associated with the first frame and the third convolutional output associated with the second frame.
[0010] Other aspects provide: a processing system configured to execute the foregoing methods and those described herein; a non-transitory computer-readable medium including instructions that, when executed by one or more processors of the processing system, cause the processing system to execute the foregoing methods and the methods described herein; a computer program product implemented on a computer-readable storage medium that includes code for executing the foregoing methods and those further described herein; and a processing system including means for executing the foregoing methods and those further described herein.
[0011] To achieve the foregoing and related purposes, one or more aspects include the features that are fully described below and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of one or more aspects. However, these features are only indicative of some of the various ways in which the principles of the various aspects may be employed, and this description is intended to cover all such aspects and their equivalents. Brief Description of the Drawings
[0013] To understand in detail the manner in which the above-described features are used, the above-briefly summarized content may be described more specifically with reference to the aspects, some of which are illustrated in the drawings. It should be noted, however, that the drawings only illustrate certain typical aspects of the disclosure and should not be considered to limit its scope, as the description may admit of other equally effective aspects.
[0014] Figure 1 Illustrates example convolution operations involving different parts of a data stream based on the increment between different parts of the data stream in accordance with aspects of the present disclosure.
[0015] Figure 2 Illustrates examples of operations for convolving different parts of a data stream based on quantization of the increment between different parts of the data stream in accordance with aspects of the present disclosure.
[0016] Figure 3 Illustrates examples of conditional incremental quantization of different parts of a data stream based on the increment between different parts of the data stream in accordance with aspects of the present disclosure.
[0017] Figure 4 Illustrates examples of per-pixel conditional incremental quantization of different parts of a data stream based on the increment between different parts of the data stream in accordance with aspects of the present disclosure.
[0018] Figure 5 Illustrates example operations for incremental quantization of different parts of a data stream in accordance with aspects of the present disclosure.
[0019] Figure 6 Illustrates an example system in which aspects of the present disclosure may be implemented.
[0020] For the sake of facilitating understanding, where possible, the same reference numerals have been used to designate the same elements common to the various figures. It is contemplated that the elements and features of one aspect may be beneficially incorporated into other aspects without further recitation.
[0021] Detailed Description
[0022] Aspects of the present disclosure provide techniques and apparatus for efficiently processing video data (or other streaming data) based on quantifying the increments or differences between different portions of the video data (or other streaming data).
[0023] As discussed, artificial neural networks can be used to perform various inference operations on video content. These inferences can be used, for example, in various compression schemes, object detection, computer vision operations, various image processing and modification operations (e.g., upscaling, denoising, etc.), and so on. However, artificial neural networks can process each video frame in the video content independently. Thus, these artificial neural networks may not exploit the various redundancies across the frames in the video content.
[0024] Using a trained neural network to process streaming data can be a computationally complex task, and the computational complexity scales with the precision of the neural network. That is, training and using an accurate model may be more computationally complex, while training and using a less accurate model may be less computationally complex. To allow for improved computational complexity while retaining the accuracy of the artificial neural network, the redundancy of the data can be exploited. For example, channel redundancy can allow for pruning weights based on various error terms, quantization can be used to represent weights using a smaller bitwidth, and singular value decomposition can be used to approximate the weight matrix in a more compact representation. In another example, spatial redundancy can be used to exploit similarities in the spatial domain. In yet another example, knowledge distillation can be used, where a student neural network is trained to match the feature outputs of a teacher neural network. However, these techniques may not exploit the temporal redundancy in video content or other content that includes a temporal component.
[0025] Aspects of the present disclosure provide techniques and apparatus for generating inferences using a neural network or other machine learning model, exploiting the temporal redundancy in video content or other content that includes a temporal component. These temporal redundancies can be represented by the differences or increments between successive portions of the video content or other content with a temporal component. By performing inferences using a neural network based on the increments from successive portions of the video content or other content with a temporal component that are quantized to a lower precision, aspects of the present disclosure can reduce the amount of data used in performing inferences using the neural network. This can accelerate the process of performing inferences using the neural network, which can reduce the number of processing cycles and memory used for these operations, reduce the amount of power used for training and inference operations, and so on.
[0026] Example Incremental Quantization of Streaming Data in an Artificial Neural Network
[0027] Generally, when processing an input floating-point tensor x, the input tensor x can be quantized to a b-bit fixed-point tensor using a quantization function where Θ represents the quantizer parameters for quantizing the input tensor x. Various quantization functions can be used to quantize x into For example, using uniform affine symmetric quantization, the quantization function can be defined as follows:
[0028]
[0029] where the quantizer parameters Θ include a scaling factor s, and where the clamp function limits the rounded value of b-1 between a lower bound of -2 b-1 and an upper bound of 2
[0030] The floating-point convolution function z = w * x can generally be performed in fixed-point by quantizing the weights and the input tensor according to the above quantization function. The rounding and clamping operations introduce a quantization error defined as where a smaller quantization error corresponds to better inference performance.
[0031] In the case where the weights are quantized and the input is not quantized such that the quantization error can be expressed by the following expression:
[0032]
[0033] The quantization error of the weights and the magnitude of this error for the input x may affect the overall quantization error. It can be seen that due to the lower variance and magnitude of the residual values, convolving these values with quantized weights may exhibit a reduced quantization error compared to convolving other data. Similarly, in the case where the input is quantized and the weights are not quantized such that the quantization error can be expressed by the following expression:
[0034]
[0035] Assuming no clipping is performed, the quantization error can be expressed by the following expression:
[0036]
[0037] and the quantization error can be restricted to the values and Since the scaling factor s can be proportional to the magnitude and variance of the input x, quantizing the residual values with a smaller variance can reduce the quantization errors Δx and ∈ x .
[0038] Figure 1Illustrated is an example convolution operation 100 that convolves different parts of a data stream based on an increment between the different parts in accordance with aspects of the present disclosure.
[0039] As Figure 1 illustrated, the data stream includes a frame 102 at time t (represented by x t ) that is convolved with a kernel w, which can be a convolutional layer of a pre-trained neural network. Although Figure 1 illustrated with respect to video frames, it should be appreciated that the techniques described herein can be used to process other types of data having a temporal component. In some aspects, the data can include one or more spatial components, such as video content, where the spatial components include a height component, a width component, and a channel component (e.g., color channels such as red / green / blue channels / alpha channel, luminance channel, etc.). The convolution of x t with w produces an output z t .
[0040] The frame at time t can be represented as the sum of the previous frame at time t - 1 and the difference (e.g., an increment frame, which can also be referred to as a residual frame) between the frame at time t and the frame at time t - 1. Thus, the sum of the convolution of the frame at time t - 1 and the convolution of the difference between the frame at time t and the frame at time t - 1 produces the convolution of the frame at time t. Thus, the same output z t can be achieved by convolving the frame 104 at time t - 1 (represented by x t-1 ) with the same kernel w and adding it to the residual frame 106 convolved with the same kernel w, where the residual frame is defined as the difference between the current frame (e.g., frame 102) and the previous frame (e.g., frame 104). In other words, the current output of the convolution (e.g., the current frame) can be constructed by reusing the previous output of the convolution (e.g., the previous frame) and adding an update given by convolving the residual frame with the kernel. Thus, using the distributive property and the sigma-delta (Σ-Δ) rule, the convolution at frame t can be represented by:
[0041] z t = w * x t = w * (x t-1 + δ t ) = z t-1 + w * δ t
[0042] where δ tRepresents the residual frame 106, which may also be referred to as the delta frame. That is, since the Σ-Δ rule allows the frame at time t to be represented as the sum of the frame at time t-1 and the delta applied to that frame, the convolution at frame t can similarly be represented as the sum of the convolution of the frame at time t-1 and the convolution of the delta between the frame at time t-1 and the frame at time t.
[0043] The quantized convolution z in fixed-point representation generated using the quantization function discussed above can be represented by the following equation:
[0044]
[0045] where represents the quantized convolution of the key frame (or key input) x k generated according to the following equation:
[0046]
[0047] In the above equation, Φ w and Φ a represent the weights and quantization parameters of the key frame (or key input) and can be shared across each key frame in the input stream of video frames. Similarly, Θ w and Θ a represent the weights and quantization parameters of the residual frame, and these weights and quantization parameters can be shared across each residual frame or adjusted dynamically based on the content of these residual frames, as discussed in further detail below.
[0048] Figure 2 Illustrates examples of operations 200 and 250 for convolving different parts of a data stream based on quantization of the deltas between different parts of the data stream, according to aspects of the present disclosure.
[0049] Generally, using quantization reduces the computational cost and thus increases the speed of performing inference when using a trained network on a processor (such as a neural signal processor (NSP)). In some cases, when quantization is used for a model of a data stream (e.g., a video stream or other data streams with a time component), different frames can be quantized independently. When independently quantizing different frames or other data with a time component, the redundancy between these different frames may not be exploited because each frame is treated as if the data in that frame had not been processed previously, even though potentially a large number of successive frames may actually remain the same (or at least substantially the same).
[0050] In some aspects, such quantization at layer l of a neural network can be performed according to the following equation:
[0051]
[0052] However, independently quantifying data with a time component ignores temporal redundancy (e.g., due to processing each frame independently, including duplicate data between frames). Thus, independent processing of successive frames wastes computational resources (e.g., processor cycles, memory utilization, etc.) due to the repeated processing of invariant data between different frames.
[0053] Aspects of the present disclosure provide techniques and apparatus for using delta quantization to exploit temporal redundancy to improve quantization of video models. In some cases, the temporal redundancy between a first frame and a second frame implies a small delta between the first frame and the second frame. For example, in a video stream captured at a frame rate of 60 frames per second (FPS), each successive frame captures only the changes that occur in the last 1 / 60th of a second. In some cases, the deltas (e.g., the changes that occur between video frames) may be very small because many parts of the scene may be invariant between frames and because the amount of change that occurs between these frames may be minimal. These small deltas may result in small quantization errors / noise. The quantization error depends on the distribution (e.g., range) of the values. In some cases, weights and / or activations can be stored (e.g., quantized) with a lower bit precision than the precision with which these values were trained. For example, floating-point values can be quantized to a set of fixed (scaled) integer values. For example, a set of continuous floating-point values can be classified into discrete bins of integer values (e.g., based on proximity to the integer values). Performing quantization reduces the memory overhead of storing tensors and reduces the complexity and cost of operations (e.g., matrix multiplication).
[0054] Delta quantization of the input at a given layer l of a neural network can be performed according to the following equation:
[0055]
[0056] As Figure 2 illustrated, delta quantization can be performed at a first precision level (e.g., full precision) according to operation 200. Full precision can refer to a data type with a wide range of possible values, such as a 32-bit floating-point data type, which supports a value range between 1.2 * 10 -38 and 3.4 * 10 38 . For full-precision delta quantization, the convolution of an initial data element (e.g., an independent frame (I-frame) in video content, also referred to as a key frame) can be calculated by convolving the initial data element with a kernel 202 to produce an output as described above with reference to Figure 1 . In some aspects, a subtraction operation 204 can be performed to calculate the difference between the initial data element and a second data element (e.g., a P-frame, or predicted frame) The difference 208. The difference 208 is also convolved with the kernel 202. In some aspects, the subtraction operation 204 may be omitted (e.g., when the data stream is compressed such that the second data element encodes the difference information and renders based on a combination of the initial data element and the second data element). An addition operation 206 may be performed on the result 210 of the convolution and the output of the I-frame, thereby producing an output to produce an output This procedure can be generalized to the t-th frame, where the difference 212 between the current frame and the previous frame can be calculated and convolved with the kernel 202. The result 214 of the convolution can be added to the output of the previous frame, thereby producing an output to produce an output According to some aspects, the previous frame can be an I-frame or a previous P-frame.
[0057] As Figure 2 explained, the incremental quantization of the P-frame can be performed at a second precision level (e.g., lower precision) according to operation 250. The lower precision can refer to a data type having fewer bits than the number of bits defining the first precision level (e.g., less than the length of a 32-bit floating point number), such as a half-precision (16-bit) floating point number (float16) or a simpler data type, such as an integer data type of the same or smaller bit width (since integer arithmetic is computationally less complex than floating point arithmetic). As described in operation 250, an independent frame (I-frame) can be calculated by convolving the first frame with the kernel 202 to produce an output as described above with reference to Figure 1 In some aspects, a subtraction operation 204 can be performed to calculate the difference 254 between the first frame and the second frame which is convolved with the kernel 252 at a precision lower than that of the kernel 202. In some aspects, the subtraction operation can be omitted (e.g., when the data stream is compressed as discussed above). An addition operation 206 can be performed on the result 256 of the convolution and the output of the I-frame to produce an output This procedure can be generalized to the t-th frame, where the difference 258 between the current frame and the previous frame can be calculated and convolved with the kernel 252. The result 260 of the convolution can be added to the output of the previous frame, thereby producing an output to produce an output According to some aspects, the previous frame can be an I-frame or a previous P-frame.
[0058] In some aspects, when quantizing both key frames and residual frames (or P-frames and I-frames), a uniform affine symmetric scheme can be used. The scaling factor s can be based on the activation range of the quantizer, which is defined to have a lower bound r min and an upper bound r max :
[0059]
[0060] At any convolutional layer in a neural network for generating a convolution z of an input x, different range setters can be used to quantize the input x and the weights w. The range defined for the weights can be based on the minimum and maximum weights such that the quantized weights can be directly computed To estimate the range of activations of the input x, a set of example inputs can be collected based on calibration samples input into the model. The example inputs can be concatenated into a batch X, and the line search space between the minimum and maximum points in the batch X has r candidate points according to the following expression:
[0061]
[0062] The range in can be searched by minimizing or at least reducing an objective function:
[0063]
[0064] Figure 3 Illustrates an example operation 300 of conditional incremental quantization of different parts based on the increment between different parts of a data stream according to aspects of the present disclosure.
[0065] In a video stream, the distribution of the increments (e.g., relative to an I-frame) changes over time. For example, the magnitude of the increment relative to an I-frame can increase as the time from the I-frame increases. For example, in a video stream with 60 FPS, each subsequent frame only captures the changes that occur in the last 1 / 60 second, which can be very small. However, as the time from the initial frame increases, the changes relative to the initial frame can increase. This property (e.g., the correlation between the time from the I-frame and the magnitude of the increment relative to the I-frame) can be exploited to reduce computational costs. For example, in some cases, computational efficiency can be improved by using computationally less expensive (e.g., complex) operations and data types when the time from the I-frame is small and gradually transitioning to using more expensive operations and larger data types as the time from the I-frame increases.
[0066] Certain aspects of the present disclosure provide techniques for using a set of quantizers with different precision levels (e.g., 8-bit, 4-bit, or 2-bit precision) for specific increments. For example, during inference, a quantizer can be dynamically selected based on the input frame for a specific increment. In some cases, the selected quantizer can be used globally for the frame.
[0067] In some cases, the weights w can be quantized to w using a certain precision b In some cases, the I-frame activation is quantized to a certain precision b using post-training quantization (PTQ). a .
[0068] In some cases, different quantizers {Q ,...,Q 1 ,...,Q n} with different precisions can be customized for different magnitudes of the increments between successive frames. In other words, as described above, quantizers with different precisions can be used for different frames, and the quantizer can be selected based on the magnitude of the difference (e.g., increment) between different frames (which can be time-related to the initial frame, as described above).
[0069] In some cases, during calibration, each quantizer Q can be independently customized by minimizing (or at least reducing) the quantization noise on the calibration set δ according to the following formula i :
[0070]
[0071] where Q i represents the i-th quantizer, represents the i-th precision (e.g., associated with the i-th quantizer), the operator represents the quantizer with the minimum precision having an error of "just small enough" (as described in more detail below), δ represents the increment between frames, w represents the weight (e.g., the kernel of a convolutional layer of a pre-trained neural network), represents the quantizer with precision selected for the increment, and represents the weight w quantized to the i-th precision (e.g., precision ).
[0072] In some cases, during inference, the quantizer selection (e.g., which can be chosen to provide sufficient resolution for the data being quantized while minimizing or at least reducing the computational cost) can be a lower-precision quantizer with an error below a threshold amount. This quantizer selection can be understood with reference to the following formula:
[0073]
[0074] The error (∈ i ) of the i-th quantizer can be approximated according to the following formula:
[0075]
[0076] In some cases, to select the precision of the quantized data and thus the quantizer to be used for convolving the data among the n configured quantizers, the difference in error Π can be thresholded to the next quantizer, as shown in the following expression:
[0077] such that ∈ i -∈ i+1 <τ
[0078] In other words, if the difference in errors is less than the threshold τ, a lower-precision quantizer can be used to provide sufficient resolution for the data being quantized (e.g., as defined based on a comparison between the error and the threshold τ) while minimizing or at least reducing the computational cost.
[0079] As Figure 3 illustrated, the operation 300 for conditional incremental quantization can include: convolving the I-frame 302 with the kernel 202 at a dynamically selected precision using a dynamically selected quantizer to convolve different elements in the data stream, and in some aspects convolving different parts of each element in the data stream. That is, the I-frame can be quantized to b 1 bits, as illustrated.
[0080] In some aspects, a subtraction operation 204 can be performed to compute the difference (e.g., the residual frame) 304 between the I-frame and the next frame (e.g., the P-frame). In some aspects, this subtraction operation can be omitted (e.g., in the case where the data stream is compressed). In some aspects, for example, the residual frame can be pre-computed in the compressed data stream. After obtaining the residual frame (e.g., via the subtraction operation), the residual frame is convolved with the kernel 202 at a dynamically selected precision level, which can be lower than the precision level for convolving the first element (e.g., the I-frame 302). In this case, the residual frame is quantized to b 1 bits, as illustrated, and the convolution result is added to the output of the previous convolution.
[0081] Next, the difference between the previous frame (e.g., the I-frame or the previous P-frame) and the current frame can be computed (or obtained) to generate a residual frame 306, which is convolved at a dynamically selected precision level. As described above, in some aspects, the subtraction operation for computing the difference can be omitted (e.g., when the data stream is compressed). After obtaining the residual frame 306, the residual frame is convolved with the kernel 202 at a dynamically selected precision level. In this case, the residual frame 306 is quantized to b 2 bits, as illustrated, and the convolution result is added to the output of the previous convolution (after summation).
[0082] Next, the difference between a previous frame (e.g., an I-frame or a previous P-frame) and the current frame can be computed (or obtained) to produce a residual frame 308, which is convolved at a dynamically selected level of precision. As described above, in some aspects, the subtraction operation for computing the difference can be omitted (e.g., when the data stream is compressed). After obtaining the residual frame 308, the residual frame is convolved with the kernel 202 at a dynamically selected level of precision. In this case, the residual frame 308 is quantized to b 3 bits, as illustrated, and the convolution result is added to the output of the previous convolution (after summation).
[0083] Figure 4 An example operation 400 of per-pixel conditional delta quantization of different portions of a data stream based on deltas between the different portions in accordance with aspects of the present disclosure is illustrated.
[0084] In a video, different regions of a delta frame (e.g., different groups of one or more pixels) may carry different amounts of information. In other words, different portions of a video frame may have different levels of redundancy relative to a previous frame. For example, deltas in a foreground region (e.g., which may focus on a moving object) may be different from deltas in a background region (e.g., which may be time invariant), and deltas in a moving region may be different from deltas in a stationary region. That is, in some cases, the background region of a video frame may have greater redundancy than the foreground region of the video frame when compared to a previous video frame. In some cases, a set of quantizers 310 1 - 310 n can be used to quantize different regions of the delta frame. Some regions may require relatively high precision (e.g., a moving region). Some regions may require relatively low precision (e.g., a stationary region).
[0085] As Figure 4 illustrated, the operation 400 for per-pixel conditional delta quantization of individual frames having one or more pixel portions can include: convolving the I-frame 302 with the kernel 202 at a dynamically selected level of precision (as described above) using a dynamically selected quantizer to convolve different elements in the data stream, and in some aspects different portions of each element in the data stream. That is, the I-frame can be quantized to b 1 bits, as illustrated.
[0086] In some aspects, a subtraction operation 204 can be performed to calculate the difference between an I-frame and the next frame (e.g., a P-frame) (e.g., residual frame 402). In some aspects, this subtraction operation can be omitted (e.g., when the data stream is compressed). In some aspects, for example, the residual (e.g., delta) frame can be pre-computed in the compressed data stream. After obtaining the residual frame 402 (e.g., via the subtraction operation), various dynamically selected precision levels can be used to convolve the various parts of the residual frame with the kernel 202, and these dynamically selected precision levels can be lower than the precision level for convolving the first element (e.g., I-frame 302). In this case, part 402a of the residual frame 402 is quantized to b 2 bits, and part 402b of the residual frame 402 is quantized to b 3 bits, as illustrated. The convolution result can then be added to the output of the previous convolution via an addition operation 206.
[0087] Next, the difference between the previous frame (e.g., an I-frame or a previous P-frame) and the current frame can be calculated (or obtained) to produce a residual frame 404, and the various parts of this residual frame 404 can be convolved at various dynamically selected precision levels. As described above, in some aspects, the subtraction operation for calculating the difference can be omitted (e.g., when the data stream is compressed). As illustrated in this example, part 404a of the residual frame 404 is quantized to b 2 bits, part 404b of the residual frame is quantized to b 1 bits, and part 404c of the residual frame is quantized to b 3 bits, as illustrated. The convolution result can then be added to the output of the previous convolution.
[0088] Next, the difference between the previous frame (e.g., an I-frame or a previous P-frame) and the current frame can be calculated (or obtained) to produce a residual frame 406, and the various parts of this residual frame 406 can be convolved at various dynamically selected precision levels. As described above, in some aspects, the subtraction operation for calculating the difference can be omitted (e.g., when the data stream is compressed). As illustrated in this example, part 406a of the residual frame is quantized to b 2 bits, part 406b of the residual frame is quantized to b 1 bits, and part 406c of the residual frame is quantized to b 3 bits, as illustrated. The convolution result can then be added to the output of the previous convolution.
[0089] Example operations for differential quantization of video data
[0090] Figure 5 Illustrates an example of a method 500 for differential quantization for video processing according to aspects of the present disclosure. In some examples, method 500 can be performed by a computing device (such asFigure 6 executed by the processing system 600 as illustrated therein.
[0091] As illustrated, method 500 begins at block 505: receiving image data including at least a first frame and a second frame.
[0092] Method 600 then proceeds to block 510: generating a first convolutional output based on the first frame using a machine learning model.
[0093] Method 500 then proceeds to block 515: generating a second convolutional output based on the difference between the first frame and the second frame using one or more quantizers of the machine learning model.
[0094] Method 500 then proceeds to block 520: generating a third convolutional output associated with the second frame according to a combination of the first convolutional output associated with the first frame and the second convolutional output associated with the difference between the first frame and the second frame.
[0095] Method 500 then proceeds to block 525: performing image processing based on the first convolutional output associated with the first frame and the third convolutional output associated with the second frame. These operations may include, for example, encoding content into a latent space for compression, object detection in video content, subject pose estimation and motion prediction, semantic segmentation of video content into different segments, and so on.
[0096] In some aspects, the first convolutional output associated with the first frame is generated using a higher precision level than the precision level used to generate the second convolutional output based on the difference between the first frame and the second frame.
[0097] In some aspects, the higher precision level includes a floating-point representation of a given bit size. In some cases, the precision level used to generate the second convolutional output (based on the difference between the first frame and the second frame) may include an integer representation of a given bit size. In other cases, the precision level used to generate the second convolutional output may include an integer representation of a bit size smaller than the given bit size.
[0098] In some aspects, the first frame includes a key frame, which may be an initial frame that includes data for each spatial component in the initial frame of the video. Generally, the key frame may include uncompressed data for each pixel (or spatial location) in the initial frame. In this case, generating the first convolutional output (associated with the first frame) may involve using the full precision defined for the machine learning model to generate the output.
[0099] In some aspects, the second frame includes a delta frame that includes information about the difference between the first frame and the second frame. The size of the delta frame may be smaller than the key frame (or the first frame) because the delta frame does not need to include information shared between the key frame (or the first frame) and the delta frame.
[0100] In some aspects, generating a second convolutional output includes: selecting a quantizer from one or more quantizers of a machine learning model based on the magnitude of the difference between a first frame and a second frame, where each of the one or more quantizers is associated with a different level of precision; and quantizing the difference between the first frame and the second frame using the selected quantizer.
[0101] In some aspects, selecting a quantizer involves selecting a minimum quantizer, where the difference in quantization error between the minimum quantizer and a larger quantizer is less than a threshold.
[0102] In some aspects, generating a second convolutional output includes performing the following for each corresponding part of the first frame and each corresponding part of the second frame: selecting a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the corresponding part of the first frame and the corresponding part of the second frame, where each quantizer is associated with a different level of precision; and quantizing the difference between the corresponding part of the first frame and the corresponding part of the second frame using the selected quantizer.
[0103] In some aspects, the corresponding part of the first frame and the corresponding part of the second frame correspond to pixels (or a plurality of pixels) at the same location in the first frame and the second frame.
[0104] Example processing system for incremental quantization for video processing
[0105] Figure 6 Depicts an example processing system 600 for incremental quantization for video processing (such as, by way of example, as described herein with respect to Figure 5 ).
[0106] Processing system 600 includes a central processing unit (CPU) 602, which in some examples can be a multi-core CPU. Instructions executed at the CPU 602 can be loaded, for example, from a program memory associated with the CPU 602, or can be loaded from a memory 624.
[0107] Processing system 600 also includes additional processing components customized for specific functions, such as a graphics processing unit (GPU) 604, a digital signal processor (DSP) 606, a neural processing unit (NPU) 608, a multimedia processing unit 610, and a wireless connectivity component 612.
[0108] An NPU (such as NPU 608) is generally configured as a dedicated circuit for implementing the control and arithmetic logic for executing machine learning algorithms (such as algorithms for processing artificial neural networks (ANNs), deep neural networks (DNNs), random forests (RFs), etc.). An NPU may sometimes alternatively be referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligent processing unit (IPU), a vision processing unit (VPU), or a graphics processing unit.
[0109] The NPU (such as NPU 608) is configured to accelerate the execution of common machine learning tasks such as image classification, machine translation, object detection, and various other prediction models. In some examples, multiple NPUs may be instantiated on a single chip (such as a system-on-chip (SoC)), while in other examples, multiple NPUs may be part of a dedicated neural network accelerator.
[0110] The NPU can be optimized for training or inference, or in some cases configured to balance performance between the two. For NPUs capable of performing both training and inference, these two tasks can generally still be performed independently.
[0111] An NPU designed to accelerate training is generally configured to accelerate the optimization of new models, which is a highly computationally intensive operation involving inputting an existing dataset (usually labeled or tagged), iterating over the dataset, and then adjusting model parameters (such as weights and biases) to improve model performance. Generally speaking, optimizing based on incorrect predictions involves passing back through the layers of the model and determining gradients to reduce prediction errors.
[0112] An NPU designed to accelerate inference is generally configured to operate on a complete model. Such an NPU can thus be configured to: input a new data segment and quickly process the segment through the already trained model to generate a model output (e.g., an inference).
[0113] In some implementations, the NPU 608 is part of one or more of the CPU 602, GPU 604, and / or DSP 606.
[0114] In some examples, the wireless connectivity component 612 may include sub-components for, for example, third-generation (3G) connectivity, fourth-generation (4G) connectivity (e.g., 4G LTE), fifth-generation connectivity (e.g., 5G or NR), Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. The wireless connectivity component 612 is further coupled to one or more antennas 614.
[0115] The processing system 600 may also include one or more sensor processing units 616 associated with sensors in any manner, one or more image signal processors (ISPs) 618 associated with image sensors in any manner, and / or a navigation processor 620, which may include satellite-based positioning system components (e.g., GPS or GLONASS) and inertial positioning system components.
[0116] The processing system 600 may also include one or more input and / or output devices 622, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, speakers, microphones, and so on.
[0117] In some examples, one or more processors of the processing system 600 may be based on the ARM or RISC-V instruction set.
[0118] The processing system 600 also includes a memory 624, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 624 includes computer-executable components that can be executed by one or more of the aforementioned processors of the processing system 600.
[0119] Specifically, in this example, the memory 624 includes a full-precision convolution output generation component 624A, a quantized convolution output generation component 624B, and an image processing component 624C. The depicted components and other non-depicted components may be configured to perform various aspects of the methods described herein.
[0120] Generally, the processing system 600 and / or its components may be configured to perform the methods described herein.
[0121] It is noted that in other aspects, aspects of the processing system 600 may be omitted, such as when the processing system 600 is a server computer, etc. For example, in other aspects, the multimedia processing unit 610, the wireless connectivity component 612, the sensor processing unit 616, the ISP 618, and / or the navigation processor 620 may be omitted. Additionally, aspects of the processing system 600 may be distributed, such as training a model and using the model to generate inferences, such as user authentication predictions.
[0122] Example clauses
[0123] Examples of various aspects of the present disclosure are described in the following numbered clauses.
[0124] Clause 1: A computer-implemented method, comprising: receiving image data including at least a first frame and a second frame; generating a first convolutional output based on the first frame using a machine learning model; generating a second convolutional output based on the difference between the first frame and the second frame using one or more quantizers of the machine learning model; generating a third convolutional output associated with the second frame according to a combination of the first convolutional output and the second convolutional output; and performing image processing based on the first convolutional output associated with the first frame and the third convolutional output associated with the second frame.
[0125] Clause 2: The method according to Clause 1, wherein the first convolutional output is generated using a higher precision level than the precision level used to generate the second convolutional output.
[0126] Clause 3: The method according to Clause 2, wherein: the higher precision level includes a floating-point representation of a given bit size; and the precision level used to generate the second convolutional output includes an integer representation of the given bit size.
[0127] Clause 4: The method according to Clause 2, wherein: the higher precision level includes a floating-point representation of a given bit size; and the precision level used to generate the second convolutional output includes an integer representation of a bit size smaller than the given bit size.
[0128] Clause 5: The method according to any one of Clauses 1-4, wherein: the first frame includes a key frame, and generating the first convolutional output includes: generating an output using the full precision defined for the machine learning model.
[0129] Clause 6: The method according to Clause 5, wherein: the second frame includes a delta frame, and the delta frame includes information about the difference between the first frame and the second frame.
[0130] Clause 7: The method according to any one of Clauses 1-6, wherein generating the second convolutional output includes: selecting a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the first frame and the second frame, wherein each quantizer in the one or more quantizers is associated with a different precision level; and quantizing the difference between the first frame and the second frame using the selected quantizer.
[0131] Clause 8: The method according to Clause 7, wherein selecting the quantizer includes selecting a minimum quantizer, and the difference in quantization error between the minimum quantizer and a larger quantizer is less than a threshold.
[0132] Clause 9: A method as in any of Clauses 1 - 8, wherein generating the second convolutional output comprises performing the following for each respective part of the first frame and each respective corresponding part of the second frame: selecting a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the respective part of the first frame and the respective corresponding part of the second frame, wherein each quantizer is associated with a different level of precision; and quantizing the difference between the respective part of the first frame and the respective corresponding part of the second frame using the selected quantizer.
[0133] Clause 10: A method as in Clause 9, wherein the respective part of the first frame and the respective corresponding part of the second frame correspond to pixels in the same positions in the first frame and the second frame.
[0134] Clause 11: A processing system, comprising: a memory including computer - executable instructions; and one or more processors configured to execute the computer - executable instructions and cause the processing system to perform a method as in any of Clauses 1 to 10.
[0135] Clause 12: A processing system, comprising means for performing a method as in any of Clauses 1 to 10.
[0136] Clause 13: A non - transitory computer - readable medium including computer - executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform a method as in any of Clauses 1 - 10.
[0137] Clause 14: A computer program product embodied on a computer - readable storage medium, comprising code for performing a method as in any of Clauses 1 - 10.
[0138] Additional Considerations
[0139] The foregoing description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not intended to limit the scope, applicability, or aspects set forth in the claims. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made to the functionality and arrangement of the elements discussed without departing from the scope of the disclosure. Various examples may appropriately omit, substitute, or add various procedures or components. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, the features described with reference to some examples may be combined in some other examples. For example, any number of aspects set forth herein may be used to implement a device or practice a method. Additionally, the scope of the disclosure is intended to cover such devices or methods practiced using other structures, functionality, or structures and functionality that supplement or are different from the aspects of the disclosure set forth herein. It should be understood that any aspect of the disclosure disclosed herein may be implemented by one or more elements of the claims.
[0140] As used herein, the term "exemplary" means "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" need not be construed as superior to or better than other aspects.
[0141] As used herein, a phrase that recites "at least one of" a list of items refers to any combination of those items, including a single member. As an example, "at least one of a, b, or c" is intended to cover: a, b, c, a - b, a - c, b - c, and a - b - c, as well as any combination having multiple identical elements (e.g., a - a, a - a - a, a - a - b, a - a - c, a - b - b, a - c - c, b - b, b - b - b, b - b - c, c - c, and c - c - c, or any other ordering of a, b, and c).
[0142] As used herein, the term "determine" encompasses a variety of actions. For example, "determine" may include calculating, computing, processing, deriving, researching, looking up (e.g., looking up in a table, database, or other data structure), ascertaining, and the like. Moreover, "determine" may include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. Also, "determine" may further include parsing, selecting, choosing, establishing, and the like.
[0143] The various methods disclosed herein include one or more steps or acts for implementing the methods. These method steps and / or acts may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of the steps or acts is specified, the order and / or use of the specific steps and / or acts may be altered without departing from the scope of the claims. Additionally, the various operations of the above methods may be performed by any suitable means capable of performing the corresponding functions. These means may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. Generally, where there are operations illustrated in the figures, these operations may have corresponding paired means-plus-function components with similar numbers.
[0144] The following claims are not intended to be limited to the aspects shown and / or described herein, but should be accorded the full scope consistent with the claim language. In the claims, the reference to a singular element is not intended to mean "one and only one" (unless specifically so stated), but rather "one or more." Unless specifically stated otherwise, the term "some / a" means one or more. No element of any claim should be construed under the provisions of 35 U.S.C. § 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, the element is recited using the phrase "step for." Elements of the various aspects described throughout this disclosure that are presently known or later come to be known to those of ordinary skill in the art as all structural and functional equivalents are expressly incorporated herein by reference and are intended to be covered by the claims. Additionally, nothing disclosed herein is intended to be dedicated to the public, whether or not such disclosure is expressly recited in the claims.
Claims
1. A computer-implemented method, comprising: receiving image data including at least a first frame and a second frame; using a machine learning model to generate a first convolutional output based on the first frame; using one or more quantizers of the machine learning model to generate a second convolutional output based on the difference between the first frame and the second frame; generating a third convolutional output associated with the second frame based on a combination of the first convolutional output and the second convolutional output; and performing image processing based on the first convolutional output associated with the first frame and the third convolutional output associated with the second frame.
2. The method according to claim 1, wherein the first convolutional output is generated using a higher precision level than the precision level used to generate the second convolutional output.
3. The method according to claim 2, wherein: the higher precision level includes a floating-point representation of a given bit size; and the precision level used to generate the second convolutional output includes an integer representation of the given bit size.
4. The method according to claim 2, wherein: the higher precision level includes a floating-point representation of a given bit size; and the precision level used to generate the second convolutional output includes an integer representation of a bit size smaller than the given bit size.
5. The method according to claim 1, wherein: the first frame includes a key frame, and generating the first convolutional output includes: generating an output using the full precision defined for the machine learning model.
6. The method according to claim 5, wherein the second frame includes a delta frame, and the delta frame includes information about the difference between the first frame and the second frame.
7. The method according to claim 1, wherein generating the second convolutional output includes: selecting a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the first frame and the second frame, wherein each quantizer in the one or more quantizers is associated with a different precision level; and using the selected quantizer to quantize the difference between the first frame and the second frame.
8. The method according to claim 7, wherein selecting the quantizer includes selecting a minimum quantizer, wherein the difference in quantization error between the minimum quantizer and a larger quantizer is less than a threshold.
9. The method according to claim 1, wherein generating the second convolutional output includes performing the following operations for each corresponding part of the first frame and each corresponding part of the second frame: selecting a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the corresponding part of the first frame and the corresponding part of the second frame, wherein each quantizer is associated with a different precision level; and using the selected quantizer to quantize the difference between the corresponding part of the first frame and the corresponding part of the second frame.
10. The method according to claim 9, wherein The corresponding portions of the first frame and the corresponding corresponding portions of the second frame correspond to pixels at the same positions in the first frame and the second frame.
11. A system, comprising: a memory storing executable instructions; and at least one processor configured to execute the executable instructions to cause the system to: receive image data including at least a first frame and a second frame; generate a first convolutional output based on the first frame using a machine learning model; generate a second convolutional output based on the difference between the first frame and the second frame using one or more quantizers of the machine learning model; generate a third convolutional output associated with the second frame based on a combination of the first convolutional output and the second convolutional output; and perform image processing based on the first convolutional output associated with the first frame and the third convolutional output associated with the second frame.
12. The system of claim 11, wherein the first convolutional output is generated using a higher level of precision than the level of precision used to generate the second convolutional output.
13. The system of claim 12, wherein: the higher level of precision includes a floating-point representation of a given bit size; and the level of precision used to generate the second convolutional output includes an integer representation of the given bit size.
14. The system of claim 12, wherein: the higher level of precision includes a floating-point representation of a given bit size; and the level of precision used to generate the second convolutional output includes an integer representation of a bit size smaller than the given bit size.
15. The system of claim 11, wherein: the first frame includes a key frame, and to generate the first convolutional output, the at least one processor is configured to cause the system to generate an output using the full precision defined for the machine learning model.
16. The system of claim 15, wherein the second frame includes an incremental frame, and the incremental frame includes information about the difference between the first frame and the second frame.
17. The system of claim 11, wherein to generate the second convolutional output, the at least one processor is configured to cause the system to: select a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the first frame and the second frame, wherein each quantizer of the one or more quantizers is associated with a different level of precision; and quantize the difference between the first frame and the second frame using the selected quantizer.
18. The system of claim 17, wherein to select the quantizer, the at least one processor is configured to cause the system to select a minimum quantizer, wherein the difference in quantization error between the minimum quantizer and a larger quantizer is less than a threshold.
19. The system of claim 11, wherein To generate the second convolutional output, the at least one processor is configured to cause the system to perform the following for each respective portion of the first frame and each respective corresponding portion of the second frame: Select a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the respective portion of the first frame and the respective corresponding portion of the second frame, wherein each quantizer is associated with a different precision level; And Quantize the difference between the respective portion of the first frame and the respective corresponding portion of the second frame using the selected quantizer.
20. The system of claim 19, Wherein, The respective portion of the first frame and the respective corresponding portion of the second frame correspond to a set of pixels at the same position in the first frame and the second frame.
21. A system, Comprising: Means for receiving image data comprising at least a first frame and a second frame; Means for generating a first convolutional output based on the first frame using a machine learning model; Means for generating a second convolutional output based on the difference between the first frame and the second frame using one or more quantizers of the machine learning model; Means for generating a third convolutional output associated with the second frame according to a combination of the first convolutional output and the second convolutional output; And Means for performing image processing based on the first convolutional output associated with the first frame and the third convolutional output associated with the second frame.
22. The system of claim 21, Wherein, The means for generating the first convolutional output is configured to: use a higher precision level than the precision level that the means for generating the second convolutional output is configured to use.
23. The system of claim 22, Wherein: The higher precision level includes a floating-point representation of a given bit size; and The precision level that the means for generating the second convolutional output is configured to use includes an integer representation of the given bit size.
24. The system of claim 22, Wherein: The higher precision level includes a floating-point representation of a given bit size; and The precision level for generating the second convolutional output includes an integer representation of a bit size smaller than the given bit size.
25. The system of claim 21, Wherein: The first frame includes a key frame, and The means for generating the first convolutional output includes: means for generating an output using the full precision defined for the machine learning model.
26. The system of claim 25, Wherein, The second frame includes an incremental frame, and the incremental frame includes information about the difference between the first frame and the second frame.
27. The system of claim 21, Wherein, The means for generating the second convolutional output includes: Means for selecting a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the first frame and the second frame, wherein each of the one or more quantizers is associated with a different precision level; and Apparatus for quantizing the difference between the first frame and the second frame using a selected quantizer.
28. The system according to claim 27, wherein, the apparatus for selecting the quantizer includes an apparatus for selecting a minimum quantizer, wherein the difference in quantization error between the minimum quantizer and a larger quantizer is less than a threshold.
29. The system according to claim 21, wherein, the apparatus for generating the second convolutional output includes an apparatus for performing the following operations for each corresponding part of the first frame and each corresponding counterpart part of the second frame: selecting a quantizer from the one or more quantizers of the machine learning model based on the magnitude of the difference between the corresponding part of the first frame and the corresponding counterpart part of the second frame, wherein each quantizer is associated with a different precision level; and quantizing the difference between the corresponding part of the first frame and the corresponding counterpart part of the second frame using the selected quantizer.
30. A non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, perform operations including the following: receiving image data including at least a first frame and a second frame; generating a first convolutional output based on the first frame using a machine learning model; generating a second convolutional output based on the difference between the first frame and the second frame using one or more quantizers of the machine learning model; generating a third convolutional output associated with the second frame based on a combination of the first convolutional output and the second convolutional output; and performing image processing based on the first convolutional output associated with the first frame and the third convolutional output associated with the second frame.