System, method and device for optimizing spatial block-level pixel activity extraction using motion vectors
By using motion estimation data and previously calculated block-level pixel activity estimation to generate pixel activity estimates for new video frame blocks, the problem of time-consuming and high power consumption in block-level pixel activity calculation in videos in the prior art is solved, and more efficient video processing is achieved.
Patent Information
- Application Number
- CN201980059492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-09-28
- Filing Date
- 2019-09-19
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2039-09-19
AI Technical Summary
The prior art is time-consuming and power consumption when calculating block-level pixel activity in video, especially when calculating pixel co-occurrence from scratch, which requires processing of the entire frame, resulting in high processing requirements.
By receiving the new video frames of the video stream, the reference video frame block corresponding to the new video frame block is identified using the motion estimation data, and an estimate of the pixel activity of the new video frame block is generated based on the previously calculated block-level pixel activity.
Reduces unnecessary pixel activity calculations, reduces calculation costs and power consumption, and improves video processing efficiency.
Smart Images

Figure CN112673631B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] DESCRIPTION OF THE RELATED ART
[0002] Various applications perform encoding and decoding of image or video content. For example, video transcoding, desktop sharing, cloud gaming, and game viewing are some applications that include support for encoding and decoding of content. Pixel activity calculations within blocks are typically performed in different video processing and analysis algorithms. For example, for certain applications, block-level pixel activity can be used to determine the type of texture in an image. Examples of blocks include coding tree blocks (CTBs) for the High Efficiency Video Coding (HEVC) standard or macroblocks for the H.264 standard. Other types of blocks that can be used with other types of standards are also possible.
[0003] Different methods can be used to calculate the pixel activity of blocks of an image or video frame. For example, in one embodiment, a gray-level co-occurrence matrix (GLCM) is calculated for each block of a frame. GLCM data shows the frequency with which pixel brightness changes differently within a block. GLCM data can be calculated for pixels at a distance d and an angle θ. In another embodiment, a two-dimensional (2D) spatial average gradient is calculated for a given block. This gradient can capture vertical and horizontal edges. In other embodiments, wavelet transforms or other types of transforms are used to measure activity parameters of a given block. Thus, as used herein, the terms "pixel activity metric," "pixel activity," or "plural pixel activities" are defined as the GLCM, gradient, wavelet transform, or other metric or summary statistic of a block. Note that the terms "pixel activity" and "block-level pixel activity" can be used interchangeably herein. In some cases, pixel activity is represented using a matrix. In other cases, pixel activity is represented using one or more values.
[0004] Calculating block-level pixel activity for each block of each frame of a video can be a time-consuming and / or unnecessary power-consuming operation depending on the type of activity metric to be used. For example, calculating pixel co-occurrence from scratch requires evaluating pixel pairs of all pixels within a block. This type of calculation requires processing the entire frame. Although parallel calculations can be performed in some cases, the processing requirements are still high. SUMMARY OF THE INVENTION
[0005] The present invention provides a system for generating block-level activity in a video, comprising: an interface configured to receive a new video frame of a video stream; control logic coupled to the interface, wherein the control logic is configured to: generate an estimate of the block-level pixel activity of the new video frame based on: motion estimation data of the new video frame, wherein the motion estimation data is used to identify a block of a reference video frame corresponding to a block of the new video frame; and previously computed block-level pixel activity from the reference video frame, wherein the previously computed block-level pixel activity is at a first granularity that is finer than a second granularity for generating the estimate of the block-level pixel activity of the new video frame, wherein the first granularity represents a smaller block size than the second granularity; an encoder configured to generate an encoded video frame based on the estimate, wherein the encoded video frame represents the new video frame.
[0006] The present invention also provides a method for generating block-level activity in a video, comprising: receiving, by a server, a new video frame of a video stream; and generating, by control logic, an estimate of the block-level pixel activity of the new video frame of the video stream based on: motion estimation data of the new video frame, wherein the motion estimation data is used to identify a block of a reference video frame corresponding to a block of the new video frame; and previously computed block-level pixel activity from the reference video frame, wherein the previously computed block-level pixel activity is at a first granularity that is finer than a second granularity for generating the estimate of the block-level pixel activity of the new video frame, wherein the first granularity represents a smaller block size than the second granularity; and generating, by an encoder, an encoded video frame based on the estimate, wherein the encoded video frame represents the new video frame.
[0007] The present invention also provides a device for generating block-level activity in a video, comprising: a memory; an encoder coupled to the memory; and control logic coupled to the encoder, wherein the control logic is configured to generate an estimate of the block-level pixel activity of a new video frame based on: motion estimation data of the new video frame, wherein the motion estimation data is used to identify a block of a reference video frame corresponding to a block of the new video frame; and previously computed block-level pixel activity from the reference video frame stored in the memory, wherein the previously computed block-level pixel activity is at a first granularity that is finer than a second granularity for generating the estimate of the block-level pixel activity of the new video frame, wherein the first granularity represents a smaller block size than the second granularity; wherein the encoder is configured to generate an encoded video frame based on the estimate, wherein the encoded video frame represents the new video frame. Description of the Drawings
[0008] The advantages of the methods and mechanisms described herein can be better understood by reference to the following description in conjunction with the accompanying drawings, in which:
[0009] Figure 1 is a block diagram of an embodiment of a system for encoding and decoding content.
[0010] Figure 2 is a block diagram of an embodiment of a server.
[0011] Figure 3 is a block diagram of an embodiment of a set of motion vectors for a series of video frames.
[0012] Figure 4 is a general flowchart of an embodiment of a method for implementing spatial block-level pixel activity extraction optimization using motion vectors.
[0013] Figure 5 is a general flowchart of another embodiment of a method for generating block-level pixel activity for blocks of a new frame.
[0014] Figure 6 is a general flowchart of an embodiment of a method for determining a pixel activity metric generation scheme.
[0015] Figure 7 is a block diagram of an embodiment of consecutive frames of a video sequence and corresponding motion vectors.
[0016] Figure 8 is a block diagram of a frame with different block metric granularity levels.
[0017] Figure 9 is a general flowchart of an embodiment of a method for calculating pixel activity metrics at different granularities. Detailed Description
[0018] In the following description, numerous specific details are set forth to provide a thorough understanding of the methods and mechanisms presented herein. However, one of ordinary skill in the art will recognize that various embodiments may be practiced without these specific details. In some instances, well-known structures, components, signals, computer program instructions, and techniques have not been shown in detail to avoid obscuring the methods described herein. It should be understood that, for the sake of simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements.
[0019] The present disclosure relates to systems, devices, and methods for implementing spatial block-level pixel activity extraction optimization techniques. In one embodiment, a system includes an encoder and control logic coupled to the encoder. In one embodiment, the control logic utilizes motion estimation data to avoid unnecessary and costly pixel activity calculations. In this embodiment, the control logic calculates an approximation of the pixel activity. For example, in one embodiment, the control logic uses local motion vectors and block-level pixel activity calculated for a reference frame to generate an estimate of the pixel activity for blocks of a new frame. As used herein, the term "reference frame" is defined as a frame of a video stream that is used to define and / or encode one or more future frames of the video stream. In one embodiment, the control logic processes the new frame on a block-by-block basis. In one embodiment, the pixel activity is a gradient calculated for a block of the new frame. In another embodiment, the pixel activity includes texture analysis of a block of the new frame. In other embodiments, the pixel activity includes any one of a variety of summary statistics or other mathematical quantities (e.g., mean, maximum, variance).
[0020] In one embodiment, the control logic calculates the difference (i.e., the sum of the difference costs) between a block of the new frame and the corresponding block of the reference frame, where the corresponding block of the reference frame is identified by a motion vector. If the difference (or "error") is below a first threshold, the control logic copies the pixel activity from the corresponding block of the reference frame instead of recalculating the pixel activity for the block of the new frame. If the error is greater than or equal to the first threshold but less than a second threshold, the control logic uses the motion estimation data to extrapolate from the pixel activity of the corresponding block of the reference frame to generate an estimate of the pixel activity for the block of the new frame. Otherwise, if the error is greater than or equal to the second threshold, the control logic uses a conventional method to calculate the pixel activity for the block of the new frame.
[0021] Now referring to Figure 1 , a block diagram of one embodiment of a system 100 for encoding and decoding content is shown. System 100 includes a server 105, a network 110, a client 115, and a display 120. In other embodiments, system 100 includes a plurality of clients connected to server 105 via network 110, where the plurality of clients receive the same bitstream or different bitstreams generated by server 105. System 100 may also include more than one server 105 for generating a plurality of bitstreams for the plurality of clients.
[0022] In one embodiment, system 100 implements encoding and decoding of video content. In various embodiments, system 100 implements different applications such as video game applications, cloud game applications, virtual desktop infrastructure applications, or screen sharing applications. In other embodiments, system 100 executes other types of applications. In one embodiment, server 105 renders video or image frames, generates a pixel activity metric for a block of the frames, encodes the frames into a bitstream, and then transmits the encoded bitstream via network 110 to client 115. Client 115 decodes the encoded bitstream and generates video or image frames to drive to display 120 or to a display compositor.
[0023] In one embodiment, server 105 generates an estimate of the pixel activity metric for a block of a frame instead of calculating the pixel activity metric from scratch. For example, a block (i,j) in frame f is represented by block (f,i,j). In one embodiment, the pixel activity metric may be used for a block (f1,i,j) in a reference frame (f1). In one embodiment, motion vectors are calculated for all blocks within the current frame f. One method of calculating motion estimation data is to determine the sum of absolute differences (SAD) between pixel samples of a reference block in a reference image and another candidate block in the current image. Candidates for the current block are multiple positions in a "search region" - each position having a horizontal displacement dx and a vertical displacement dy. Motion estimation finds the vector (dx,dy) with the minimum SAD in the search region. In this case, SAD is referred to as the "cost". Other cost functions are possible. The reference image may be different from the current image (temporal), or the reference image may be the same image (spatial). The motion vector MV(dx,dy,i,j,C) represents a motion vector for blocks i and j with displacements dx and dy and a cost C. The pixel activity metric for block (i,j,f) is calculated based on the pixel activity metric of block (i - dx,j - dy,f1), where f1 is the reference frame with the motion vector (dx,dy,i,j,C).
[0024] In one embodiment, if C < Threshold1, the two blocks are similar, and the pixel activity metric of the new block is estimated as the pixel activity metric of the relevant block in the reference frame. If C ≥ Threshold1 and C < Threshold2, the two blocks are similar but different. In this case, according to requirements, the cost C is mapped to a correction factor using a transfer function. The pixel activity metric of the reference block is updated using the correction factor to generate an estimate of the pixel activity metric of the current block. Alternatively, a machine learning solution is used to collect some information, including the motion estimation cost and the reference block pixel activity metric, to calculate the pixel activity metric of the new block. If C > Threshold2, the two blocks are different, and the pixel activity metric of the block is calculated from scratch.
[0025] Network 110 represents any type of network or combination of networks, including wireless connections, direct local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), intranets, the Internet, wired networks, packet-switched networks, fiber optic networks, routers, storage area networks, or other types of networks. Examples of LANs include Ethernet, Fiber Distributed Data Interface (FDDI) networks, and Token Ring networks. In various embodiments, network 110 includes Remote Direct Memory Access (RDMA) hardware and / or software, Transmission Control Protocol / Internet Protocol (TCP / IP) hardware and / or software, routers, forwarders, switches, grids, and / or other components.
[0026] Server 105 includes any combination of software and / or hardware for rendering video / image frames and encoding the frames into a bitstream. In one embodiment, server 105 includes one or more software applications executed on one or more processors of one or more servers. Server 105 also includes network communication capabilities, one or more input / output devices, and / or other components. The processors of server 105 include any number and type of processors (e.g., Graphics Processing Unit (GPU), Central Processing Unit (CPU), Digital Signal Processor (DSP), Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC)). The processors are coupled to one or more memory devices storing program instructions executable by the processors. Similarly, client 115 includes any combination of software and / or hardware for decoding the bitstream and driving the frames to display 120. In one embodiment, client 115 includes one or more software applications executed on one or more processors of one or more computing devices. In various embodiments, client 115 is a computing device, gaming console, mobile device, streaming media player, or other type of device.
[0027] Turning now to Figure 2 , a block diagram of one embodiment of the software components of server 200 for encoding video frames is shown. Note that in other embodiments, server 200 includes other components and / or is different from Figure 2Arranged in other suitable ways as shown. The server 200 receives a new frame 205 of video on the interface 208, and the frame is coupled to the motion vector unit 210, the control logic 220, and the encoder 230. Depending on the implementation, the interface 208 is a bus interface, a memory interface, or an interconnection to a communication fabric and / or other types of devices. Each of the motion vector unit 210, the control logic 220, and the encoder 230 is implemented using any suitable combination of hardware and / or software. The motion vector unit 210 generates motion vectors 215 for blocks of the new frame 205 based on a comparison of the new frame 205 with a reference frame 207. In one implementation, the reference frame 207 is stored in the memory 240. The memory 240 represents any number and type of memory or cache devices for storing data and / or instructions associated with the encoding process.
[0028] The motion vectors 215 are provided to the control logic 220 and the encoder 230. The control logic 220 generates an estimated pixel activity 225 based on the pixel activity 222 calculated from the reference frame according to the motion vectors 215. For example, in one implementation, the control logic 220 processes the new frame 205 on a block-by-block basis. For each block, the control logic 220 retrieves the sum of the calculated difference costs between the block and the corresponding block in the reference frame 207 identified by the corresponding motion vector 215. If the sum of the difference costs of the block is less than a first threshold, the control logic 220 generates the estimated pixel activity 225 of the block as the previously calculated pixel activity 222 of the corresponding block in the reference frame 207.
[0029] If the sum of the difference costs of the block is greater than or equal to the first threshold but less than a second threshold, the control logic 220 generates the estimated pixel activity 225 of the block by extrapolating from the pixel activity 222 of the corresponding block in the reference frame 207 based on the block's motion vector 215. For example, in one implementation, the control logic 220 uses a transfer function to map the sum of the difference costs to a correction factor that is used to update the pixel activity 222 of the corresponding block according to the reference frame 207. Then, this updated pixel activity from the corresponding block in the reference frame 207 is used as the estimated pixel activity 225 of the block in the new frame 205. In one implementation, a look-up table is used to apply the transfer function. If the sum of the difference costs of the block is greater than or equal to the second threshold, the control logic 220 generates the pixel activity of the block from scratch using conventional techniques.
[0030] Now refer to Figure 3, a block diagram showing an embodiment of a set of motion vectors 315A-C for a series of video frames 305A-D is presented. Frames 305A-D represent consecutive frames of a video sequence. Block 310A represents individual pixel blocks within frame 305A. Block 310A may also be referred to as a macroblock. Arrow 315A represents the known motion of the image within block 310A as the video sequence moves from frame 305A to 305B. The known motion shown by arrow 315A can be defined by a motion vector. It should be noted that although motion vectors 315A-C point in the direction of motion of block 310 in subsequent frames, in another embodiment, the motion vector can be defined to point in the direction opposite to the motion of the image. For example, in some compression standards, the motion vector associated with a macroblock points to the source of the block in a reference frame. The reference frame can be forward or backward in time. It should also be noted that in certain embodiments, the motion vector can represent entropy.
[0031] In one embodiment, motion vectors 315A-C can be used to track blocks 310B-D in subsequent frames. For example, motion vector 315A indicates the change in position of block 310B in frame 305B compared to block 310A in frame 305A. Similarly, motion vector 315B indicates the change in position of block 310C in frame 310C compared to block 310B in frame 305B. Additionally, motion vector 315C indicates the change in position of block 310D in frame 310D compared to block 310C in frame 305C. In another embodiment, the motion vector is defined to track the reverse motion of a block from a given frame back to a previous frame.
[0032] In one embodiment, when an encoder needs to generate various pixel activity metrics for blocks in a new frame, the encoder generates an estimate of the pixel activity metric based on the previously calculated pixel activity metrics of the corresponding blocks in the previous frame. In one embodiment, the encoder uses motion vectors 315A-C to identify which corresponding block in the previous frame matches a given block in the new frame. Then, the encoder uses the previously calculated pixel activity metric for the identified block in the previous frame to help generate an estimate of the pixel activity metric for the block in the new frame. In one embodiment, the encoder uses the previously calculated pixel activity metric without modification as an estimate of the pixel activity metric for the block in the new frame. In another embodiment, the encoder extrapolates from the previously calculated pixel activity metric by mapping the sum of the differential costs or the sum of any other costs between blocks to a correction factor using a transfer function. Then the correction factor is applied to the previously calculated pixel activity to generate an estimate of the pixel activity metric for the block in the new frame.
[0033] Now turning to Figure 4 , an embodiment of a method 400 for implementing spatial block-level pixel activity extraction optimization using motion vectors is presented. For purposes of discussion, the steps andFigures 5 to 6 Those steps. However, it should be noted that in various embodiments of the described method, one or more of the described elements are performed simultaneously, performed in a different order than shown, or omitted entirely. Other additional elements may also be performed as needed. Any of the various systems or devices described herein are configured to implement method 400.
[0034] The motion vector unit calculates the sum of the absolute difference costs of the blocks of the new frame in the video stream compared to the corresponding blocks of one or more reference frames (block 405). In one embodiment, the cost is the sum of the absolute difference costs. In other embodiments, the cost is other types of costs or errors calculated based on the blocks of the new frame and the corresponding blocks of one or more reference frames. In one embodiment, the motion vector unit is part of an encoder, where the encoder also includes control logic. In one embodiment, the encoder is implemented on a system having at least one processor coupled to at least one memory device. In one embodiment, the encoder is implemented on a server as part of a cloud computing environment. The motion vector unit generates a motion vector for the block of the new frame based on a comparison with the corresponding block of the reference frame (block 410). The encoder control logic receives the cost and motion vector of the block of the new frame (block 415).
[0035] The control logic also receives the previously calculated block-level pixel activity of the blocks of the reference frame (block 420). Then, the control logic generates an estimate of the block-level pixel activity of the block of the new frame based on the cost and motion vector of the block of the new frame and based on the previously calculated block-level pixel activity of the blocks of the reference frame (block 425). After block 425, method 400 ends. An example of an embodiment for performing block 425 is described in further detail below in the Figure 5 discussion. In various embodiments, the block-level pixel activity for the blocks of the new frame aids in the encoding of the new frame, for classifying the new frame and / or for performing further analysis of the new frame. In one embodiment, the block-level pixel activity for the blocks of the new frame is used to classify the new frame into one of a plurality of categories to identify what type of objects are present in the new frame, and / or to perform other actions.
[0036] Now turning to Figure 5, which shows an embodiment of method 500 for generating block-level pixel activity for blocks of a new frame. Encoder control logic receives a new frame to be encoded (block 505). The encoder control logic is implemented using any suitable combination of hardware and / or software. The control logic analyzes the new frame on a block-by-block basis (block 510). For the purposes of discussion, it is assumed that when the control logic receives a new frame to be encoded, motion estimation data (e.g., motion vectors) and costs have already been calculated for the blocks of the new frame. In embodiments where the motion estimation data or costs have not been calculated, the control logic calculates these values in response to receiving the new frame to be encoded. Any type of cost can be calculated according to the embodiment.
[0037] For each block of the new frame, if the cost of the block is less than a first threshold (conditional block 515, "yes" branch), the control logic generates an estimate of the pixel activity of the block to be equal to the previously calculated pixel activity of the corresponding block in one or more reference frames (block 520). If the cost of the block is greater than or equal to the first threshold (conditional block 515, "no" branch), the control logic determines whether the cost of the block is less than a second threshold (conditional block 525). If the cost of the block is less than the second threshold (conditional block 525, "yes" branch), the control logic generates an estimate of the pixel activity of the block by extrapolating from the previously calculated pixel activity of the corresponding block in the reference frame based on the block's motion estimation data (block 530). If the cost of the block is greater than or equal to the second threshold (conditional block 525, "no" branch), the control logic calculates the pixel activity of the block independently of the previously calculated pixel activity of the corresponding block in the reference frame and the block's motion estimation data (block 535). In other words, in block 535, the control logic calculates the pixel activity of the block from scratch in a conventional manner. After blocks 520, 530, and 535, method 500 ends.
[0038] Now turning to Figure 6 , which shows an embodiment of method 600 for determining a pixel activity metric generation scheme. The encoder generates a cost histogram of the blocks of the new frame compared to the corresponding blocks of one or more reference frames (block 605). If at least a given percentage of the blocks have a cost greater than a threshold (conditional block 610, "yes" branch), the encoder calculates the pixel activity metric of the blocks of the new frame in a conventional manner (block 615). The values of the given percentage and the threshold can vary according to the embodiment. If less than the given percentage of the blocks have a cost greater than the threshold (conditional block 610, "no" branch), the encoder generates an estimate of the pixel activity metric of the blocks of the new frame based on the previously calculated pixel activity metrics, costs, and motion vectors of the blocks of the reference frame (block 620). An example of performing block 620 is described in the discussion related to Figure 5 . After blocks 615 and 620, method 600 ends.
[0039] Now refer to Figure 7 , a block diagram of an embodiment of successive frames of a video stream and corresponding motion vectors is shown. In some embodiments, a pixel activity metric may not be available for any reference block. For example, activity may only be available at the macroblock granularity, which may mean that only reference blocks at positions (16*i, 16*j) are available, where i and j are integers. Note that the example of a 16 pixel by 16 pixel block size only indicates one embodiment. In other embodiments, the macroblock or coding unit may have other sizes. For other motion vectors, if the error is acceptable, an estimate of the pixel activity metric value may be estimated based on the block from which the pixel is sourced. In a particular case, a pixel may be from one block, while in other cases, a pixel may be from up to four blocks. In Figure 7 the example shown, the pixel is from two blocks.
[0040] Figure 7 Two successive video frames 705 and 710 at times 't - 2' and 't - 1' are shown respectively. Within the video frame 705 at time 't - 2', block (0, 2, t - 2) has a pixel activity metric A 0,2, while block (1, 2, t - 2) has a pixel activity metric A 1,2 . Within the video frame 710 at time 't - 1', block (1, 2, t - 1) has a pixel activity metric N 1,2 . In one embodiment, if the cost of block (1, 2, t - 1) compared to the matching pixels of frame 705 is less than a first threshold, the motion vector 715 match is considered a close match and thus the following estimation method can be used. For discussion purposes, assume that block (0, 2, t - 2) contributes w 0 percent of its pixels to block (1, 2, t - 1) and block (1, 2, t - 2) contributes w 1 percent of its pixels to block (1, 2, t - 1). The possible error is inversely proportional to the maximum of w 0 and w 1 . Based on this possible error, the blocks for which activity needs to be calculated rather than just generating an estimate can be determined. In one embodiment, the estimate can be the higher of A 0,2 or A 1,2 or an interpolated value, for example N 1,2 = w 0 * A 0,2 + w 1 * A 1,2 , where w 0 + w 1 = 1.
[0041] In some cases, candidate estimation regions can be calculated occasionally without considering possible errors. For example, when a randomly generated value x is greater than a random threshold, a candidate estimation region can be calculated. This is to ensure that the estimated pixel activity metric does not deviate too far from the actual signal. In one scenario, if the random threshold is set to.9, this means that the actual calculation of the pixel activity metric will be skipped 90% of the time. If higher precision is required, the value of the random threshold can be decreased.
[0042] Now turning to Figure 8 , a block diagram of an embodiment of a frame with different block metric granularity levels is shown. Many pixel activity metrics (such as discrete gradients) are accumulated within a region. In one embodiment, by storing the metrics at a finer granularity, by storing a summed area table or using other acceleration data structures that allow for fast calculation of metrics with specific mathematical properties (e.g., homogeneous, additive), these pixel activity metrics are constructed into a fast estimate. By summing appropriate values in these data structures, an estimated value can be calculated quickly. For example, by adding together the pixel activity metrics of regions 815A - D of a high metric granularity frame 805, the pixel activity metric of region 830 of a low metric granularity frame 810 is estimated. This estimated value is considered acceptable because the additional regions included in regions 815B and 815D are considered insignificant (i.e., the error calculation is acceptable). Similarly, the regions missed adjacent to regions 815A and 815C are also considered trivial. If higher precision is required, correction can be made by calculating and adjusting the contributions of the metrics in the additional required and unwanted regions.
[0043] If the estimate of the high metric granularity frame 805 is considered not precise enough, the estimate can be corrected. For example, in one embodiment, a correction 825A shown in the extended view 820 is calculated for the contribution of the metric in the corresponding marked region. Another correction 825B is calculated for the contribution of the metric in the corresponding marked region. Then, the pixel activity measure of the reference block is estimated by adding together the metrics of regions 815A - D plus correction 825A and subtracting correction 825B. In this example, the required correction is attributed to the error caused by a horizontal shift. Under different alignments, corrections may be required to handle vertical shift errors; or for other alignments, corrections may be required to adjust for both vertical and horizontal shifts.
[0044] Now referring to Figure 9, which shows an embodiment of a method 900 for calculating pixel activity metrics at different granularities. The encoder calculates pixel activity metrics for blocks of a reference frame at a first granularity (block 905). The encoder generates an estimate of the pixel activity metrics for blocks of a new frame at a second granularity, where the first granularity is a finer granularity than the second granularity, and where the estimate is generated based on the pixel activity metrics of blocks of the reference frame (block 910). As part of generating an estimate of the pixel activity metrics for blocks of the new frame at the second granularity, the encoder identifies a plurality of blocks in the reference frame corresponding to each block of the new frame based on motion vectors (block 915). The encoder then generates a cumulative estimate of the pixel activity metrics for each block of the new frame by summing the pixel activity metrics of the plurality of blocks from the reference frame (block 920). After block 920, method 900 ends.
[0045] In various embodiments, the methods and / or mechanisms described herein are implemented using program instructions of a software application. For example, program instructions executable by a general or special purpose processor are contemplated. In various embodiments, such program instructions may be represented by a high-level programming language. In other embodiments, the program instructions may be compiled from a high-level programming language into binary form, intermediate form, or other form. Alternatively, program instructions may be written to describe the behavior or design of hardware. Such program instructions may be represented by a high-level programming language such as C. Alternatively, a hardware design language (HDL) such as Verilog may be used. In various embodiments, the program instructions are stored on any of a variety of non-transitory computer-readable storage media. During use, a computing system may access the storage media to provide the program instructions to the computing system for program execution. Generally speaking, such a computing system includes at least one or more memories and one or more processors configured to execute the program instructions.
[0046] It should be emphasized that the above embodiments are merely non-limiting examples of embodiments. Once the above disclosure is fully understood, numerous variations and modifications will become apparent to those skilled in the art. The following claims are intended to be construed to cover all such variations and modifications.
Claims
1. A system for generating block-level activity in a video, which comprises: An interface configured to receive a new video frame of a video stream; Control logic coupled to the interface, wherein the control logic is configured to: Generate an estimate of the block-level pixel activity of the new video frame based on: Motion estimation data of the new video frame, wherein the motion estimation data is used to identify a block of a reference video frame corresponding to a block of the new video frame; And Previously calculated block-level pixel activity from the reference video frame, wherein the previously calculated block-level pixel activity is at a first granularity that is finer than a second granularity for generating the estimate of the block-level pixel activity of the new video frame, and wherein the first granularity represents a block size smaller than the second granularity; An encoder configured to generate an encoded video frame based on the estimate, wherein the encoded video frame represents the new video frame.
2. The system according to claim 1, wherein for each block of the new video frame, the control logic is configured to: Compare the cost of the block of the new video frame relative to the corresponding block of the reference video frame with a first threshold; and Generate an estimate of the pixel activity of the block, wherein if the cost is less than the first threshold, the estimate is equal to the previously calculated pixel activity of the corresponding block in the reference video frame.
3. The system according to claim 2, wherein for each block of the new video frame, the control logic is further configured to: Compare the cost of the block with a second threshold; In response to determining that the cost is greater than or equal to the first threshold and less than the second threshold: Map the cost to a correction factor using a transfer function; and Generate an estimate of the pixel activity of the block by applying the correction factor to the previously calculated pixel activity of the corresponding block in the reference video frame.
4. The system according to claim 3, wherein the control logic is further configured to: if the cost is greater than or equal to the second threshold, calculate the pixel activity of the block independently of the previously calculated pixel activity.
5. The system according to claim 1, wherein the control logic is configured to classify the new video frame based on the estimate, and wherein the block-level pixel activity represents a gradient, a co-occurrence matrix, or any other metric, wherein the result of a mathematical operation performed on a first pixel value and a second pixel value at a relative displacement defined relative to the first pixel value is aggregated over a block of the new video frame.
6. The system according to claim 1, wherein the previously calculated block-level pixel activity is at a first granularity, and wherein the encoder is configured to generate an estimate of the block-level pixel activity of the new video frame at a second granularity, and wherein the first granularity is a finer granularity than the second granularity.
7. The system according to claim 6, wherein the control logic is further configured to apply a correction to one or more of the estimates based on an alignment error.
8. A method for generating block-level activity in a video, which comprises: Receiving a new video frame of a video stream by a server; And The control logic generates an estimate of the block - level pixel activity of the new video frame for the video stream based on the following: Motion - estimation data of the new video frame, where the motion - estimation data is used to identify a block of a reference video frame corresponding to a block of the new video frame; And Previously - calculated block - level pixel activity from the reference video frame, where the previously - calculated block - level pixel activity is at a first granularity that is finer than a second granularity used to generate the estimate of the block - level pixel activity of the new video frame, and where the first granularity represents a smaller block size than the second granularity; And The encoder generates an encoded video frame based on the estimate, where the encoded video frame represents the new video frame.
9. The method according to claim 8, further comprising: Comparing the cost of the block of the new video frame relative to the corresponding block of the reference video frame with a first threshold; And Generating an estimate of the pixel activity of the block, where if the cost is less than the first threshold, the estimate is equal to the previously - calculated pixel activity of the corresponding block in the reference video frame.
10. The method according to claim 9, further comprising: Comparing the cost of the block with a second threshold; In response to determining that the cost is greater than or equal to the first threshold and less than the second threshold: Mapping the cost to a correction factor using a transfer function; And Generating an estimate of the pixel activity of the block by applying the correction factor to the previously - calculated pixel activity of the corresponding block in the reference video frame.
11. The method according to claim 10, further comprising: If the cost is greater than or equal to the second threshold, calculating the pixel activity of the block independently of the previously - calculated pixel activity.
12. The method according to claim 8, further comprising classifying the new video frame based on the estimate, and where the block - level pixel activity represents a gradient, a co - occurrence matrix, or any other metric, where the result of a mathematical operation on a first pixel value and a second pixel value at a relative displacement defined relative to the first pixel value is aggregated over a block of the new video frame.
13. The method according to claim 8, where the control logic is further configured to apply a correction to one or more of the estimates based on an alignment error.
14. The method according to claim 13, further comprising generating a cumulative estimate of the block - level pixel activity of each block of the new video frame by summing the pixel activities of multiple blocks from the reference video frame.
15. A device for generating block - level activity in a video, which comprises: A memory; An encoder coupled to the memory; And A control logic coupled to the encoder, where the control logic is configured to generate an estimate of the block - level pixel activity of a new video frame based on: Motion - estimation data of the new video frame, where the motion - estimation data is used to identify a block of a reference video frame corresponding to a block of the new video frame; And Previous calculated block-level pixel activities from the reference video frame stored in the memory, where the previous calculated block-level pixel activities are at a first granularity that is finer than the second granularity of the estimated block-level pixel activities for generating the new video frame, and where the first granularity represents a block size smaller than the second granularity; where the encoder is configured to generate an encoded video frame based on the estimate, and where the encoded video frame represents the new video frame.
16. The apparatus according to claim 15, wherein for each block of the new video frame, the control logic is configured to: compare the cost of the block of the new video frame relative to the corresponding block of the reference video frame with a first threshold; and generate an estimate of the pixel activity of the block, wherein if the cost is less than the first threshold, the estimate is equal to the previously calculated pixel activity of the corresponding block in the reference video frame.
17. The apparatus according to claim 16, wherein for each block of the new video frame, the control logic is further configured to: compare the cost of the block with a second threshold; in response to determining that the cost is greater than or equal to the first threshold and less than the second threshold: map the cost to a correction factor using a transfer function; and generate an estimate of the pixel activity of the block by applying the correction factor to the previously calculated pixel activity of the corresponding block in the reference video frame.
18. The apparatus according to claim 17, wherein the control logic is further configured to: if the cost is greater than or equal to the second threshold, calculate the pixel activity of the block independently of the previously calculated pixel activity.
19. The apparatus according to claim 15, wherein the control logic is configured to classify the new video frame based on the estimate, and wherein the block-level pixel activity represents a gradient, a co-occurrence matrix, or any other metric, where the results of mathematical operations performed on a first pixel value and a second pixel value at a relative displacement defined relative to the first pixel value are aggregated over the blocks of the new video frame.
20. The apparatus according to claim 15, wherein the previously calculated block-level pixel activities are at a first granularity, and wherein the encoder is configured to generate an estimate of the block-level pixel activities of the new video frame at a second granularity, and wherein the first granularity is a finer granularity than the second granularity.
Citation Information
Patent Citations
On-camera image processing based on image activity data
US20170339390A1
Determining variance of a block of an image based on a motion vector for the block
US20180109804A1