Video processing method and video processing apparatus

CN116437083BActive Publication Date: 2026-09-15MEDIATEK INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211457059.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-08-22
Filing Date
2022-11-16
Publication Date
2026-09-15
Estimated Expiration
2042-11-16

AI Technical Summary

Benefits of technology

[0008] One object of this disclosure is to provide schemes, concepts, designs, techniques, methods, and apparatus related to precoding video picture frames in a video stream using pixel-based filtering (e.g., motion-compensated temporal filtering). It is believed that various embodiments of this disclosure achieve benefits including improved precoding latency, higher encoding/decoding gain, and/or reduced hardware overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116437083B_ABST
    Figure CN116437083B_ABST
Patent Text Reader

Abstract

A video processing method and related devices are provided. The video processing method includes determining a plurality of target pictures based on a filter interval, including a first subset of a plurality of source pictures of a video, each of the plurality of source pictures having a temporal identifier identifying a temporal position of the respective source picture in a temporal sequence; determining, for each of the plurality of target pictures, a reference picture, each of the reference pictures being a source picture of a second subset of the plurality of source pictures; generating a plurality of filtered pictures, each filtered picture corresponding to a respective one of the plurality of target pictures and generated by performing a pixel-based filtering based on the reference picture corresponding to the respective target picture; and encoding the video into a bitstream using the filtered pictures and the second subset of the plurality of source pictures. The video processing method and related devices improve pre-encoding delay.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to video encoding and decoding, and more specifically, to methods and apparatus for encoding video containing frames employing motion-compensated temporal filtering. [Background Technology]

[0002] Unless otherwise stated herein, the methods described in this section are not prior art to the claims and are not included in this section as prior art.

[0003] Video encoding and decoding typically involve encoding video (i.e., the original video) into a bitstream via an encoder, transmitting the bitstream to a decoder, and then parsing and processing the bitstream to decode the video from it, producing a reconstructed video. Encoders can use various encoding / decoding modes or tools to encode video with the aim of reducing the total size of the bitstream transmitted to the decoder while still providing the decoder with enough information about the original video so that the decoder can generate a reconstructed video that is satisfactorily faithful to the original. Therefore, in addition to the video data, the bitstream may also include some information about the encoding / decoding tools used, which the decoder needs to successfully reconstruct the video from the bitstream.

[0004] In addition to encoding the raw video to reduce the bitstream size, the encoder can also preprocess the video before the actual encoding operation occurs. That is, the encoder can examine the image frames of the raw video to understand certain features of the video, and then manipulate or otherwise adjust certain aspects of the image frames based on the examination results before performing the encoding operation on the raw video. Preprocessing can provide benefits such as further reducing the bitstream size achieved at the encoder's output and / or enhancing certain features of the reconstructed video at the decoder. [Summary of the Invention]

[0005] The following overview is illustrative only and is not intended to be limiting in any way. That is, it is provided to introduce the concepts, key points, benefits, and advantages of the novel and progressive techniques described herein. The alternative implementations are further described in the detailed description below. Therefore, the following summary is not intended to identify essential features of the claimed subject matter, nor is it intended to define the scope of the claimed subject matter.

[0006] This invention provides a video processing method, comprising: determining multiple target images based on a filtering interval, including a first subset of multiple source images of a video, each of the multiple source images having a time identifier to identify the time position of the corresponding source image in a time series; determining a reference image for each of the multiple target images, each of the reference images being a source image of a second subset of the multiple source images; generating multiple filtered images, each filtered image corresponding to a corresponding one of the multiple target images, and generating the filtered images by performing pixel-based filtering based on the reference images corresponding to each target image; and encoding the video into a bitstream using the filtered images and the second subset of the multiple source images.

[0007] The present invention also provides a video processing apparatus, comprising: a processor configured to receive a video comprising a plurality of source images in a time series, each of the plurality of source images having a time identifier identifying a time position of a corresponding source image in the time series; a target image buffer configured to store a plurality of target images determined by the processor based on a filtering interval, the plurality of target images comprising a first subset of the plurality of source images; a reference image buffer configured to store one or more reference images determined by the processor for each of the plurality of target images, each of the one or more reference images being a source image of a second subset of the plurality of source images, wherein the second subset of the plurality of source images comprises a plurality of source images not in the first subset; a motion compensation (MC) module configured to generate a plurality of filtered images by performing pixel-based filtering on a corresponding target image of the plurality of target images based on one or more reference images corresponding to a corresponding target image, each of the plurality of filtered images being generated by the MC module; and a video encoder configured to encode the plurality of filtered images and the second subset of the plurality of source images into a bitstream representing a video.

[0008] One object of this disclosure is to provide schemes, concepts, designs, techniques, methods, and apparatus related to precoding video picture frames in a video stream using pixel-based filtering (e.g., motion-compensated temporal filtering). It is believed that various embodiments of this disclosure achieve benefits including improved precoding latency, higher encoding / decoding gain, and / or reduced hardware overhead. [Attached Image Description]

[0009] The accompanying drawings are included to provide a further understanding of this disclosure, and are incorporated in and constitute a part of this disclosure. The drawings illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the disclosure. It is understood that the drawings are not necessarily drawn to scale, as some components may be shown out of proportion to actual dimensions in order to clearly illustrate the concepts of the disclosure.

[0010] Figure 1This is a diagram designed according to an example of an embodiment of this disclosure.

[0011] Figure 2 Explain how to do it Figure 1 Perform an integer pixel search on the target image.

[0012] Figure 3 This describes a video with a time sequence, which includes multiple mixed images.

[0013] Figure 4 This is a diagram designed according to an example of an embodiment of this disclosure.

[0014] Figure 5 This includes diagrams illustrating example designs based on embodiments of this disclosure.

[0015] Figure 6 This includes a diagram of another example design based on an embodiment of this disclosure.

[0016] Figure 7 An example utilizing hardware parallelism is shown.

[0017] Figure 8 An example design according to an embodiment of this disclosure is illustrated.

[0018] Figure 9 An example design according to an embodiment of this disclosure is illustrated.

[0019] Figure 10 An example video encoder is shown.

[0020] Figure 11 This describes the part of the video encoder that implements the pre-encoding processing module.

[0021] Figure 12 An example video decoder is shown.

[0022] Figure 13 An example process according to an embodiment of the present disclosure is illustrated.

[0023] Figure 14 An example process according to an embodiment of the present disclosure is illustrated.

[0024] Figure 15 An example process according to an embodiment of the present disclosure is illustrated.

[0025] Figure 16 An electronic system that implements some embodiments of the present disclosure is conceptually illustrated.

Detailed Implementation Methods

[0026] The following description represents the best mode for carrying out the invention. This description is intended to illustrate the general principles of the invention and should not be construed as limiting. The scope of the invention is determined by reference to the appended claims.

[0027] This document discloses detailed embodiments and implementations of the claimed subject matter. However, it should be understood that the disclosed embodiments and implementations are merely illustrative of the claimed subject matter, which can be embodied in various forms. This disclosure may be embodied in many different forms and should not be construed as limited to the exemplary embodiments and implementations set forth herein. Rather, these exemplary embodiments and implementations are provided so that the description of this disclosure will be thorough and complete, and will fully convey the scope of this disclosure to those skilled in the art. In the following description, details of well-known features and techniques may be omitted to avoid unnecessarily obscuring the presented embodiments and implementations.

[0028] Embodiments of this disclosure relate to various techniques, methods, schemes, and / or solutions related to encoding video using motion-compensated temporal filtering (MCTF) precoding. According to this disclosure, multiple possible solutions can be implemented individually or in combination. That is, although these possible solutions may be described individually below, two or more of these possible solutions may be implemented in one combination or another.

[0029] I. Precoding using MCTF

[0030] As mentioned above, instead of directly encoding image frames of the source video, the encoder can process source video frames before actually encoding the video. Figure 1 This is an example design based on an embodiment of the present disclosure, wherein a video encoder 130 is shown having a pre-encoding processing module 132 and an encoding module 134 (a video encoder in the conventional sense). The video encoder 130 is configured to receive a video stream comprising a time series 190 of multiple source video frames from a video source 105. The pre-encoding processing module 132 is configured to alter or adjust certain characteristics of the source video frames of the time series 190 by performing a pre-encoding process, enabling the encoding module 134 to perform a more efficient encoding process. This may result in superior encoding outcomes, such as a smaller encoded video size of the bitstream 195 and / or higher subjective / objective video quality of the video decoded by the video decoder accessing the bitstream 195.

[0031] Typically, video consists of multiple pictures or "frames" presented in a time sequence. That is, a series of pictures, when captured or displayed in a specific temporal order, is called video. For example, a camera or camcorder can use a series of picture frames to capture video of a moving object over a period of time. Each picture frame contains a "snapshot" of the object at a different moment—a record of the object at a specific instant within a given time period. When displayed in the same chronological order as the camera recorded the object, video is a faithful reproduction of the moving object within that time period.

[0032] Video can be generated using methods other than cameras. For example, video games or cartoon animations may consist of a series of images generated by computer graphics algorithms or human drawing. In some embodiments, multiple sources can be combined to generate a video. Regardless of its creator or generation method, a video consists of multiple images that have a temporal relationship. That is, a video has multiple images presented or recorded in chronological order. In order to render the images as video (e.g., on a display), the temporal relationship between the images must be maintained. That is, the order of the images, i.e., the temporal position of each image in a time series, will be recorded or otherwise specified.

[0033] like Figure 1 As shown, the precoding module 132 performs certain operations on the time series 190, thereby modifying the time series 190 into another time series 193, which is then passed to the encoding module 134. A portion of the time series 190 remains unchanged in the time series 193, while other portions are modified by the precoding module 132, as described in detail below.

[0034] Figure 1 The time-series recording schemes used by time series 190 and 193 are also shown, where the temporal relationships between images can be recorded. For example... Figure 1 As shown, time series 190 includes a series of images, such as images 100, 101, 102, 103, ..., 106, 107, 108, 109, and 110, which are presented in time series 190, where there are temporal relationships between the images. The temporal relationship is expressed as the order of the images in time series 190. For example, image 100 is the first image in time series 190. That is, image 100 represents the first frame of time series 190 as a video presentation (e.g., recording or display). In time series 190, image 101 is followed by image 102, then image 103, and so on. Similarly, image 106 is followed by image 107, image 107 by image 108, then image 109, then image 110, and so on. Furthermore, Figure 1Each image in the video has a time identifier called a "picture order count (POC)," which is an integer used to record or otherwise identify the time position of the corresponding image in the time series 190. For example... Figure 1 As shown, image 100 has a corresponding time identifier that is specified or otherwise recorded as POC=0, while the POC of image 101 is specified as POC=1. Similarly, the POC values ​​of images 102, 103…106, 107, 108, 109, and 110 are specified as POC=2, 3…6, 7, 8, 9, and 10, respectively. Figure 1 As shown. This scheme records the temporal relationships between images. The POC value of a specific image identifies its temporal position in the video time series. Each image in the time series has a unique POC value, and the first image with a POC value less than the second image must precede the second image in the video time series. POC information is important for the encoder to perform the pre-encoding process, as will be disclosed in detail elsewhere below.

[0035] The general idea of ​​the precoding process according to this disclosure is as follows. In this disclosure, the terms "frame," "picture frame," and "source picture" are used interchangeably, referring to a picture in the video before the precoding process is performed, but not limited to recorded pictures or pictures otherwise generated by the camera. Figure 1As shown, the video encoder 130 receives video from video source 105, including source images 100, 101, ..., 109, 110, etc., in time series 190. Some of these source images (e.g., source images 100 and 108) are selected as "target images," and the precoding processing module 132 will perform certain precoding processing operations on them. The remaining source images in time series 190, i.e., those not selected as target images, are called "reference images" (e.g., source images 101, 102, ..., 107 and 109, 110, etc.). Specifically, the precoding processing unit 132 generates a "processed image," or sometimes called a "filtered image," for each target image, such as images 180 and 188. One or more of the reference images may be used in the precoding process or otherwise referenced for generating the processed image. The precoding process can use different subsets of the reference images to generate processed images for different target images. The precoding process performed by the precoding processing unit 132 alters or otherwise modifies time series 190 into time series 193, which includes filtered images 180 and 188, and reference images 101, 102, ..., 107 and 109, 110. Notably, the temporal relationships of the images in time series 190 are inherited or otherwise maintained in time series 193. That is, the processed images 180 and 188 retain the same POC values ​​as the target images 100 and 108, respectively.

[0036] The precoding process according to this disclosure involves applying motion-compensated temporal filtering (MCTF) to certain reference images to generate a processed image for a target image. Therefore, in this disclosure, the terms "processed image (processed image)" and "filtered image (also referred to as filtered image, filtered image, filtered picture, etc.)" are used interchangeably. The main concept of MCTF is to generate a filtered image for a target image by performing a discrete wavelet transform with motion compensation (MC) on a set of reference images associated with or otherwise related to the target image. Typically, the reference images used or otherwise referenced by MCTF are adjacent frames of the target image. For example, if image 108 is the target image, the corresponding set of reference images to be used by MCTF may include images 106, 107, 109, and 110. In alternative embodiments, the corresponding reference images for target image 108 may include images 107 and 109, or images 106 and 107, or images 109 and 119. In some embodiments, MCTF may refer to only one reference image, namely image 107 or image 109, to generate a filtered image as target image 108.

[0037] After selecting or otherwise determining a corresponding reference image for the target image, MCTF involves finding a resemblance of the target image in each reference image. This can be performed using block-based motion estimation (ME) and motion compensation (MC) techniques commonly used in inter-frame coding and decoding, especially those using block matching algorithms.

[0038] Specifically, the target image can be divided into multiple non-overlapping prediction blocks, each of which is a rectangular region of the target image. For each prediction block of the target image, a best-matching block of the same size is to be found in each reference image. An integer pixel search algorithm can be used to find the best-matching block within a specific search range of the reference images. As the word "search" suggests, the encoder examines all candidate blocks within this search range and then finds the candidate block that has the smallest amount of difference (e.g., lowest distortion) compared to the prediction block of the target image. For reference images adjacent to the target image, the candidate blocks are typically displacement versions of the prediction blocks. Each candidate block in the reference image has the same size as the prediction block (i.e., width and height). For integer pixel search, candidate blocks differ from each other by one pixel in either the horizontal or vertical direction.

[0039] To find the best-matching block, the encoder calculates the difference between each candidate block and the predicted block. A loss value is used to represent this difference; a smaller loss value indicates greater similarity. In some embodiments, an error matrix can be used to calculate the loss value, such as the sum of squared differences (SSD) or the sum of absolute differences (SAD) of all block pixels across a particular candidate block. The candidate block with the minimum loss value is the best-matching block and therefore the optimal matching block. Thus, the integer pixel search algorithm determines a corresponding MC result from each reference image for each predicted block, where the MC result includes the optimal matching block itself and the loss value associated with it.

[0040] Take the target image 108 as an example. Figure 2 Explain how to do it Figure 1 Perform an integer pixel search on the target image of 108. For example... Figure 2As shown, the target image 108 is divided into multiple prediction blocks, such as prediction blocks 211, 212, 213, 214, 215, 216, and 217. Furthermore, a set of reference images, namely source images 106, 107, 109, and 110, has been determined to generate a filtered image 208 for the target image 108. That is, source images 106, 107, 109, and 110 are reference images corresponding to the target image 108. The filtered image 208 can be one embodiment of the filtered image 188. For each prediction block of the target image 108, the encoder finds a corresponding block from each reference image. The corresponding block will be a block that is very similar to the prediction block. In fact, the corresponding block will be the block on the corresponding image that best matches the prediction block within a specific search range thereon.

[0041] For example, during precoding, the encoder searches for a single best-matching block that corresponds to the predicted block 217 on each of reference images 106, 107, 109, and 110. Specifically, to find each best-matching block, the encoder searches a rectangular region on each of reference images 106, 107, 109, and 110, where the rectangular region corresponds to the predicted block 217 and its surrounding area. The rectangular region, referred to as the "search range," is the same for each of reference images 106, 107, 109, and 110. Figure 2 The search range 269 on reference image 106 is searched to find the best-matching block on reference image 106 with the predicted block 217 of target block 108. Similarly, the search ranges 279 on reference image 107, 299 on reference image 10, and 209 on reference image 110 are searched respectively to find the best-matching blocks on reference images 107, 109, and 110. Accordingly, the best-matching blocks 263, 273, 293, and 203 are determined by integer pixel search algorithms for reference images 106, 107, 109, and 110, respectively. Figure 2 As shown, each of the best matching blocks 263, 273, 293 and 203 is within the corresponding search range, even though its position in the corresponding reference image may differ from the exact position of the predicted block 217 in the target image 108.

[0042] Typically, the search range will be larger than the size of the prediction block. Assuming that prediction block 217 has a size of 32x32 (i.e., 32 pixels wide and 32 pixels high), each of the search ranges 269, 279, 289, 299, and 209 can have a size of (32+delta) pixels multiplied by (32+delta) pixels, such as 43x43 or 50x50.

[0043] During precoding, the motion compensation step follows the integer pixel search step. This includes storing or otherwise buffering the results of the integer pixel search for best-matching blocks 263, 273, 293, and 203 for use in the filtering step following the motion compensation step. For example, the motion compensation step stores the MC result 262 output by the integer pixel search algorithm, since the algorithm has determined the best-matching block 263. The MC result 262 includes the best-matching block 263 itself (i.e., its pixel value), and a loss value 264 used to quantize or otherwise represent the difference between the predicted block 217 and the best-matching block 263. Various matrices can be used to compute the loss value 264, such as SSD, SAD, etc., as described above. Similarly, the motion compensation step also stores MC results 272, 292, and 202, which are output by the integer pixel search algorithm when the algorithm determines the best-matching blocks 273, 293, and 203, respectively. Figure 2 As shown, MC result 272 includes the best-matching block 273 and a loss value 274, which quantizes the difference between prediction block 217 and the best-matching block 273. Similarly, MC result 292 includes the best-matching block 293 and a loss value 294, which quantizes the difference between prediction block 217 and the best-matching block 293, while MC result 202 includes the best-matching block 203 and a loss value 204, which quantizes the difference between prediction block 217 and the best-matching block 203. All loss values, namely loss values ​​264, 274, 294, and 204, are calculated using the same loss calculation matrix and therefore can be meaningfully compared with each other.

[0044] The filtering step following the motion compensation step takes the MC results 262, 272, 292, and 202 as inputs and generates a filter block 287 of the filtered image 208 accordingly. The filtering step can employ pixel-based bilateral filtering (290), where each pixel of the filter block 287 can be calculated according to a weighted sum equation, as follows:

[0045]

[0046] Among them, F n Let B represent the value of the nth pixel in the filtered block 287, k represent the total number of reference images (i.e., reference images 106, 107, 109, and 110) corresponding to the target image 108, and B represent the value of the nth pixel in the filtered block 287. i,n w represents the value of the nth pixel of the best matching block found in the i-th reference image. i It is a real number, representing when B i,n The weighted sum of pixel values ​​from the i-th reference image. The weight is a real number. In some embodiments, the weight w iIt was determined using loss values ​​(i.e., loss values ​​264, 274, 294, and 204), and not all weighted w... i They can have the same value.

[0047] like Figure 2 As shown, the filtered image 208 includes multiple filtering blocks, such as filtering block 287, and filtering blocks 281, 282, 283, 284, 285, and 286. Each of the filtering blocks is generated in a manner similar to how filtering block 287 was generated as described above, thereby generating a filtered image 208 corresponding to the target image 108.

[0048] In some embodiments, time series 190 may include multiple target images, i.e., target images other than image 108. For each target image, a corresponding set of reference images may be determined, which are then used by the encoder to generate a filtered image corresponding to the respective target image in a manner similar to how the filtered image 208 was generated for target image 108.

[0049] In some embodiments, the encoder may determine a filtering interval and may select target images based on the filtering interval. Specifically, the filtering interval determines how often or how frequently filtered images are generated in a time series with multiple source images. For example, the encoder may determine an eight-frame filtering interval for time series 190. Accordingly, the precoding process will use an eight-frame increment to select target images. For example, for time series 190, source image 100 may also be selected as a target image in addition to source image 108. When selecting target images, POC numbers may be used in conjunction with the filtering interval, and each source image whose POC value is a multiple of the filtering interval is selected as the target image for applying MCTF. For example, an 8-frame filtering interval may result in source images with POC = 0, 8, 16, 24, 32, 40, 48, 56, 64, ... becoming target images. Similarly, a 10-frame filtering interval may result in source images with POC = 0, 10, 20, 30, 40, 50, 60, ... becoming target images. That is, the difference in POC between any two consecutive target images is equal to the filtering interval. In some embodiments, the encoder can use a default value or a predetermined value instead of based on any algorithm or any details of the time series to determine the filtering interval. For example, the encoder can use a default filtering interval of 8 frames for any time series without determining the filtering interval based on any details of the time series 190.

[0050] In some embodiments, a hierarchical pixel search method, including integer pixel search and fractional pixel search, can be used in the pre-encoding process. That is, one or more additional fractional pixel search steps can follow the integer pixel search steps, enabling the encoder to find better matching blocks compared to using only integer pixel search. Fractional pixel search operates similarly to integer pixel search, except that candidate blocks differ from each other by a fraction of a pixel in the horizontal or vertical direction. Furthermore, the search range can be adaptively adjusted to include the best matching block found from the integer pixel and its surrounding area. For example, if the encoder wants to perform a fractional pixel search after finding the best matching block 263, the search range used for the subsequent fractional pixel search can be adjusted to a rectangular area containing the best matching block 263 and some of its surrounding area.

[0051] In some embodiments, the size of the prediction block used for fractional pixel search can be smaller than the size of the prediction block used for integer pixel search. For example, each prediction block of target image 108 can be further divided into prediction blocks for fractional pixel search. Assuming prediction block 217 is 32x32 in size, when the encoder performs a fractional pixel search, prediction block 217 can be divided into four smaller blocks, each 16x16 in size, and each smaller block can be individually processed through motion compensation to find the best matching block of size 16x16 in each of reference images 106, 107, 109, and 110. To perform the fractional pixel search, the encoder needs to use interpolation techniques to generate fractional pixel values ​​using the pixel values ​​of the reference images. For example, if a 1 / 2 pixel (i.e., half-pixel) search follows an integer pixel search, the encoder will generate half-pixels in the reference images by interpolating the integer pixel values ​​of the reference images. Accordingly, candidate blocks differ from each other by 1 / 2 pixel in either the horizontal or vertical direction. Furthermore, if a 1 / 4 pixel search is followed by a 1 / 2 pixel search, the encoder will generate 1 / 4 pixel values ​​by interpolating the integer pixel values ​​and 1 / 2 pixel values ​​of the reference image, where candidate blocks differ from each other by 1 / 4 pixel in either the horizontal or vertical direction.

[0052] After pre-encoding, the encoder continues to encode the video into a bitstream. Instead of directly encoding the original source images of the video, the encoder encodes reference images and filtered images into the bitstream, leaving the target images in the bitstream. That is, the filtered images replace the target images during the encoding process. For example, when the encoder encodes time series 190 into a bitstream, the filtered image 208 will replace the source image 108. The filtered images will recover or otherwise inherit the POC value of their corresponding target images so that they replace the target images at their respective time positions in the video. For example, the filtered image 208 will therefore have a POC = 8, the same POC value as its corresponding target image—the source image 108.

[0053] As described above, replacing the target image with a corresponding filtered image generated from the precoding process helps achieve more efficient video coding, which typically manifests as a smaller bitstream size and / or higher subjective / objective video quality. In some embodiments of the MCTF precoding process described above, up to 6% codec gain can be achieved for 4K resolution video.

[0054] II. Sub-image-based precoding processing

[0055] In embodiments according to this disclosure, higher encoding / decoding gain and / or shorter processing time can be achieved for videos containing mixed source images. The mixed images include regions of natural images (NI) and regions of screen content images (SCI). Typically, natural images are images containing objects that exist in the real world, while screen content images contain computer-generated objects or text, such as screenshots, web pages, video game images, and other computer graphics.

[0056] Various video use cases now feature hybrid images within videos, where NI (National Interest) content and SCI (Single Content Context) content are presented in the same image. For example, during a televised sports broadcast, the television screen might display the stadium, athletes' performances, audience cheers, clouds in the sky, etc.—this is the NI content on the screen. Additionally, there might be SCI content, such as player statistics, scoreboards, advertising messages, and real-time text commentary from television viewers, presented simultaneously on the same screen as the NI content.

[0057] Applying MCTF indiscriminately to both the NI and SCI content in a frame is not the most efficient way to perform the precoding process described above. Our experimental data show that MCTF provides decent and satisfactory encoding / decoding gains for the NI content, while the gains from the MCTF-filtered SCI content are negligible. Therefore, a more efficient precoding process is achieved when MCTF uses more reference images for the NI content and fewer for the SCI content. In some embodiments, MCTF can apply reference images only to the NI content and not to the SCI content. That is, MCTF is applied only to the NI portion of the target image and not to its SCI portion, thus saving processing time, power consumption, and hardware overhead associated with performing MCTF on the SCI content, since performing these operations on the SCI content would result in minimal encoding / decoding gains.

[0058] In some embodiments, the savings (e.g., in terms of processing time and / or hardware overhead) from not applying MCTF to SCI content can be used for NI content to improve the resulting encoded video. For example, when applying MCTF to the NI content of a video, more frames can be included as reference pictures. Generally, the more reference frames used, the better matching blocks can be found, resulting in higher codec gain. Alternatively or additionally, the search range for integer or fractional pixel searches (e.g., search ranges of 269, 279, 299, and 209) can be larger, which also increases the chance of finding better matching blocks.

[0059] In a composite image, NI content and SCI content are typically presented in separate sub-images. A sub-image is a portion of an image. A sub-image can be one or more tiles and / or slices of an image. During video encoding and decoding, an image is typically divided into multiple codec tree blocks (CTBs), each CTB being a rectangular region of the image used as the basic unit for encoding or decoding. A slice is a fragment of an image formed by correlated (in raster scan order) CTBs, while a tile is a rectangular partition of the image containing multiple adjacent CTBs. Each tile or slice can be encoded, decoded, or otherwise processed independently. A composite image can include one or more NI sub-images and one or more SCI sub-images.

[0060] Figure 3 This is a diagram illustrating an example design based on an embodiment of this disclosure. Specifically, Figure 3The video, with a time series of 390, includes multiple mixed images, such as source images 305, 306, 307, 308, and 309. For pre-coding, the encoder determines that source image 308 is the target image. The encoder further determines that target image 308 is used with reference to source images 305, 306, 307, and 309 for MCTF (Multi-Channel Filtering). That is, source images 305, 306, 307, and 309 are a set of reference images used to generate the MCTF-filtered image for target image 308, as described above. Figure 2 As described. Figure 3 As shown, each of source images 305, 306, 307, 308, and 309 includes a first sub-image with NI content and a second sub-image with SCI content. For example, source image 305 has NI sub-image 315 and SCI sub-image 325. Similarly, source image 306 has NI sub-image 316 and SCI sub-image 326; source image 307 has NI sub-image 317 and SCI sub-image 327; source image 308 has NI sub-image 318 and SCI sub-image 328; and source image 309 has NI sub-image 319 and SCI sub-image 329. Figure 3 This is the filtered image 388 corresponding to the target image 308. The encoder generates the filtered image 388 by referencing only the NI sub-images (i.e., sub-images 315, 316, 317, and 319) of reference images 306, 307, 308, and 309, without referencing any of their SCI sub-images (i.e., sub-images 325, 326, 327, and 329). Specifically, when using the combination as described above... Figure 2 When applying the MCTF process as described, only the prediction blocks from NI sub-images 315, 316, 317, and 319 undergo the motion compensation and bilateral filtering steps. That is, the MC result with the bilateral filtering step as input comes only from the motion compensation performed on those prediction blocks of NI sub-images 315, 316, 317, and 319. Accordingly, the filtered image 388 includes the filtered NI sub-image 358, which is generated by the MCTF using reference sub-images 315, 316, 317, and 319 and the unfiltered SCI sub-image 328. In other words, the SCI content of the filtered image 388 is a direct copy of the SCI content of the original target image 308.

[0061] exist Figure 3In the example design shown, each image has one NI sub-image and one SCI sub-image, but the embodiments of the present invention are not limited to this. That is, a target image can have multiple NI sub-images and / or multiple SCI sub-images, and MCTF is applied to all of the multiple sub-images, but not to any single SCI sub-image. Since the MCTF precoding process is applied only to the NI sub-images and not the entire frame, the precoding process becomes more efficient, using less computational power, with less processing latency and / or shorter processing time. Furthermore, the hardware required to store the MC results (e.g., memory buffers) is smaller because the size of the best-matching block is only the same as the NI content, not the entire frame.

[0062] In some embodiments according to this disclosure, the encoder determines the number of reference images for the target image based on a filtering interval. Assume the encoder has determined the filtering interval to be N, meaning that every Nth image in the video stream is selected as the target image. If the target image contains only NI content, the encoder can determine how many N neighboring frames to reference when applying the MCTF to the target image. However, if the target image is a hybrid image containing both NI and SCI sub-images, the encoder can disable the MCTF for the SCI sub-images of the target image while maintaining the MCTF for the NI sub-images, which can be referenced by N neighboring frames when applying the MCTF to the target image. Alternatively, the encoder can even increase the number of reference images for the NI sub-images from N to N+k, where k is a positive integer.

[0063] III. Hardware Considerations for MCTF

[0064] As mentioned above, the precoding process using MCTF involves block-based pixel search, motion estimation, and motion compensation operations. As discussed later, these operations are also essential for the actual encoding of the video, particularly those performed by its inter-picture prediction module. Given that some of these similar functions are performed in both the precoding and actual encoding stages, and that the precoding process occurs before actual encoding, some hardware components within the inter-picture prediction module can be shared with it. That is, some hardware within the inter-picture prediction module, such as the integer motion estimation (IME) kernel and fractional motion estimation (FME) kernel, can be used instead of having separate hardware dedicated to MCTF precoding.

[0065] To share the IME and FEM designed for inter-picture prediction in video encoding, certain constraints may need to be imposed on the MCTF precoding process, some of which will be described in this section. In some embodiments, different numbers of reference frames may be determined for different types of target pictures in the video. Typically, video frames to be encoded into a bitstream fall into one of three frame types: intra-coding frames (also known as I-frames), predicted frames (also known as P-frames), and bi-directional predicted frames (also known as B-frames). I-frames use only spatial compression and no temporal information. That is, I-frames use only information within themselves and not information from other frames for motion estimation and motion compensation. Therefore, I-frames occupy the most bits in the video stream. In contrast, P-frames predict what has changed compared to previous (i.e., past) frames with smaller POC values, resulting in a combination of spatial and temporal compression. Therefore, P-frames provide better compression than I-frames. B-frames are similar to P-frames in using a combination of spatial and temporal compression, except that B-frames go further and reference past and future (in terms of POC) frames for motion estimation and motion compensation. Therefore, B-frames typically offer the highest compression ratio and occupy the fewest bits in a video stream compared to P-frames and I-frames.

[0066] In some embodiments according to this disclosure, the encoder determines the number of reference images for the target image based on the filtering interval. Assume the encoder has determined the filtering interval to be N, meaning every Nth image in the video stream is selected as the target image. Based on the frame type of each target image, the encoder determines the corresponding number of reference images to be used for the MCTF step. In some embodiments, an I-frame target image has N adjacent frames as its MCTF reference images, while a P-frame target image has N / 2 adjacent frames as its MCTF reference images. Furthermore, for B-frame target images, MCTF is disabled. That is, the encoder does not pre-encode B-frames, and therefore no reference images are determined for them. For example, if the video filtering interval is determined to be every eight images (i.e., N = 8), then source images with POC = 0, 8, 16, 24, 32, 40, 48, 56, 64, ... are selected as target images. Further assume that the frame with POC = 32 is an I-frame, the frame with POC = 16 is a P-frame, and the frames with POC = 8 and 24 are B-frames. Therefore, the encoder can determine that the target image with POC=32 has 8 MCTF reference images, namely, adjacent frames with POC=28, 29, 30, 31, 33, 34, 35, and 36. Furthermore, the encoder can determine that the target image with POC=16 has four MCTF reference images, namely, adjacent frames with POC=14, 15, 17, and 18. The encoder can further determine that MCTF is not applied to the target frames with POC=8 and 24, and therefore their number of reference images is zero.

[0067] In some embodiments, the determination of reference images for the target image can be customized to improve the hardware overhead and / or processing latency caused by the MCTF precoding process. Figure 4 This is a diagram illustrating an example design based on an embodiment of this disclosure. Reference Figure 4Figure 410 illustrates the timing diagram of a video stream with MCTF precoding disabled. The first row of Figure 410 represents the source pictures in the timing diagram, labeled POC = 0, 1, 2, 3, 4, 5, ..., 32. The second row of Figure 410 shows that MCTF precoding is disabled. Furthermore, the third row of Figure 410 shows that the encoder is idle and does not begin actual encoding of picture frames until a frame with POC = 32 has been received. This is because the group-of-pictures (GOP) size of the video is 32 frames. That is, frames with POC = 0-31 belong to the same GOP. A GOP is a collection of consecutive pictures in an encoded video stream. The encoded video stream consists of consecutive GOPs of the same size, and each GOP is an independent encoding and decoding unit because motion estimation and motion compensation for all inter-coded frames only refer to the pictures within the GOP, plus the first frame of the next GOP. Therefore, the encoder cannot begin encoding a GOP consisting of frames with POC = 0–31 until all frames with POC = 0–32 have been acquired by the encoder. A GOP has one I-frame, which is the first frame of the GOP. Therefore, the GOP size of a video stream is the distance between two consecutive I-frames (measured in frames). From the decoder's perspective, encountering a new GOP in an encoded video stream means that the decoder does not need any previous frames to decode the subsequent frames.

[0068] Figure 420 illustrates a timing diagram of the same video stream as in Figure 410, but with MCTF precoding enabled. For the MCTF process, the filtering interval is 8 frames. As shown in Figure 420, the picture frames POC = 0, 8, 16, 24, 32, and 40 are differently labeled to indicate that they are target pictures (i.e., filtering interval = 8 frames). Furthermore, one or more adjacent frames around each picture frame are also differently labeled to show the corresponding set of reference pictures for the MCTF reference of the respective target picture. For example, target picture POC = 0 has two frames POC = 1 and 2 as its reference pictures, while each of the remaining target pictures in Figure 420 has four reference pictures. For example, target picture POC = 8 has frames POC = 6, 7, 9, and 10 as its reference pictures, as labeled in Figure 420. The second row of Figure 420 shows the activity of the MCTF precoding process for the video. MCTF is idle when the encoder receives frames POC = 0–2 as camera input. Upon receiving frames POC=0–2, the encoder can begin MCTF processing for the target image POC=0 because this is the earliest time when all the reference images required to process the POC=0 frame are in the encoder, i.e., frames POC=1 and 2 are already in the encoder. As shown in Figure 420, the MCTF hardware takes four frames to complete MCTF processing for the POC=0 target image. That is, when the camera input sends frames POC=3–6, the encoder is processing the target frame POC=0 for MCTF, and completes processing when it sends the frame POC=6. Subsequently, during the period when the encoder receives frames POC=7–10, the MCTF hardware is idle again because the encoder needs information from frames POC=6, 7, 9, and 10 to process the target image POC=8, as these frames are its reference images. The target image POC=8 requires 8 frames to complete MCTF, twice the time required for the target image POC=0. This is because the target image POC=0 has only two reference images, while the target image POC=8 has four.

[0069] Comparing Figures 410 and 420 reveals a 10-frame delay, as described below. In Figure 410 with MCTF precoding disabled, the encoder can begin actual video encoding after receiving frames POC=32. However, when MCTF precoding is enabled, encoding must begin much later. Assuming POC=32 is the target image, actual video encoding cannot begin until the corresponding filtered image for the target image POC=32 is generated. As shown in Figure 420, MCTF is being performed on the target image POC=32 when source images POC=35-42 are received. Therefore, the earliest the encoder can begin encoding a GOP containing frames POC=0–31 is the next cycle, i.e., when frames POC=43 are being received. Compared to Figure 410, the 10-frame delay 412 is introduced due to the enabling of MCTF precoding. The delay caused by MCTF directly translates into hardware costs. That is, a memory buffer is therefore needed to temporarily store these frames during the delay period when they arrive, so that they can be referenced in the buffer later when needed. For example, in the scenario depicted in Figure 420, when the encoder receives source images with POC = 33–42, it needs a memory buffer to temporarily store these source images because the MCTF hardware is being pre-encoded with target images with POC = 24 and 32 during this period. The memory buffer holds the source images with POC = 33–42 until the MCTF hardware is ready to use them, for example, when the MCTF hardware processes the target images with POC = 32 and 40.

[0070] IV. Considerations for MCTF Delay

[0071] Various modifications can be made to the MCTF precoding process described elsewhere above to reduce the resulting latency / delay and thereby infer the required memory buffer size. In some embodiments, the encoder may determine a smaller number of reference images for one or more target images near the end of the GOP. Figure 5 This includes diagrams illustrating example designs based on embodiments of this disclosure. Figure 5Figure 520 illustrates an embodiment where a target image whose POC value is a multiple of the video's GOP size has fewer reference images compared to another target image whose POC value is not a multiple of the GOP size. Specifically, the video shown in Figure 520 has a GOP size of 32 frames, and the filtering interval is determined to be every 8 frames. For target images with POC = 8, 16, and 24, since their POC values ​​are not multiples of the GOP size, each target image has four reference images, including two past frames and two future frames. For the target image with POC = 32, there are only two reference images, namely frames with POC = 31 and 33. Therefore, MCTF processing for the target frame with POC = 32 can be completed in just four frames. Compared to the scenario in Figure 410, this embodiment pulls the actual encoding start of the frame with POC = 0 to be aligned with the reception of the source image with POC = 39. This embodiment reduces the latency from 10 frames (represented by latency 412) to 6 frames (represented by latency 512). The reduced latency also translates directly to a smaller memory buffer required, since the memory buffer for the scenario described in Figure 520 only needs to buffer frames with POC = 33–38, instead of frames with POC = 33–42 as shown in Figure 420.

[0072] Figure 6 This includes a diagram of another example design based on an embodiment of this disclosure. See also... Figure 6 Figure 620 illustrates an embodiment where a target image whose POC value is a multiple of the video's GOP size has fewer reference images compared to another target image whose POC value is not a multiple of the GOP size. Specifically, the video shown in Figure 620 has a GOP size of 32 frames, and the filtering interval is determined to be every 8 frames. For target images with POC = 8, 16, and 24, since their POC values ​​are not multiples of the GOP size, each target image has four reference images, including two past frames and two future frames. For the target image with POC = 32, only two past frames serve as its reference images, namely frames with POC = 30 and 31. Thus, MCTF processing for the target frame with POC = 32 only requires four frames to complete. Compared to the scenario in Figure 410, this embodiment pulls the actual encoding of the frame with POC = 0 to be aligned with the received image of the source image with POC = 39, where the actual encoding of the frame with POC = 0 is aligned with the frame with POC = 43. This embodiment reduces latency from 10 frames (represented by latency 412) to 6 frames (represented by latency 612). The reduced latency also translates directly to a smaller memory buffer required, since the memory buffer for the scenario depicted in Figure 620 only needs to buffer frames with POC = 33–38, instead of frames with POC = 33–42 as shown in Figure 420.

[0073] Another way to reduce latency / wait time caused by MCTF precoding is to employ hardware parallelism, with two implementations in... Figure 7 As shown in the image. Figure 7 An example utilizing hardware parallelism is shown. (Reference) Figure 7 Figure 710 illustrates MCTF processing using twice the hardware parallelism for the same target image and corresponding reference image set presented in Figure 420. This doubled hardware parallelism halves the processing time for generating the MCTF-filtered image. Comparing Figures 420 and 710, due to hardware parallelism, the time spent processing the target frame with POC=0 is reduced from four frames to two. Similarly, the time spent processing MCTF for each of the target frames with POC=8, 16, 24, and 32 is also halved, from eight frames in Figure 420 to four frames in Figure 710. Therefore, the actual video encoding start time is reduced by four frames, from receive alignment with POC=43 as shown in Figure 420 to receive alignment with POC=39 as shown in Figure 710. Compared to Figure 410 with MCTF disabled, using hardware parallelism reduces the latency caused by MCTF in Figure 710 to six frames, as shown in Delay 711.

[0074] Similarly, Figure 720 illustrates MCTF processing using twice the hardware parallelism for the same target images and corresponding reference image sets presented in Figure 620. Comparing Figures 620 and 720, due to hardware parallelism, the time spent processing each of the target frames POC=0 and 32 is reduced from four frames to two frames. Likewise, the time spent processing each of the target frames POC=8, 16, and 24 is also halved, from eight frames in Figure 620 to four frames in Figure 720. Therefore, the start of actual video encoding, from receive alignment with POC=39 as shown in Figure 420, to receive alignment with POC=35 as shown in Figure 720, is reduced by four frames. Compared to Figure 410 with MCTF disabled, using hardware parallelism reduces the latency caused by MCTF in Figure 720 to two frames, as shown in latency 712.

[0075] V.MCTF codec gain considerations

[0076] As shown in each of Figures 420, 520, 620, 710, and 720, the MCTF hardware is not always busy but has intermittent idle times. In some embodiments, the encoder can reduce the idle time of the MCTF hardware by including more reference images when generating a filtered image for the target image. This approach can achieve more comprehensive MCTF hardware utilization without introducing additional latency. By including more reference images, a better filtered image can be obtained, thereby improving the encoder's coding gain.

[0077] In some embodiments, the encoder is designed to include additional reference images for the target image with POC=0. Figure 8 Two such embodiments are illustrated, as shown in Figures 810 and 820, respectively. Figure 810 is identical to Figure 420, except that the target image with POC=0 includes an additional reference image, a frame with POC=3. That is, the MCTF hardware references three reference images, namely frames with POC=1-3, when generating the filtered image for the target image with POC=0. As shown in Figure 810, the start of actual video encoding is still aligned with the reception of the frame with POC=43, as in Figure 420, thus no additional latency is introduced. Given that the number of reference images for the target image with POC=0 increases from 2 to 3, a better filtered image is expected to be encoded into the bitstream instead of the target image with POC=0, thus the bitstream is expected to have better (i.e., greater or more) encoding / decoding gain. This also reduces the MCTF hardware idle time immediately following the processing of the target image with POC=0, from four frames shown in Figure 420 to only one frame shown in Figure 810.

[0078] Similarly, Figure 820 is identical to Figure 710, except that the target image with POC=0 includes two additional reference images, namely frames with POC=3 and 4. That is, the MCTF hardware references four reference images, namely frames with POC=1-4, when generating the filtered image for the target image with POC=0. As shown in Figure 820, the start of actual video encoding remains aligned with the reception of frame POC=39, as in Figure 710, thus introducing no additional latency. Given that the number of reference images for the target image with POC=0 increases from 2 to 4, a better filtered image is expected to be encoded into the bitstream instead of the target image with POC=0, thus the bitstream is expected to have even better (i.e., larger or more) encoding / decoding gain than in Figure 810. This also reduces the MCTF hardware idle time immediately following the processing of the target image with POC=0, from six frames shown in Figure 710 to two frames shown in Figure 820.

[0079] In some embodiments according to this disclosure, the encoder determines the number of reference images for the target image based on a filtering interval. Assume the encoder has determined the filtering interval to be N, meaning that every Nth image in the video stream is selected as the target image. Specifically, for target images other than frames with POC = 0, the encoder can determine N adjacent frames as MCTF reference images for the target image, where half of the N reference images have a POC value less than the target image's POC value, and the other half have a POC value greater than the target image's POC value. For a target image with POC = 0, the encoder can determine (N / 2 + k) frames following the target image with POC = 0 as its MCTF reference images, where k is a positive integer.

[0080] In some embodiments, the encoder aims to improve target encoding / decoding gain by including more relevant reference frames for the target image, especially when a scene change occurs relatively close to the target image. In a video stream, frames presented before a theme change have a different "theme" than frames that appear after a theme change. Therefore, the content of frames before a theme change is generally completely unrelated to the content of frames after a theme change. Due to the uncorrelation between the two sets of frames, applying MCTF to a set of reference frames that includes both before and after a theme change does not result in much encoding / decoding gain, as frames before a theme change do not help predict any frames that appear after a theme change, and vice versa. Depending on whether the target image is presented before or after a theme change, MCTF will be more effective in referencing frames that are presented only before or after the theme change. That is, if a theme change occurs between the target image and one of its reference images, MCTF will be less effective. On the other hand, MCTF is effective when there is no theme change between the target image and all of its reference images.

[0081] Figure 9 An example design according to an embodiment of this disclosure is illustrated. References Figure 9Figure 910 illustrates the same embodiment as that of Figure 710, except that the video in Figure 910 includes a scene change 915 immediately preceding the target image at POC=8. In Figure 710, the encoder identifies frames POC=6, 7, 9, and 10 as the MCTF reference images for the target image at POC=8. However, this is not ideal for the scene in Figure 910 because the content in frames POC=6 and 7 is not very relevant to the target image at POC=8 due to scene change 915. Instead, as shown in Figure 910, the encoder identifies frames POC=9, 10, 11, and 12 as the MCTF reference images for the target image at POC=8 because there is no scene change between the target image at POC=8 and the reference images at POC=9, 10, 11, and 12. This results in better encoding / decoding gain.

[0082] Similarly, Figure 920 illustrates the same embodiment as the one in Figure 710, except that the video in Figure 920 includes a scene change 925 immediately following the target image with POC=8. In Figure 710, the encoder identifies frames with POC=6, 7, 9, and 10 as MCTF reference images for the target image with POC=8. However, this is not ideal for the scene in Figure 910 because the content in frames with POC=9 and 10 is not very relevant to the target image with POC=8 due to scene change 925. Instead, as shown in Figure 920, the encoder identifies frames with POC=4, 5, 6, and 7 as MCTF reference images for the target image with POC=8 because no scene change is introduced between the target image with POC=8 and the reference images with POC=4, 5, 6, and 7. This results in better encoding / decoding gain.

[0083] VI. Implementation of the Example

[0084] Figure 10An example video encoder 1000 is shown. As shown, the video encoder 1000 receives an input video signal from a video source 1005 and encodes the signal into a bitstream 1095. The video encoder 1000 has several components or modules for encoding the signal from the video source 1005, including at least some components selected from the precoding processing module 1080, transform module 1010, quantization module 1011, inverse quantization module 1014, inverse transform module 1015, intra-picture estimation module 1020, intra-prediction module 1025, motion compensation module 1030, motion estimation module 1035, loop filter 1045, reconstructed picture buffer 1050, motion vector (MV) buffer 1065, MV prediction module 1075, and entropy encoder 1090. The motion compensation module 1030 and motion estimation module 1035 are part of the inter-frame prediction module 1040. Inter-frame prediction module 1040 may include an integer motion estimation (IME) kernel configured to perform integer pixel search and a fractional motion estimation (FME) kernel configured to perform fractional pixel search. Both integer pixel search and fractional pixel search are fundamental functions of motion compensation module 1030 and motion estimation module 1035. Precoding processing module 1080 may be an embodiment of precoding processing module 132, while the remainder of video encoder 1000 may collectively embody video encoder 134.

[0085] In some embodiments, modules 1010-1090 as listed above are software instruction modules executed by one or more processing units (e.g., processors) of a computing device or electronic device. In some embodiments, modules 1010-1090 are hardware circuit modules implemented by one or more integrated circuits (ICs) of an electronic device. Although modules 1010–1090 are shown as separate modules, some modules may be combined into a single module.

[0086] Video source 1005 provides the raw video signal, which presents the pixel data of each video frame without compression. That is, video source 1005 provides a video stream comprising source images presented in a time series. Precoding processing module 1080 takes the video stream as input and performs precoding MCTF processing according to one or more embodiments described elsewhere above. The processed video data, including all source images not selected as target images for precoding MCTF processing, and filtered images that replace the target images in the time series, is sent to other modules of video encoder 1000 for actual video encoding.

[0087] Subtractor 1008 calculates the difference between the processed video data generated by precoding module 1080 and the predicted pixel data 1013 from motion compensation module 1030 or intra-frame prediction module 1025. Transform module 1010 converts the difference (or residual pixel data or residual signal 1009) into transform coefficients 1016 (e.g., by performing a discrete cosine transform, also known as DCT). Quantization module 1011 quantizes the transform coefficients into quantized data (or quantized coefficients) 1012, which is encoded into a bitstream 1095 by entropy encoder 1090.

[0088] The inverse quantization module 1014 dequantizes the quantized data (or quantized coefficients) 1012 to obtain transform coefficients, and the inverse transform module 1015 performs an inverse transform on the transform coefficients to generate a reconstruction residual 1019. The reconstruction residual 1019 is added to the predicted pixel data 1013 to generate reconstructed pixel data 1017. In some embodiments, the reconstructed pixel data 1017 is temporarily stored in a line buffer (not shown) for intra-image prediction and spatial MV prediction. The reconstructed pixels are filtered by a loop filter 645 and stored in a reconstructed image buffer 1050. In some embodiments, the reconstructed image buffer 1050 is external memory to the video encoder 1000. In some embodiments, the reconstructed image buffer 1050 is internal memory to the video encoder 1000.

[0089] Intra-image estimation module 1020 performs intra-frame prediction based on reconstructed pixel data 1017 to generate intra-frame prediction data. The intra-frame prediction data is provided to entropy encoder 1090 to be encoded into a bitstream 1095. The intra-frame prediction data is also used by intra-frame prediction module 1025 to generate predicted pixel data 1013.

[0090] The motion estimation module 1035 performs inter-frame prediction by generating MVs (Motion Values) to reference pixel data of previously decoded frames stored in the reconstructed image buffer 1050. These MVs are provided to the motion compensation module 1030 to generate predicted pixel data.

[0091] Instead of encoding the complete actual MV in the bitstream, the video encoder 1000 uses MV prediction to generate a predicted MV, and the difference between the MV used for motion compensation and the predicted MV is encoded as residual motion data and stored in the bitstream 1095.

[0092] The MV prediction module 1075 generates a predicted MV based on a reference MV generated for encoding a previous video frame, i.e., a motion-compensated MV used to perform motion compensation. The MV prediction module 1075 retrieves the reference MV from the previous video frame from the MV buffer 1065. The video encoder 1000 stores the MV generated for the current video frame in the MV buffer 1065 as a reference MV for generating the predicted MV.

[0093] The MV prediction module 1075 uses a reference MV to create a predicted MV. The predicted MV can be calculated by spatial MV prediction or temporal MV prediction. The entropy encoder 1090 encodes the difference (residual motion data) between the predicted MV of the current frame and the motion-compensated MV (MC MV) into the bitstream 1095.

[0094] The entropy encoder 1090 encodes various parameters and data into a bitstream 1095 using entropy encoding techniques such as context-adaptive binary arithmetic codec (CABAC) or Huffman coding. The entropy encoder 1090 encodes various header elements, flags, along with quantized transform coefficients 1012 and residual motion data, as syntax elements into the bitstream 1095. The bitstream 1095 is then stored in a storage device or transmitted to the decoder via a communication medium such as a network.

[0095] The loop filter 1045 performs filtering or smoothing operations on the reconstructed pixel data 1017 to reduce encoding / decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering operation includes sample adaptive offset (SAO). In some embodiments, the filtering operation includes adaptive loop filter (ALF).

[0096] Figure 11This section describes a portion of the precoding processing module 1080 implemented by the video encoder 1000. As shown, the precoding processing module 1080 includes a processor 1110, a target image buffer 1105, a reference buffer 1120, a motion compensation module 1130, an MC result buffer 1140, a bilateral filtering module 1150, and a filtered image buffer 1160. The processor 1110 is configured to receive and analyze the raw video stream of raw pixel frames from the video source 1005 to identify or parse certain parameters of the video stream, such as GOP size. The processor 1110 can also identify whether each frame includes NI sub-images and SCI sub-images. The processor 1110 can further identify the location of scene change events in the video stream (if any). Based on the analysis of the raw video data, the processor 1110 can determine the MCTF filtering interval and multiple target images to which MCTF is to be applied. The processor 1110 can store multiple target images in the target image buffer 1105. Additionally, the processor 1110 can determine one or more reference images to be used for MCTF for each target image. The processor 1110 can store one or more reference images for each target image in the reference image buffer 1120.

[0097] Motion compensation module 1130 is configured to access reference image buffer 1120 and perform motion estimation (ME) and motion compensation (MC) operations on the reference image to produce MC results for the target image, as described elsewhere above herein. Before performing the ME and MC operations, motion compensation module 1130 may divide each target image and each corresponding reference image into multiple prediction blocks. Motion compensation module 1130 may include an integer motion estimation (IME) kernel 1132 configured to perform an integer pixel search to find the best matching block for the prediction blocks in the reference image. Motion compensation module 1130 may also include a fractional motion estimation (FME) kernel 1134 configured to perform a fractional pixel search (e.g., a 1 / 2 pixel search or a 1 / 4 pixel search) to find the best matching block for the prediction blocks in the reference image. Motion compensation module 1130 may perform ME and MC operations by involving IME kernel 1132 and / or FME kernel 1134. In some embodiments, the video encoder 1000 may share or reuse the same circuitry or hardware used as the IME kernel 1132 and the IME kernel within the inter-frame prediction module 1040. Similarly, the video encoder 1000 may share or reuse the same circuitry or hardware used as the FME kernel 1134 and the FME kernel within the inter-frame prediction module 1040. The MC results generated by the motion compensation module 1130 may be stored in the MC result buffer 1140.

[0098] The bilateral filtering module 1150 can access the MC result buffer 1140 and perform pixel-by-pixel bilateral filtering on the MC result accordingly, thereby generating a filtered image for each target image, as described elsewhere above. The generated filtered image can be stored in the filtered image buffer 1160. The reference image stored in the reference image buffer 1120 and the filtered image stored in the filtered image buffer 1160 are then encoded into the bitstream 1095 by other modules of the video encoder 1000.

[0099] Figure 12 An example video decoder 1200 is shown. As shown, the video decoder 1200 is an image decoding or video decoding circuit that receives a bitstream 1295 and decodes the contents of the bitstream 1295 into pixel data of video frames for display. The video decoder 1200 has several components or modules for decoding the bitstream 1295, including some components selected from the inverse quantization module 1211, inverse transform module 1210, intra-frame prediction module 1225, motion compensation module 1230, loop filter 1245, decoded image buffer 1250, MV buffer 1265, MV prediction module 1275, and parser 1290. The motion compensation module 1230 is part of the inter-frame prediction module 1240.

[0100] In some embodiments, modules 1210-1290 are software instruction modules executed by one or more processing units (e.g., processors) of a computing device. In some embodiments, modules 1210-1290 are hardware circuit modules implemented by one or more ICs of an electronic device. Although modules 1210-1290 are shown as separate modules, some modules may be combined into a single module.

[0101] Parser (e.g., entropy decoder) 1290 receives bitstream 1295 and performs initial parsing according to the syntax defined by the video codec or image codec standard. The parsed syntactic elements include various header elements, flags, and quantized data (or quantization coefficients) 1212. Parser 1290 parses the various syntactic elements using entropy coding techniques, such as context-adaptive binary arithmetic codec (CABAC) or Huffman coding.

[0102] The inverse quantization module 1211 dequantizes the quantized data (or quantized coefficients) 1212 to obtain transform coefficients, and the inverse transform module 1210 performs an inverse transform on the transform coefficients 1216 to generate a reconstructed residual signal 1219. The reconstructed residual signal 12112 is added to the predicted pixel data 1213 from the intra-frame prediction module 1225 or the motion compensation module 1230 to generate decoded pixel data 1217. The decoded pixel data is filtered by the loop filter 1245 and stored in the decoded image buffer 1250. In some embodiments, the decoded image buffer 1250 is external memory to the video decoder 1200. In some embodiments, the decoded image buffer 1250 is internal memory to the video decoder 1200.

[0103] Intra-frame prediction module 1225 receives intra-frame prediction data from bitstream 1295 and accordingly generates predicted pixel data 1213 from decoded pixel data 1217 stored in decoded image buffer 1250. In some embodiments, the decoded pixel data 1217 is also stored in a line buffer (not shown) for intra-image prediction and spatial MV prediction.

[0104] In some embodiments, the contents of the decoded image buffer 1250 are used for display. The display device 1255 either retrieves the contents of the decoded image buffer 1250 for direct display or retrieves the contents of the decoded image buffer into a display buffer. In some embodiments, the display device receives pixel values ​​from the decoded image buffer 1250 via pixel transfer.

[0105] The motion compensation module 1230 generates predicted pixel data 1213 from the decoded pixel data 1217 stored in the decoded image buffer 1250 based on the motion compensation MV (MC MV). These motion compensation MVs are decoded by adding the residual motion data received from the bitstream 1295 to the predicted MV received from the MV prediction module 1275.

[0106] The MV prediction module 1275 generates a predicted MV based on a reference MV generated for decoding a previous video frame, such as a motion-compensated MV used to perform motion compensation. The MV prediction module 1275 retrieves the reference MV of the previous video frame from the MV buffer 1265. The video decoder 1200 stores the motion-compensated MV generated for decoding the current video frame in the MV buffer 1265 as a reference MV for generating the predicted MV.

[0107] The loop filter 1245 performs filtering or smoothing operations on the decoded pixel data 1217 to reduce encoding / decoding artifacts, particularly at pixel block boundaries. In some embodiments, the filtering operation includes Sample Adaptive Offset (SAO). In some embodiments, the filtering operation includes Adaptive Loop Filter (ALF).

[0108] VII. Process Description

[0109] Figure 13 Example processes 1300 and 1305 according to embodiments of the present disclosure are illustrated. Processes 1300 and 1305 may each represent an aspect of implementing the various proposed designs, concepts, schemes, systems, and methods described above. More specifically, processes 1300 and 1305 may each represent an aspect of proposed concepts and schemes related to a precoding process of a video stream according to the present disclosure. Processes 1300 and 1305 may each include one or more operations, actions, or functions, as shown in one or more of blocks 1310, 1320, 1330, 1340, 1350, 1360, 1370, and 1380. Depending on the desired implementation, process 1300 or process 1305 may be divided into additional blocks, combined into fewer blocks, or eliminated. Furthermore, the blocks / sub-blocks of processes 1300 and 1305 may be arranged in... Figure 13 The processes can be executed in the order shown, or in a different order. Furthermore, one or more blocks / subblocks of process 1300 or process 1305 can be executed repeatedly or iteratively. Each of processes 1300 and 1305 can be implemented by or within device 1000 and any variations thereof. For illustrative purposes only and without limitation, processes 1300 and 1305 are described in the context of device 1000 as precoding processing unit 1080. Process 1300 may begin with block 1310. Process 1305 may also begin with block 1310.

[0110] At 1310, processes 1300 and 1305 may involve the processor 1110 of apparatus 1080 receiving a video stream having source images presented in a time sequence (e.g., recorded or displayed). Each source image may be associated with a temporal identifier (e.g., a POC value) that identifies the time position of the source image in the time sequence. The processor 1110 may accordingly determine a filtering interval, expressed as repeating at certain intervals of frames, based on which MCTF is applied. Process 1300 may proceed from 1310 to 1320. Process 1305 may also proceed from 1310 to 1320.

[0111] At 1320, processes 1300 and 1305 may involve processor 1110 determining or selecting multiple target images based on a filtering interval. Each target image is a source image of the video stream. The target images may be stored in a target image buffer 1105. Process 1300 can proceed from 1320 to 1330. Process 1305 can proceed from 1320 to 1360.

[0112] At 1330, process 1300 may involve processor 1110 analyzing target images stored in target image buffer 1105 and, for each target image, searching for or otherwise identifying regions containing natural image (NI) and regions containing screen content image (SCI). In some implementations, a region may be one or more sub-images. A sub-image may be one or more slices or tiles composed of multiple adjacent CTUs. Process 1300 may proceed from 1330 to 1340.

[0113] At 1340, process 1300 may involve processor 1110 identifying the GOP size of the video stream. Process 1300 can proceed from 1340 to 1350.

[0114] At 1350, process 1300 may involve the processor identifying locations of scene changes in the video stream (if any). Process 1300 can proceed from 1350 to 1360.

[0115] At 1360, processes 1300 and 1305 may involve processor 1110 determining one or more reference images for each target image. Process 1300 may proceed from 1360 to 1370. Process 1305 may also proceed from 1360 to 1370.

[0116] In some implementations, when determining one or more reference images for each target image, the processor 1110 can generate different numbers of reference images for different target images based on whether the POC value of the target image is a multiple of the GOP size. If the POC value is a multiple of the GOP size, the processor 1110 can determine fewer reference images for the target image. If the POC value is not a multiple of the GOP size, the processor 1110 can determine more reference images for the target image.

[0117] In some implementations, when the POC value of the target image is a multiple of the GOP size, the processor 1110 can determine that the reference image for the target image includes only past frames compared to the target image, i.e., frames with a POC value less than the POC value of the target image. In other words, only frames with a time position earlier than the target image in the time series of the video stream can be considered reference images for the target image.

[0118] In some implementations, if a scene change occurs immediately before the target image in the video's time series, the processor 1110 can determine that the reference image for the target image includes only future frames compared to the target image, i.e., frames with a POC value greater than the target image's POC value. In other words, only frames with a time position later than the target image in the video stream's time series can be reference images for the target image.

[0119] In some implementations, if a scene change occurs immediately following the target image in the video's time sequence, the processor 1110 can determine that the reference image for the target image includes only frames that are earlier than the target image, i.e., frames with a POC value less than the target image's POC value. In other words, only frames with a time position earlier than the target image in the video stream's time sequence can be considered reference images for the target image.

[0120] In some implementations, the processor 1110 may determine a reference image for the target image such that no scene change event occurs between the target image and the reference image determined for the target image.

[0121] At 1370, processes 1300 and 1305 may involve apparatus 1080 generating a filtered image for each target image. Generating the filtered image involves apparatus 1080 performing pixel-based filtering (e.g., MCTF) using a reference image corresponding to the respective target image. Processes 1300 and 1305 may also involve apparatus 1080 storing the filtered image generated by bilateral filtering module 1150 into a filtered image buffer 1160. Process 1300 can proceed from 1370 to 1380. Process 1305 can also proceed from 1370 to 1380.

[0122] In some implementations, when generating a filtered image, the apparatus 1080 applies pixel-based filtering only to the NI sub-images and not to the SCI sub-images. That is, pixel-based filtering applies only to the NI sub-images of the target image, and not to any SCI sub-images of the target image.

[0123] In some implementations, when generating each filtered image, the device 1080 applies pixel-based filtering to both the NI and SCI sub-images. However, the device 1080 references more reference images for the NI sub-image of the target image, but fewer reference images for the SCI sub-image.

[0124] In some implementations, when generating the filtered image, processes 1300 and 1305 may involve motion compensation module 1130 dividing each target image into multiple prediction blocks. The prediction blocks may have the same size or different sizes. Processes 1300 and 1305 may also involve motion compensation module 1130 determining multiple MC results for each prediction block, wherein each MC result is determined based on a corresponding reference image corresponding to a reference image of the target image. Processes 1300 and 1305 may also involve motion compensation module 1130 performing bilateral filtering on each pixel of each prediction block, wherein the bilateral filtering is performed based on the determined MC results.

[0125] In some implementations, each MC result of the prediction block includes a best-matching block and a loss value. The best-matching block may have the same width and height as the corresponding prediction block. Furthermore, the loss value may represent the difference between the best-matching block and the corresponding prediction block. Additionally, processes 1300 and 1305 may involve a bilateral filtering module 1150 performing bilateral filtering by calculating a weighted sum for each pixel of the prediction block based on the corresponding pixel value of the best-matching block and the loss value of the MC.

[0126] In some implementations, when determining the MC result for each prediction block, processes 1300 and 1305 may involve motion compensation module 1130 performing an integer pixel search, fractional pixel search, or both based on the corresponding prediction block and one or more reference images.

[0127] At 1380, processes 1300 and 1305 may involve video encoder 1000 encoding video into bitstream 1095. Specifically, video encoder 1000 may encode filtered images generated by device 1080 and stored in filtered image buffer 1160. Video encoder 1000 may also encode source images that are not determined to be target images into bitstream 1095.

[0128] Figure 14 An example process 1400 according to an embodiment of the present disclosure is illustrated. Process 1400 may represent one aspect of implementing the various proposed designs, concepts, schemes, systems, and methods described above. More specifically, process 1400 may represent one aspect of proposed concepts and schemes related to a precoding process for a video stream according to the present disclosure. Process 1400 may include one or more operations, actions, or functions as shown in one or more of blocks 1410, 1420, 1430, 1440, 1450, 1460, 1470, and 1480. Although shown as discrete blocks, the individual blocks of process 1400 may be divided into additional blocks, combined into fewer blocks, or eliminated, depending on the desired implementation. Furthermore, the blocks / sub-blocks of process 1400 may be arranged in... Figure 14 The process can be executed in the order shown, or in a different order. Furthermore, one or more blocks / sub-blocks of process 1400 can be executed repeatedly or iteratively. Process 1400 can be implemented by or within device 1080 or any variant thereof. For illustrative purposes only and without limitation, process 1400 is described in the context of device 1000 as precoding processing unit 1080. Process 1400 may begin with block 1410.

[0129] At 1410, process 1400 may involve the processor 1110 of device 1080 determining the filtering interval for each N frames. Process 1400 may proceed from 1410 to 1420.

[0130] At 1420, process 1400 may involve motion compensation module 1130 retrieving the target image from target image buffer 1105. Process 1400 can proceed from 1420 to 1430.

[0131] At 1430, process 1400 may involve processor 1110 determining whether the target image is an I-frame. If the target image is an I-frame, process 1400 can proceed from 1430 to 1440. If the target image is not an I-frame, process 1400 can proceed from 1430 to 1450.

[0132] At 1440, process 1400 may involve device 1080 performing MCTF on N reference images, where N equals the filtering interval. Process 1400 can proceed from 1440 to 1480.

[0133] At 1450, process 1400 may involve processor 1110 determining whether the target image is a P-frame. If the target image is a P-frame, process 1400 can proceed from 1450 to 1460. If the target image is not a P-frame, process 1400 can proceed from 1450 to 1470.

[0134] At 1460, process 1400 may involve device 1080 performing MCTF with N / 2 reference images, where N equals the filtering interval. Process 1400 can proceed from 1460 to 1480.

[0135] At 1470, process 1400 may involve processor 1110 copying the target image to filter image buffer 1160 for storage as its filtered image.

[0136] At 1480, process 1400 may involve processor 1110 storing the filtered image into filtered image buffer 1160.

[0137] Figure 15 An example process 1500 according to an embodiment of the present disclosure is illustrated. Process 1500 may represent one aspect of implementing the various proposed designs, concepts, schemes, systems, and methods described above. More specifically, process 1500 may represent one aspect of proposed concepts and schemes related to a precoding process for a video stream according to the present disclosure. Process 1500 may include one or more operations, actions, or functions as shown in one or more of blocks 1510, 1520, 1530, 1540, 1550, and 1560. Although shown as discrete blocks, depending on the desired implementation process, the individual blocks of 1500 may be divided into additional blocks, combined into fewer blocks, or eliminated. Furthermore, the blocks / sub-blocks of process 1500 may be arranged in... Figure 15The process can be executed in the order shown, or in a different order. Furthermore, one or more blocks / sub-blocks of process 1500 can be executed repeatedly or iteratively. Process 1500 can be implemented by or within device 1080 or any variant thereof. For illustrative purposes only and without limitation, process 1500 is described in the context of device 1000 as precoding processing unit 1080. Process 1500 may begin with block 1510.

[0138] At 1510, process 1500 may involve processor 1110 of apparatus 1080 determining the filtering interval per N frames. Processor 1110 may further determine the GOP size of the video. Process 1500 may proceed from 1510 to 1520.

[0139] At 1520, process 1500 may involve motion compensation module 1130 retrieving the target image from target image buffer 1105. Process 1500 can proceed from 1520 to 1530.

[0140] At 1530, process 1500 may involve processor 1110 determining whether the target image has a POC value that is a multiple of the GOP size. If the target image has a POC value that is a multiple of the GOP size, process 1500 can proceed from 1530 to 1540. If the target image has a POC value that is not a multiple of the GOP size, process 1500 can proceed from 1530 to 1550.

[0141] At 1540, process 1500 may involve apparatus 1080 performing MCTF with N / 2 reference images, where N is equal to the filtering interval. In some embodiments, each of the N / 2 reference images has a POC value smaller than that of the target image. Process 1500 can proceed from 1540 to 1560.

[0142] At 1550, process 1500 may involve apparatus 1080 performing MCTF with N reference images, where N is equal to the filtering interval. In some embodiments, half of the N reference images have a POC value less than the POC value of the target image, while the other half of the N reference images have a POC value greater than the POC value of the target image. Process 1500 can proceed from 1550 to 1560.

[0143] At 1560, process 1500 may involve processor 1110 storing the filtered image into filtered image buffer 1160.

[0144] VIII. Example Electronic Systems

[0145] Many of the features and applications described above are implemented as a software process that specifies a set of instructions recorded on a computer-readable storage medium (also known as a computer-readable medium). When these instructions are executed by one or more computing or processing units (e.g., one or more processors, processor cores, or other processing units), they cause the processing unit to perform the actions indicated in the instructions. Examples of computer-readable media include, but are not limited to, CD-ROMs, flash drives, random access memory (RAM) chips, hard disk drives, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc. Computer-readable media do not include carrier waves and electronic signals transmitted via wireless or wired connections.

[0146] In this specification, the term "software" is intended to include firmware residing in read-only memory or an application stored in magnetic storage that can be read into memory for processing by a processor. Furthermore, in some embodiments, multiple software inventions may be implemented as sub-parts of a larger program while retaining distinct software inventions. In some embodiments, multiple software inventions may also be implemented as separate programs. Finally, any combination of separate programs that collectively implement the software inventions described herein is within the scope of this disclosure. In some embodiments, when a software program is installed to run on one or more electronic systems, one or more specific machine implementations are defined to execute and perform the operations of the software program.

[0147] Figure 16 An electronic system 1600 implementing some embodiments of the present disclosure is conceptually illustrated. The electronic system 1600 may be a computer (e.g., a desktop computer, personal computer, tablet computer, etc.), a telephone, a PDA, or any other type of electronic device. Such an electronic system includes various types of computer-readable media and interfaces for various other types of computer-readable media. The electronic system 1600 includes a bus 1605, a processing unit 1610, a graphics processing unit (GPU) 1615, system memory 1620, a network 1625, read-only memory 1630, permanent storage device 1635, input device 1640, and output device 1645.

[0148] Bus 1605 collectively represents all system, peripheral, and chipset buses that communicate with the numerous internal devices of electronic system 1600. For example, bus 1605 communicatively connects processing unit 1610 to GPU 1615, read-only memory 1630, system memory 1620, and permanent storage device 1635.

[0149] From these various memory units, the processing unit 1610 retrieves instructions to be executed and data to be processed in order to perform the processes of this disclosure. In different embodiments, the processing unit may be a single processor or a multi-core processor. Some instructions are passed to the GPU 1615 and executed thereon. The GPU 1615 may offload various computations or supplement image processing provided by the processing unit 1610.

[0150] Read-only memory (ROM) 1630 stores static data and instructions used by processing unit 1610 and other modules of the electronic system. On the other hand, permanent storage device 1635 is a read-write storage device. This device is a non-volatile storage unit that stores instructions and data even when the electronic system 1600 is turned off. Some embodiments of this disclosure use mass storage devices (e.g., magnetic disks or optical disks and their corresponding disk drives) as permanent storage device 1635.

[0151] Other embodiments use removable storage devices (e.g., floppy disks, flash memory devices, etc., and their corresponding disk drives) as permanent storage devices. Like permanent storage device 1635, system memory 1620 is a read-write memory device. However, unlike storage device 1635, system memory 1620 is volatile read-write memory, such as random access memory. System memory 1620 stores some instructions and data used by the processor during runtime. In some embodiments, processes according to this disclosure are stored in system memory 1620, permanent storage device 1635, and / or read-only memory 1630. For example, various memory units include instructions for processing multimedia clips according to some embodiments of this disclosure. From these various memory units, processing unit 1610 retrieves instructions to be executed and data to be processed in order to perform the processes of some embodiments.

[0152] Bus 1605 is also connected to input and output devices 1640 and 1645. Input device 1640 enables a user to communicate information and select commands to the electronic system. Input device 1640 includes an alphanumeric keypad and a pointing device (also known as a "cursor control device"), a camera (e.g., a webcam), a microphone, or similar devices for receiving voice commands. Output device 1645 displays images or other output data generated by the electronic system. Output device 1645 includes printers and display devices, such as cathode ray tube (CRT) or liquid crystal display (LCD), and speakers or similar audio output devices. Some embodiments include devices used as input and output devices, such as touchscreens.

[0153] Finally, as Figure 16As shown, bus 1605 also couples electronic system 1600 to network 1625 via a network adapter (not shown). In this way, the computer can be part of a computer network (e.g., a local area network (“LAN”), a wide area network (“WAN”), or an intranet, or a network of networks. Any or all components of electronic system 1600 can be used in conjunction with this disclosure.

[0154] Some embodiments include electronic components, such as microprocessors, memories, and readable storage media, that store computer program instructions in a machine-readable or computer-readable medium (or, as referred to, a computer-readable storage medium, machine-readable medium, or machine-readable medium). Examples of such computer-readable media include RAM, ROM, read-only optical disc (CD-ROM), recordable optical disc (CD-R), rewritable optical disc (CD-RW), read-only digital versatile optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), various recordable / rewritable DVDs (e.g., DVD-RAM, DVD-RW, DVD+RW, etc.), flash memory (e.g., SD card, mini-SD card, microSD card, etc.), magnetic and / or solid-state hard disk drives, read-only and recordable Blu-ray discs... Optical discs, high-density optical discs, any other optical or magnetic media, and floppy disks. Computer-readable media may store computer programs that can be executed by at least one processing unit and include a set of instructions for performing various operations. Examples of computer programs or computer code include machine code generated by a compiler, and files that include high-level code executed by a computer, electronic components, or a microprocessor using an interpreter.

[0155] While the foregoing discussion primarily concerns microprocessors or multi-core processors that execute software, many of the aforementioned features and applications are implemented by one or more integrated circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, such integrated circuits execute instructions stored on the circuit itself. Furthermore, some embodiments execute software stored in programmable logic devices (PLDs), ROM, or RAM devices.

[0156] As used in this specification and any claim of this application, the terms "computer," "server," "processor," and "memory" refer to electronic or other technical devices. These terms do not include people or groups of people. For the purposes of this specification, the term "display" or "show" means display on an electronic device. As used in this specification and any claim of this application, the terms "computer-readable medium," "computer-readable medium," and "machine-readable medium" are entirely limited to tangible physical objects that store information in a form readable by a computer. These terms do not include any wireless signals, wired download signals, or any other fleeting signals.

[0157] Although this disclosure has been described with reference to many specific details, those skilled in the art will recognize that this disclosure may be implemented in other particular forms without departing from the spirit of this disclosure.

[0158] The subjects described herein sometimes illustrate different components contained within or connected to other different components. It should be understood that the architectures depicted are merely exemplary, and many other architectures can actually be implemented to achieve the same functionality. Conceptually, any arrangement of components achieving the same function is effectively “associated” to achieve the desired function. Therefore, any two components combined in this document to obtain a particular function can be considered “associated” with each other to achieve the desired function, regardless of the architecture or intermediate components. Similarly, any two such associated components can also be considered “operably connected” or “operably coupled” with each other to achieve the desired function, and any two components that can be suchly associated can also be considered “operably coupled” with each other to achieve the desired function. Specific examples of “operably coupled” include, but are not limited to: physically connectable and / or physically interacting components, and / or wirelessly interactable and / or logically interactable components.

[0159] Furthermore, regarding the use of virtually any plural and / or singular terms in the text, those skilled in the art may convert plural to singular and / or singular to plural, provided that it is appropriate for the context and / or application.

[0160] Those skilled in the art will understand that, generally, the terms used herein, particularly those used in the appended claims (e.g., the subjects in the appended claims), are intended as “open-ended” terms (e.g., the term “comprising” should be interpreted as “comprising but not limited to”, the term “having” should be interpreted as “having at least”, the term “comprising” should be interpreted as “comprising but not limited to”, etc.). Those skilled in the art will also understand that if a specific number of the objects described in the appended claims is intended, such intention will be explicitly stated in the claims; in the absence of such a statement, such intention does not exist. For example, to aid understanding, the appended claims may include the use of introductory phrases such as “at least one” and “one or more” to introduce the objects of the claims. However, the use of such phrases should not be interpreted as limiting any claim containing such an indefinite article "a (a) or an" to an invention containing only one such claim, even if the same claim contains the introductory phrases "one or more" or "at least one" and indefinite articles such as "a (a)" or "an" (e.g., "a (a)" and / or "an" should generally be interpreted as meaning "at least one" or "one or more"); the same applies to the use of definite articles to introduce the claim. Furthermore, even if a specific number of the claimed claims is explicitly stated, those skilled in the art will recognize that such a statement should generally be interpreted as meaning at least the stated number (e.g., a statement containing only "two claims" without other modifiers generally means at least two claims, or two or more claims). Furthermore, when using idioms such as "at least one of A, B, and C," such a structure is generally intended to convey the meaning of the idiom as understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" would include, but is not limited to, systems having a single A, a single B, a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). When using idioms such as "at least one of A, B, or C," such a structure is generally intended to convey the meaning of the idiom as understood by a person skilled in the art (e.g., "a system having at least one of A, B, or C" would include, but is not limited to, systems having a single A, a single B, a single C, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). A person skilled in the art will further understand that, whether in the specification, claims, or drawings, virtually arbitrary extractives and / or phrases representing two or more alternative terms should be understood to consider the possibility of including one, any, or all two terms.For example, the phrase “A or B” should be understood as including the possibility of “A”, “B”, or “A and B”.

[0161] Although various methods, apparatuses, and systems have been used to describe and illustrate exemplary techniques herein, those skilled in the art will understand that various other modifications and equivalent substitutions can be made without departing from the claimed subject matter. Furthermore, many modifications can be made to adapt particular situations to the teachings of the claimed subject matter without departing from the central concept described herein. Therefore, it is intended that the claimed subject matter is not limited to the specific examples disclosed, and that such claimed subject matter may also include all implementations and their equivalents falling within the scope of the appended claims.

Claims

1. A video processing method, comprising: Multiple target images are determined based on the filtering interval. These multiple target images include a first subset of multiple source images in the video. The video includes the multiple source images in a time series. Each of the multiple source images has a time identifier that identifies the time position of the corresponding source image in the time series. For each of the plurality of target images, one or more reference images are determined, each of the one or more reference images being a source image of a second subset of the plurality of source images, wherein the second subset of the plurality of source images includes a plurality of source images that are not in the first subset; Multiple filtered images are generated, each of which corresponds to a specific one of the multiple target images, and the multiple filtered images are generated by performing pixel-based filtering based on the one or more reference images corresponding to each target image; The video is encoded into a bitstream using the second subset of the multiple filtered images and the multiple source images; as well as Apply one of the first, second, third, or fourth conditions, or perform the first or second operation. The first condition includes: At least one of the multiple target images includes a first sub-image belonging to the first group and a second sub-image belonging to the second group. Generating a corresponding filtered image for the at least one target image includes performing pixel-based filtering on the first sub-image based on a first number of reference images corresponding to the at least one target image. Generating a corresponding filtered image for the at least one target image further includes performing pixel-based filtering on the second sub-image based on a second number of reference images corresponding to the at least one target image and one or more reference images. The first quantity is different from the second quantity; The second condition includes: At least one of the multiple target images includes a first sub-image of a natural image and a second sub-image of a screen content image, and Generating a corresponding filtered image for the at least one target image includes performing pixel-based filtering on the first sub-image based on the one or more reference images corresponding to the at least one target image, but not on the second sub-image; The third condition includes: At least one of the multiple target images includes a first sub-image of a natural image and a second sub-image of a screen content image. Generating a corresponding filtered image for the at least one target image includes performing pixel-based filtering on the first sub-image based on a first number of reference images corresponding to the at least one target image. Generating a corresponding filtered image for the at least one target image further includes performing pixel-based filtering on the second sub-image based on a second number of reference images corresponding to the at least one target image and one or more reference images. The first quantity is greater than the second quantity; The fourth condition includes that the video does not include scene changes in the time series between each of the plurality of target images and the one or more reference images determined for the corresponding target image; The first operation includes determining the size of the image group in the video, wherein: The multiple target images include multiple first-type target images and multiple second-type target images. Each first-type target image has a corresponding time identifier, which is a multiple of the image group size. Each second-type target image has a corresponding time identifier, which is not a multiple of the image group size. Determining the one or more reference images includes specifying a first number of source images from the second subset for each of the first type of target images. The determination of the one or more reference images also includes specifying a second number of source images from the second subset for each of the second type of target images, and The first quantity is less than the second quantity; The second operation includes determining the size of the image group in the video, wherein: The multiple target images include multiple first-type target images and multiple second-type target images. Each first-type target image has a corresponding time identifier, which is a multiple of the image group size. Each second-type target image has a corresponding time identifier, which is not a multiple of the image group size. For each of the first type of target images, the corresponding one or more reference images include only one or more source images that have a time position earlier than the corresponding first type of target image in the time series.

2. The video processing method as described in claim 1, wherein, Performing pixel-based filtering includes: Each target image is divided into multiple prediction blocks; For each predicted block of the corresponding target image, multiple motion compensation results are determined, each motion compensation result being based on a corresponding one of the one or more reference images corresponding to that target image; and Perform bilateral filtering based on the motion compensation results for each pixel in each of the multiple prediction blocks.

3. The video processing method as described in claim 2, wherein: Each of the multiple motion compensation results includes a best-matching block with the same width and height as the corresponding predicted block, and a loss value representing the difference between the best-matching block and the corresponding predicted block. The execution of bilateral filtering involves calculating a weighted sum for each pixel of the prediction block, based on the corresponding pixel value of the best-matching block among the plurality of motion compensation results and the loss value.

4. The video processing method of claim 2, wherein determining the plurality of motion compensation results includes performing an integer pixel search, a fractional pixel search, or both based on the corresponding prediction block and the one or more reference images.

5. The video processing method as described in claim 1, wherein: The multiple target images include a special target image, wherein scene changes are introduced in the time series immediately preceding the special target image, and The one or more reference images corresponding to the specific target image include only one or more source images that are at a time position later than the time position of the specific target image in the time series.

6. A video processing apparatus, comprising: The processor is configured to receive a video comprising a plurality of source images in a time series, each of the plurality of source images having a time identifier that identifies the time position of the corresponding source image in the time series. The target image buffer is configured to store a plurality of target images determined by the processor based on a filtering interval, the plurality of target images comprising a first subset of the plurality of source images; A reference image buffer is configured to store one or more reference images determined by the processor for each of the plurality of target images, each of the one or more reference images being a source image of a second subset of the plurality of source images, wherein the second subset of the plurality of source images includes the plurality of source images that are not in the first subset; The motion compensation module is configured to generate multiple filtered images by performing pixel-based filtering on the corresponding target image among a plurality of target images based on one or more reference images corresponding to the corresponding target image, each of the plurality of filtered images being generated by the motion compensation module; A video encoder is configured to encode the second subset of the plurality of filtered images and the plurality of source images into a bitstream representing the video; and Apply one of the first, second, third, or fourth conditions, or perform the first or second operation. The first condition includes: At least one of the multiple target images includes a first sub-image belonging to the first group and a second sub-image belonging to the second group. Generating a corresponding filtered image for the at least one target image includes performing pixel-based filtering on the first sub-image based on a first number of reference images corresponding to the at least one target image. Generating a corresponding filtered image for the at least one target image further includes performing pixel-based filtering on the second sub-image based on a second number of reference images corresponding to the at least one target image and one or more reference images. The first quantity is different from the second quantity; The second condition includes: At least one of the multiple target images includes a first sub-image of a natural image and a second sub-image of a screen content image, and Generating a corresponding filtered image for the at least one target image includes performing pixel-based filtering on the first sub-image based on the one or more reference images corresponding to the at least one target image, but not on the second sub-image; The third condition includes: At least one of the multiple target images includes a first sub-image of a natural image and a second sub-image of a screen content image. Generating a corresponding filtered image for the at least one target image includes performing pixel-based filtering on the first sub-image based on a first number of reference images corresponding to the at least one target image. Generating a corresponding filtered image for the at least one target image further includes performing pixel-based filtering on the second sub-image based on a second number of reference images corresponding to the at least one target image and one or more reference images. The first quantity is greater than the second quantity; The fourth condition includes that the video does not include scene changes in the time series between each of the plurality of target images and the one or more reference images determined for the corresponding target image; The first operation includes determining the size of the image group in the video, wherein: The multiple target images include multiple first-type target images and multiple second-type target images. Each first-type target image has a corresponding time identifier, which is a multiple of the image group size. Each second-type target image has a corresponding time identifier, which is not a multiple of the image group size. Determining the one or more reference images includes specifying a first number of source images from the second subset for each of the first type of target images. The determination of the one or more reference images also includes specifying a second number of source images from the second subset for each of the second type of target images, and The first quantity is less than the second quantity; The second operation includes determining the size of the image group in the video, wherein: The multiple target images include multiple first-type target images and multiple second-type target images. Each first-type target image has a corresponding time identifier, which is a multiple of the image group size. Each second-type target image has a corresponding time identifier, which is not a multiple of the image group size. For each of the first type of target images, the corresponding one or more reference images include only one or more source images that have a time position earlier than the corresponding first type of target image in the time series.

Citation Information

Patent Citations

  • Wavelet based coding using motion compensated filtering based on both single and multiple reference frames

    US20030202597A1

  • Non-local bilateral filter

    US20190082176A1

  • Adaptive GOP structure with future reference frame in random access configuration for video coding

    US20190098301A1