Method and system for motion compensated temporal filtering for high efficiency video coding
By adopting an MCTF method that can adapt to reference frame selection based on video content analysis in video encoding and decoding, the problem of low noise removal efficiency in the prior art is solved, and higher image quality and encoding efficiency are achieved.
Patent Information
- Application Number
- CN202411532039.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-10-30
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively remove noise during the video encoding and decoding process, resulting in low encoding quality and efficiency. The conventional motion compensation time filter (MCTF) cannot fully adapt to the image content and introduces large distortion.
Using an MCTF method that can adapt to reference frame selection based on video content analysis, weights are calculated by robust measurement of image data distortion between multiple reference frames and the current frame, highly accurate correlation measurements are provided, and reference frame selection and weight calculation are adjusted according to encoding parameters, scene changes and image content.
It significantly improves image quality and compression rate, enhances encoding and decoding efficiency, and can more accurately distinguish noise and image content, adapt to the encoding needs of different scenarios.
Smart Images

Figure CN120075437A_ABST
Abstract
Description
Background Art
[0001] As the use of video encoding / decoding and streaming becomes more and more common, the demand for high-quality video is also continuously increasing. In the encoding / decoding process in which a video stream is encoded, sent to a remote computing device, and decoded, specific preprocessing operations are performed before encoding to better ensure the quality of the resulting decompressed and displayed video images and improve the encoding / decoding efficiency. This may include performing denoising before the video is encoded or compressed for transmission to another device. Denoising involves removing noise from an image, where the noise is in the form of unwanted variations in the pixel image data, which may blur or obscure the image and cause color changes or brightness errors where the pixels have incorrect image values. This may occur due to poor lighting, malfunctioning or low-quality camera sensors, other camera equipment, and / or other reasons.
[0002] To perform denoising, a motion-compensated temporal filter (MCTF) technique that generates filtered pixel image values can be used. The MCTF technique compares the image data of the original frame with the image data of a motion-compensated reference frame. Then, the motion-compensated blocks of the image data from the reference frame are weighted to form filtered image data. However, due to noise and motion vector errors, such techniques have still proven to be insufficient. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The materials described herein are illustrated in the drawings by way of example and not limitation. For simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Additionally, where considered appropriate, reference numerals have been repeated in the drawings to indicate corresponding or similar elements. In the drawings:
[0004] Figure 1 is a schematic diagram of an example image processing system for performing motion-compensated temporal filtering for video encoding / decoding according to at least one implementation described herein;
[0005] Figure 2 is according to at least one implementation described herein Figure 1 schematic diagram of an example reference frame selection unit of the system;
[0006] Figure 3 is a flowchart of a method for motion-compensated temporal filtering for video encoding / decoding according to at least one implementation described herein;
[0007] Figure 4 is a detailed flowchart of a method for reference frame selection according to at least one implementation described herein;
[0008] Figures 5A - 5C is a detailed flowchart for generating filter weights according to at least one implementation among the implementations of the present disclosure;
[0009] Figure 6 is a schematic diagram of an example system;
[0010] Figure 7 is a schematic diagram of another example system; and
[0011] Figure 8 illustrates another example device all arranged according to at least some implementations of the present disclosure. Detailed implementation
[0012] Now, one or more implementations will be described with reference to the accompanying drawings. Although specific configurations and arrangements are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the relevant art will recognize that other configurations and arrangements can be adopted without departing from the spirit and scope of this specification. It will be apparent to those skilled in the relevant art that the techniques and / or arrangements described herein can also be used in a variety of other systems and applications other than the systems and applications described herein.
[0013] Although the following description elaborates on various implementations that can be embodied in an architecture such as a system-on-chip (SoC) architecture, the implementations of the techniques and / or arrangements described herein are not limited to a specific architecture and / or computing system and can be implemented by any architecture and / or computing system for similar purposes. For example, various architectures and / or various computing devices and / or consumer electronics (CE) devices (such as servers, network devices, set-top boxes, smart phones, tablet computers, mobile devices, computers, etc.) employing, for example, multiple integrated circuit (IC) chips and / or packages can implement the techniques and / or arrangements described herein. Additionally, although the following description may elaborate on many specific details (such as logical implementations, types and interrelationships of system components, logical partitioning / integration selections, etc.), the claimed subject matter can be practiced without such specific details. In other instances, some materials such as control structures and complete software instruction sequences may not be shown in detail so as not to obscure the materials disclosed herein.
[0014] The materials disclosed herein can be implemented in hardware, firmware, software, or any combination thereof. The materials disclosed herein can also be implemented as instructions stored on a machine-readable medium, which can be read and executed by one or more processors. A machine-readable medium can include any medium and / or mechanism for storing or transmitting information in a form readable by a machine (e.g., a computing device). For example, a machine-readable medium can include read only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, electrical, optical, acoustic, or other forms of propagated signals (e.g., carrier waves, infrared signals, digital signals, etc.). In another form, a non-transitory article (e.g., a non-transitory computer-readable medium) can be used in conjunction with any of the above examples or other examples, but does not itself include transitory signals. The non-transitory article does include those elements other than the signal itself that can temporarily store data in a "transitory" manner, such as, for example, RAM, etc.
[0015] References in the specification to "one implementation", "an implementation", "example implementation", etc., indicate that the described implementation may include a particular feature, structure, or characteristic, but each implementation may not necessarily include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same implementation. Further, when a particular feature, structure, or characteristic is described in connection with an implementation, it should be considered within the knowledge of one of ordinary skill in the art to implement such feature, structure, or characteristic in connection with other implementations, whether or not explicitly described herein.
[0016] The following description relates to systems, articles, and methods for motion-compensated temporal filtering for efficient video coding and decoding.
[0017] During video compression, noisy video content results in low encoder efficiency due to the uncorrelated nature of the noise, and specifically because: (1) the noise reduces the temporal correlation between video frames; (2) the noise reduces the spatial correlation of the pixel image data within a single frame; and (3) the noise increases the bit cost of entropy coding and decoding by increasing the number and magnitude of the residuals. This results in a high total bit cost for performing the coding and decoding. Further, the noise also limits the effectiveness of coding and decoding tools for reducing the bit rate distortion and improving the performance. It should be noted that unless otherwise mentioned, the term "correlation" as used herein generally means similarity but is not limited to a specific mathematical equation.
[0018] Conventional MCTF is typically used to attempt to reduce noise to improve coding quality and efficiency. MCTF filtering is performed as a preprocessing step before encoding, and multiple reference frames can be used for motion estimation relative to the frame or block being filtered. Specifically, motion estimation (ME) generally involves searching for a block of image data on a reference frame that matches a block (referred to as the previous block and previous frame) on the frame being filtered. The match is a candidate match represented by a motion vector (MV), and then this MV is used in motion compensation (MC) to select the best match among multiple candidate MVs and corresponding matches and for the previous block of the filtered previous frame. The matches form the resulting motion-compensated (MC) reference frame, and then this reference frame can be used to generate weights to be applied to the MC reference data of the next or current frame being filtered to generate new filtered image data for the block or frame. The filtered frame replaces the current block or current frame to be input to the encoder, and the motion-compensated frame of the current frame can be used as a reference frame for subsequent frames to be filtered.
[0019] Moreover, in conventional MCTF, the motion-compensated frame can be compared with the current frame being filtered to obtain a relatively rough noise estimate as the difference between the two frames. Then the difference or noise estimate is used to determine the weights to be applied to the motion-compensated values, and then these weights are added to the original values of the current frame to generate new filtered image data values. MCTF has been adopted by standard reference codec software, e.g., the Versatile Video Coding (VVC) Test Model (VTM), the Alliance for Open Media (AOM) for the AOMedia Video 1 (AV1) codec, and the High Efficiency Video Coding (HEVC) Test Model (HM), as well as optimized software encoders, e.g., the Scalable Video Technology (SVT) and the Versatile Video Encoder (VVENC). It should be noted that the terms image, frame, and picture may be used interchangeably herein.
[0020] However, when selecting reference frames using conventional MCTF techniques, whenever the same number of future frames and past frames are available, the algorithm only uses the same fixed number of past frames and future frames as reference frames relative to the current frame being filtered. Without some further criteria regarding which reference frames to use, reference frames with relatively large differences (or low correlations) in pixel data relative to the current frame being filtered are more likely to introduce larger distortions during temporal filtering and interfere with the temporal noise estimate. This can include capturing fluctuations from motion and complex content in the image data and misidentifying such changes as noise. Thus, conventional MCTF cannot well adapt to include only the reference frames that are most relevant or most similar to the current frame being filtered.
[0021] Moreover, regarding weight calculation for conventional MCTF techniques, either the known MCTF algorithms do not perform reference-frame based weighting at all, or they only perform simple and fixed reference-frame based weighting based on distortion. This in itself does not take into account accurate noise estimation, nor the errors in the motion vectors (MVs) from the ME. Without more accurate consideration of the distortion between the reference frame and the current frame being filtered, the motion vector (MV) errors, and the noise level, such conventional MCTF techniques cannot adequately adapt to the image content, resulting in insufficient noise reduction and still too low image quality and encoder efficiency.
[0022] To address these problems, the disclosed video coding and decoding methods and systems have preprocessing using denoising that employs MCTF with adaptable reference frame selection based on video content analysis. Additionally or alternatively, the MCTF method used herein calculates weights based on robust multiple measurements (or statistics) of the image data distortion between multiple reference frames and the current frame being filtered, thereby providing a highly accurate correlation (or distortion) measurement between the reference block or frame and the current block or frame.
[0023] More specifically, the disclosed methods and systems for operating MCTF include: selecting reference frames depending on: (1) encoder parameters, such as coding and rendering modes, and thus the group of pictures (GOP) configuration used by the encoder; (2) whether a scene change is near the current frame being filtered; and / or (3) the correlation between each initially available reference frame (or the reference frames still available after applying (1) and (2)) and the current frame being filtered. For the disclosed highly adaptable reference frame selection, the number of reference frames to be used for MCTF can vary from the current frame being filtered to the current frame or even block to block. Moreover, the number of past reference frames and future reference frames can be different, including being zero.
[0024] Thus, the disclosed adaptive reference frame selection is adaptive to video characteristics to better ensure that the selected reference frames have image data with a specific minimum correlation with the current frame (or current block) being filtered, such that MCTF on the one hand more accurately distinguishes noise and on the other hand more accurately distinguishes motion, image complexity, or other intentional image content. Also, the number and position of reference frames within a video sequence are restricted due to scene changes or due to the coding or encoder mode used. This results in a highly adaptable reference frame selection that significantly increases the coding and decoding efficiency including both image quality and compression rate.
[0025] Regarding weights, for MCTF, the disclosed methods and systems determine weights based on the statistics of block distortion, noise level, and coding parameters. Specifically, the weights are determined by aggregately or jointly considering the following: the distortion distribution between the reference frames being used for the current frame, the noise level of the image content, and the coding parameters, e.g., the base or other quantization parameters (QP) of the encoder to be used after preprocessing. In one form, the distortion distribution includes both the dispersion distribution (DD) and the distortion variance of a set of comparisons of the motion-compensated frames or blocks generated using the reference frames with the current frame or block being filtered. The weight calculation also includes a noise factor generated by calculating the estimated noise level of the motion-compensated image or frame relative to the current frame being filtered.
[0026] In the disclosed MCTF solution, reference frame weighting can also be considered highly adaptive because the weights are at least partially based on the dispersion distribution, distortion variance, encoder QP, noise level, and thus on the image content. This adaptability and the in-depth analysis of various measurements of the difference between the motion-compensated frames or blocks and the current frame or block provide very accurate and improved noise reduction and coding efficiency. Since the MCTF unit or module can be an independent module, the MCTF herein can be added before any encoder (e.g., AVC, HEVC, AV1, VP9, VVC, etc.).
[0027] Reference Figure 1 , example image processing (or video codec) system 100 receives an uncompressed video input 101 from a memory or other source and provides image content, which can be the original image content, to an initial preprocessing (PP) unit 102 that adequately formats the image data for denoising and coding. The video input 101 can be in the form of frames of one or more video sequences in any content format (whether natural camera-captured images or synthetic) and with any resolution or color scheme. The initial preprocessing can include demosaicing and / or color scheme conversion (e.g., from RGB to YUV), where each color scheme channel can be processed separately, including denoising and coding. The video sequence frames are not particularly limited by resolution or other video parameters as long as they conform to the codec being used. The frames of the video input can be provided to the initial preprocessing unit 102 in display order.
[0028] Then the initially pre - processed image data or frames are provided to a reference frame generation (RFG) unit 103 having a motion estimation (ME) unit 128 and a motion compensation (MC) unit 130 to generate a new MC reference frame, and motion estimation is performed by using other frames in the initial video sequence 101 (these reference frames have not been motion - compensated). In some systems, motion estimation can be performed outside the system 100 or can be considered as part of the pre - processing unit 102 or MCTF 104 when needed. Although motion estimation can be codec - based, motion estimation is not limited to the number of past reference frames being equal to the number of future reference frames. Specifically, motion estimation and motion compensation can be performed by many different techniques, e.g., known codecs (e.g., VTM or AOM). The block size of the output MV from the ME can be 8x8 (VTM) or 16x16 (AOM), and has 1 / 16 per - pixel accuracy (VTM) or 1 / 8 per - pixel accuracy (AOM). The interpolation filter used in the MC can be 6 - tap (VTM) or 8 - tap (AOM). This can include those techniques that use alternative candidate reference block sizes and patterns for the same pixel region on the current frame being filtered, and the MC selects the best match among the alternative blocks. It should be noted that the block sizes, shapes, and positions for the MC and ME can be completely different from those used for reference frame selection and weight calculation.
[0029] A total of M i available frames can be preset for the ME unit 128 to be used as reference frames, and encoding and decoding are performed in display order. The system 100 or MCTF 104 can have input or configuration parameters to control which frames are to be filtered and thus control which frames are to be used as the current frames for the ME and MC, rather than performing the ME and MC for each frame (although this can be done alternatively). The result is M i available MC reference frames in the past or future relative to the current frame being filtered.
[0030] Then, both the initially preprocessed image data (or frame) and the MC reference frame are provided to a temporal filter unit 104, which can also be considered a preprocessing unit and may or may not be in the same unit as the initial preprocessing unit 102. The temporal filter unit 104 (or just MCTF 104) can be a motion-compensated temporal filter (MCTF) or a denoising unit for performing denoising as described herein. Thus, MCTF 104 filtering is a preprocessing operation added before the encoder 108. In one form, the video input 101 of the source raw video frames is input to the MCTF 104 in display order. The output of the filtered video frames from the MCTF 104 can also be in display order and can be fed into the encoder for compression. The MCTF 104 receives the MC reference frame in display order and performs temporal filtering on the current frame (or block) as an enhanced bilateral filter.
[0031] For example, the preprocessed and denoised image data is then provided to the encoder 108 to output the encoded video 109 for transmission to a remote decoder. The system 100 may also have or communicate with one or more image processing application units 106, which may be involved in setting parameters at the encoder 108 and the temporal filter unit 104. The application unit 106 can provide stored parameters for an application that can use the frames after decoding the frames at a remote device. An application with the application unit 106 can be any display or image analysis application for rendering the frames for display on a television, computer monitor, or mobile device and / or for providing the frames for processing (e.g., for depth analysis, 3D modeling, and / or object recognition (e.g., for VR or AR headsets or other systems)). The application unit 106 can provide desired codec and rendering modes, such as low latency, random access, etc., as described below.
[0032] The temporal filter unit 104 may have a reference frame selection (RFS) unit 110, a block distortion unit 112, a noise estimation unit 114, and a reference frame weighting unit 116, and the reference frame weighting unit 116 has a distortion statistics unit 118 and a weight calculation unit 120. The reference frame weighting unit 116 generates a block weight 122. The temporal filter unit 104 may also have a block attenuation unit 124 and a filtering unit 126, and the filtering unit 126 has a filter application unit 132. The reference frame selection (RFS) unit 110, the block distortion unit 112, the noise unit 114, and the filtering unit 126 may all receive the initially preprocessed image or frame from the initial PP unit 102 and may receive it block by block. It will be understood that any one or more of these units may be on one physical device or at one physical location, and in addition, any one or more of these units may also be on separate remote devices and / or remote locations.
[0033] More specifically, and when referring Figure 2 thereto, the RFS unit 110 receives the image data of the current frame 202 that has been initially preprocessed from the initial PP unit 102, and this image data is received either as an entire frame or block by block when analyzing the blocks. The RFS unit 110 also receives the past MC reference frame 204 and the future MC reference frame 206 generated by the RFG unit 103 by performing motion estimation and motion compensation. The reference frames may be referred to herein as motion-compensated (MC) reference frames or simply as reference frames. Although the unfiltered frames of a video sequence may be used as reference frames for reference frame selection, MC reference frames are used in the examples herein to better ensure high image quality. The RFS 110 also receives a maximum number M i of initially available reference frames to be used with a particular current frame or with a particular current block on the current frame, such that M i may vary from block to block when needed. The maximum reference frame amount M i may be: a previously stored value, or a value obtained from the encoder or an application parameter, and / or a preset adjustable value (e.g., using firmware), etc. The maximum reference frame amount may also be used by the RFG unit 103 as described above to perform ME and MC. Here, the total number of selected frames actually to be used for ME and MC selected by the RFS unit 110 is designated as M, such that M ≤ M i .
[0034] The RFS unit 110 may have a video analysis unit 208 that generates data to be provided to the reference frame decision unit 210. The video analysis unit 208 may have a scene change detection unit (SCD) 212, or may receive and collect scene change data from an external SCD. Such scene change detection is well known and, for example, in one technique, may be based on a change in image data values from frame to frame exceeding a threshold. In one form, the SCD unit 212 may then provide before and after the scene change location relative to the current frame being analyzed and in the video sequence. For example, this may be provided as the frame identification (ID) number or timestamp of the video sequence. This may be omitted when the scene change is not closer to the current frame than Mi / 2 on either side (past or future) of the current frame, or when the scene change is not closer to the current frame than some other maximum range on one side (when the two sides do not have the same number of available reference frames).
[0035] The video analysis unit 208 may also have a correlation unit 214 to determine the correlation (or distortion or similarity) between each available reference frame and the current frame. By way of example, this may be performed as a pixel-by-pixel comparison, as described below with respect to process 400( Figure 4 ). The correlation of each comparison of the past reference frames p1 to pJ with the future reference frames f1 to fK (where J + K = M i ) may then be provided to the reference frame decision unit 210.
[0036] The reference frame decision unit 210 receives encoder or application parameters 216 regarding the encoding and rendering mode. Thus, if the mode is low latency, only past frames are used as reference frames, and if the mode is random access, both future and past reference frames may be used. Also, when the current frame is within J or K frames of a scene change, those available reference frames that form the scene change or those available reference frames of images with a different scene from the current frame are discarded. Otherwise, the reference frame decision unit 210 compares the correlation with a threshold and now considers those available reference frames that meet the threshold to form a set of selected reference frames M for the current frame, where all (or individual reference frames) of the selected reference frames are to be used to determine the weights for filtering the current frame.
[0037] In one form, the reference frame decision unit 210 indicates which of each of the specific available reference frames is the selected reference frame. However, in another form, the reference frame decision unit 210 establishes the farthest or outermost past and future reference frames (regardless of which are considered the selected reference frames) based on the current frame in the video sequence. Once these maximum outer reference frames are established, all frames closer to the current frame than these outer frames are considered the selected MC reference frames, and in one form, it is determined whether their correlation meets a threshold. In one form, the correlation test starts from the outermost reference frame and moves inward, and once a reference frame passes the correlation test, the correlation test stops. It is assumed that any closer frame will also meet the threshold compared to the more outer frame that meets the correlation threshold.
[0038] Accordingly, the reference frame decision unit 210 can provide or transmit a plurality of future reference frames m_f and past reference frames m_p extending from the current frame in video sequence order (or display order), and wherein, m_f + m_p = M, and this is the selection for the entire frame. As another option, when selecting non - consecutive reference frames, the RFS unit 110 can transmit the frame positions in video sequence.
[0039] Referring again to Figure 1 , the noise unit 114 uses the pre - processed image data to generate frequencies (or noise estimates or levels) freq_1 to freq_M by comparing the data of each motion - compensated reference frame being used with the current frame (and in one form, block - by - block) using the noise algorithm quoted below. The noise or frequency values are provided to the reference frame weighting unit 116 and the block attenuation unit 124.
[0040] The block distortion unit 112 receives the past m_p and future m_f (or other) signals or indicators from the RFS unit 110. The reference frame indicators and the current frame and thus the current block can be received in display order, or the block distortion unit 112 can access the current frame stored in, for example, a buffer, and then use the current frame and obtain the reference frames in display order. The block distortion unit 112 then obtains or receives the reference frames, or specifically, receives the MC reference blocks on the selected MC reference frames for the entire current frame or for one or more specific current blocks on the current frame, and wherein the reference frames are selected based on the past reference frame signal m_p and / or the future reference frame signal m_f. In another form, instead of providing the notification from the RFS unit 110 to the filtering unit 126 as the MC reference frame selection, only the selected MC reference frames are placed in a specific MC reference frame buffer, and the distortion unit 112 only retrieves the m_p and / or m_f frames from the buffer. Thus, when memory is involved, the reference frame M can simply be accessed in memory to provide the frame and block image data to the distortion unit 112.
[0041] Once the distortion unit 12 receives a block of the current image and obtains the MC reference blocks in the same block pattern and the same block position, the block distortion E (Equation (6)) is calculated by the algorithm described below and a single value is provided for the pixel block. In one form, the sum of squared differences (SSD) of the current block pixel data and the variance are used. The resulting block distortions dist_1 to dist_M are provided to the reference frame weighting unit 116 and the block attenuation unit 124, and specifically, to the distortion statistics unit 118 and the weight calculation unit 120. When both the past reference frame and the future reference frame are provided, this includes both the past reference frame and the future reference frame.
[0042] The distortion statistics unit 118 calculates distortion statistics based on the distortion E (dist_1 to dist_M) for a single MC block position across multiple selected MC reference frames. This is repeated for each block position on the frame. The distortion statistics unit 118 can then use the distortions dist_1 to dist_M to generate distortion statistics, such as the maximum, minimum, variance, and average of the distortions for each block and for the group of distortions 1 to M. Thus, the variance is the variance of the distortions 1 to M, and the average is the average distortion E of all the reference frames being used. Thus, if there are eight distortions E, the distortion average is the average E between the eight reference frames. The distortion statistics unit 118 can also calculate the distortion dispersion (DD) as the variance of the distortions over the distortion average. Then, the weight calculation unit 120 uses the statistics by considering the DD, the distortion variance, the noise, and the encoder parameters (e.g., the reference of the encoder 108 or other quantization parameters (QP) and several other factors described below). This generates the block weights 122, where in one form, all pixels in the block will have the same weight and will be used in the final weight equation applied by the filtering application unit 132.
[0043] The frequency or noise value from the noise unit 114 can also be used by the block attenuation unit 124 to calculate the attenuation term, which is the Euler's number e fractional exponent denominator term for the final weight equation. The attenuation also considers encoder parameters, such as the QP.
[0044] The filtering unit 126 receives an initial frame or a current frame (or accesses the initial frame or the current frame from a buffer), and may receive them in display order. The filtering unit 126 uses the current frame to calculate a filtered frame to be provided to the encoder 108. Specifically, the filtering unit 126 may cause the filter application unit 132 to obtain attenuation terms and block weights and insert them into a final weight equation, thereby considering block distortion, distortion variance, dispersion distribution, and noise between a reference frame and the current frame. The final weight equation also considers many other constants, such as filter strength, and the per-pixel difference between the reference frame and the current frame. Then, the final weight generated according to the final weight equation is used in a filtering equation for generating filtered pixel image values. Then, the filtered image is provided to the encoder 108 for compression and transmission or storage. Although not shown, the filtering unit 126 may use a reference frame buffer to save MC reference frames as needed.
[0045] The encoder 108 can be any encoder as long as the filtering unit 104 can be arranged to be compatible with the encoder 108. The encoder can use any codec such as VVC, AV1, HEVC, AVC, VP9, etc. and appropriate reference software as described above. The encoder 108 may alternatively or additionally provide codec and rendering modes to achieve low latency and random access, and may incorporate them into the methods disclosed herein. Thereafter, the encoder provides compressed image data for transmission to a remote decoder.
[0046] Reference Figure 3 , an example process 300 for motion-compensated temporal filtering for efficient video coding is arranged according to at least some implementations of the present disclosure. In the illustrated implementation, the process 300 may include one or more operations, functions, or actions shown by one or more of the uniformly numbered operations 302 to 320. By way of non-limiting example, the process 300 may be described herein with reference to the example systems or devices 100, 600, 700, and / or 800 and the discussions herein, respectively. Figure 1 and Figures 6 to 8 The process 300 may include "obtaining image data of frames of a video sequence" 302, and as described above regarding the video input 101. The video frames are not particularly limited to a specific format or parameter (e.g., resolution), and may be sufficiently pre-processed initially as described above for MCTF and encoding. The frames provide the current frame to be filtered and are the original frames for generating motion-compensated (MC) frames that are used as reference frames for subsequent current frames.
[0047]
[0048] Process 300 may include "determining one or more reference frames of the current frame of the video sequence" 304. In particular, operation 304 may include "wherein each of the reference frames has at least one motion-compensated (MC) block of image data" 306, and this may include "comparing the one or more reference frames with previous blocks of the image data of the previous current frame" 308. This operation first refers to the generation of adjacent MC reference frames, which are generated based on previous (in display order) current blocks of the previous current frame, in order to use the MC reference frames for reference frame selection and to generate weights for the current current block of the current current frame to be temporally filtered. Thus, this operation involves motion estimation that can generate motion vectors from the MC blocks to the previous current blocks on the previous current frame, and motion compensation that can use the motion vectors and the indicated MC reference blocks to finally generate new MC blocks for the current current block, as described above.
[0049] In one form, this operation further includes further determining or selecting the reference frame depending on (or considering) the following: (1) the encoding parameters of the encoder for receiving the denoised and filtered image data; (2) the proximity of the scene change to the current frame; and (3) the correlation between the image data on the current frame and the image data on the MC reference frame.
[0050] More specifically, and by way of some examples, operation 304 may include "considering the encoder frame configuration mode being used by the encoder" 310, which refers to the encoding mode associated with the reference frame or GOP dependency structure of the encoder for receiving the denoised and filtered image data. Such a mode may include low latency that does not use future reference frames relative to the current frame being filtered, and random access that uses both past and future reference frames relative to the current frame. In one form, the initial maximum number and / or position of frames available according to the particular codec may be set as the available reference frames for the current frame. The encoder mode may then be used to indicate which of those available reference frames are selected for use with the current frame.
[0051] By way of another example, operation 304 may include "considering the position of the scene change" 312, where it is determined whether the current frame is a scene change frame, or whether the current frame is within an available number of consecutive reference frames from the start or end of the scene. Such scene change detection may be performed by an algorithm that analyzes the luminance and chrominance image data (e.g., comparing the image data with past frames). Those adjacent frames that are not in the same scene as the current frame are not used as reference frames for filtering the current frame.
[0052] By way of yet another example, operation 304 may include "considering the pixel image difference between the current frame and a previously generated reference frame" 314. Here, a correlation or an initial distortion or similarity may be calculated, and it may be the sum of absolute differences (SAD) or other such distortion or correlation equations, and is between the same pixel locations on the current frame and the reference frame. Thus, in one form, those references that may still be used after encoder mode and scene change considerations may each be tested by calculating their correlation with the current frame. For example, the reference frame may be an MC reference frame that has been motion compensated by the reference frame generation (RFG) unit 103. Those MC reference frames that meet the correlation criterion (e.g., less than a threshold) may then be used as the reference frame for the current frame. Although the example has been explained in terms of the entire frame, it should be understood that such correlation calculations may be performed block by block, such that each block in a single frame may alternatively have a different MC reference frame.
[0053] Process 300 may include "generating weights that consider the noise, distortion variance, and dispersion distribution between the MC block and the current block" 316. Once the selected MC reference frame (or reference block) has been determined for the current block to be filtered, a distortion E is calculated between each MC reference block and the current block, such that for each MC reference frame being used, a distortion E is generated for the same MC block location. In one form, and as described above, as an example, the distortion E is calculated by using Equation (6) below, and a single value is provided for the pixel block. In one form, the sum of squared differences (SSD) between the current block and the reference block and the variance of the current block pixel data are used. Then, the individual distortion Es (or dist_1 to dist_M) (each for the same block location, for M reference frames) may be used to calculate distortion statistics, such as the minimum and maximum distortion, distortion variance, and distortion average of the distortion value E. Then, the dispersion distribution (DD) may be calculated as the distortion variance over the distortion average.
[0054] Then, the MCTF may calculate a weight w o (also referred to as an offset weight) based on or depending on an offset adjusted by the block distortion E, which is calculated by considering the DD, distortion variance, and noise associated with the current block.
[0055] Then, predetermined distortion factors and noise factors for the weight equation are selected depending on the values of the block distortion and noise. The selected factors are used to calculate a baseline weight (bw) and a sigma weight (sw). The baseline weight (bw) adjusts the offset weight (w o)To generate a final weight for the weight portion of the final weight equation. sw is used in the decay portion of the final weight equation. The weight equation also takes into account other constants such as filter strength based on the reference frame codec hierarchy, encoder parameters (e.g., the baseline of the encoder being used or other quantization parameters (QP)), and the difference or distortion between the MC block and the current block being analyzed. The final weight is then provided for use in the filtering equation for generating the filtered pixel values.
[0056] Specifically, process 300 may include "generating denoised filtered image data" 318, which may include "applying weights to the image data of the MC block" 320. Here, the filtering equation uses the final weight to modify the pixel values or samples of the MC block, which are added for all MC reference blocks being used for the current block, added to the current block, and then divided by the sum to obtain an average or normalized value, as described below with respect to equation (18). This generates a frame of filtered pixel values, which is then provided to the encoder.
[0057] Refer more specifically to Figure 4 An example process 400 for motion compensated temporal filtering for efficient video coding (and specifically for reference frame selection) is arranged in accordance with at least some implementations of the present disclosure. In the illustrated implementation, process 300 may include one or more operations, functions, or actions as shown by one or more of the uniformly numbered operations 402 to 422. By way of non-limiting example, reference may be made herein to Figure 1 and Figures 6 to 8 the example systems or devices 100, 600, 700, and / or 800 and the discussion herein to describe process 400.
[0058] Process 400 may include "obtaining image data of frames of a video sequence" 402, as mentioned above with respect to process 300. The uncompressed input image data of the input video frames from the video sequence may be obtained from a memory (which has raw image data from one or more camera sensors), from other sources or memories, etc., or may be streaming image data for transcoding and obtained from a decoder. Many variations are possible.
[0059] Process 400 may include "performing sufficient preprocessing for denoising" 404, and this may include performing sufficient initial preprocessing to perform denoising (and generally filtering) and then encoding. This may include Bayer demosaicing of the raw data and other preprocessing techniques. Other encoder preprocessing techniques such as image pixel linearization, shadow compensation, resolution reduction, vignetting removal, image sharpening, etc. may be applied, whether as part of the initial preprocessing before denoising or after denoising and before encoding.
[0060] Process 400 may include "setting the maximum amount of available reference frames" 404, and this refers to setting the maximum initial reference frame amount M i . As described above, the amount M i may be based on the codec standard itself used by the encoder and may have been used by the reference frame generation unit 103. By way of an example, this may include 8 or 16 consecutive frames (also referred to as neighboring frames) in display order before and / or after the current frame to be filtered, for a total of M i = 16 or 32 initial available reference frames.
[0061] Process 400 may include "setting reference frame availability depending on the encoder group of the picture configuration being used" 406. Then, an image processing application at the decoding device (e.g., the display or image data analysis application described above with respect to Figure 1 ) may provide a notification to the temporal filter regarding which encoder (or codec) and rendering mode is being used. An example provides an option for a low-latency mode, where in one form, only past reference frames are used and future reference frames are not used. Otherwise, the random access mode uses both past and future reference frames. The notification may simply provide a single bit (or several bits) for each option, such that a predetermined code or flag is provided for the mode and indicates which mode is being used, or alternatively or more specifically, which reference frames can be used to meet the temporal and access requirements for a particular mode. Thus, by one possible method, the temporal filter may have a list of encoder mode codes, which may be part of the codec standard, corresponding to indicators of which reference frames can be used.
[0062] Thus, process 400 may include asking "Random access type configuration?" 408. If it is not a random access type configuration and the encoding is using the low-latency (LD) mode, then process 400 may include "using only past frames" 410 as follows:
[0063] m p = M, m_f = 0 (1) as explained above with reference to Figure 1 . Here, there are no future frames for the LD mode and there are multiple past frames, where M ∈ M i .
[0064] If, instead, the RA mode is the mode in use, both past and future reference frames can be used. Process 400 then continues and can include "performing scene change detection on the video sequence" 412, regardless of which encoder mode is in use. In this case, the scene change can be detected by the above algorithm, and the scene change can be applied to the current frame by comparing the current frame with the previously analyzed frames. This comparison can be made between the input image data before any denoising modification, although initial preprocessing may have occurred.
[0065] Process 400 can include "limiting the available reference frames depending on the scene change" 414. When the current frame is found to be a scene change frame, the frames before the current frame cannot be used as reference frames because it may have little similarity with the current frame and the subsequent frames in display order after the current frame. In this case, the SCD output sets the position of the current frame in display order as the scene change in the video sequence. Then, if the encoding mode is the RA mode, the reference frame selection (RFS) unit 110 can dynamically derive the amount and / or positions m_p and m_f of the past and future reference frames in the video scene. For example, if the current frame is the start of a new scene, then
[0066] m_p = 0, m_f = M (2)
[0067] If, instead, the current frame is the end of the scene, then
[0068] m p = M, m_f = 0 (3)
[0069] If the current frame is in the middle of the scene, both past and future frames can be used as reference frames for the current frame. As a result, the reference frames m_p and m_f selected at this time can be set to:
[0070] m_p + m_f = M (4)
[0071] The process 400 may include "obtaining the selected MC reference frame positions for initial m_p and / or m_f depending on the scene change and the encoder GOP configuration" 416, where the frame positions of the reference frames M that are still available are each obtained to determine the correlation of each frame in the frame with the current frame. First, it should be noted that the reference frames to be used for filtering are the frames that have previously undergone motion compensation (MC) by the reference frame generation unit 103. Also, as described above, although the correlation is discussed at the frame level, this can be performed on a per-block basis, e.g., as desired, 8x8, 16x16, or other sizes, shapes, and block patterns, such that different blocks can have different MC reference frames. Also, at this time, as described above, the reference frame selection unit can obtain both the MC reference frames, e.g., from the buffer loaded by the reference frame generation unit, and the input (or initial or original) frames in display order from the input buffer.
[0072] The process 400 may include "comparing each selected MC reference frame m with the current frame to be filtered to determine the correlation" 418. Here, the selected refers to those frames that are still to be selected as candidates. The correlation can be the SAD equation, which finds the difference between the pixels at the same pixel positions within the adjacent MC reference frame and the current frame. The result is a single SAD value for the frame comparison or correlation (or for each correlation of the blocks in the frame if performed at the block level). It should be noted that the comparison (or distortion or correlation) between the MC reference frame and the current frame herein is performed between blocks having the same pixel positions on the frame (or other set positions when needed), regardless of whether the comparison is for the entire frame or at the block level. ME (with motion vector block matching search) and MC are only applied to the current frame to generate the MC reference frames for subsequent current frames by the RFG unit 103. Thus, with the methods and systems disclosed herein, a block matching search between the reference frame and the current frame is not necessary to perform reference frame selection for the current frame and weight calculation for the current frame for temporal filtering according to the methods disclosed herein.
[0073] The process 400 may include "comparing each correlation with a criterion" 420, where each correlation is compared with a threshold determined experimentally, for example. In one form, the threshold is the maximum correlation that can be used to obtain sufficiently accurate denoising to generate a good quality image. In one form, as an example, the correlation threshold can be an SAD that is less than or equal to approximately 6% of the maximum pixel value in the block.
[0074] The process 400 may include "maintaining each MC reference frame of the current frame that meets the criterion" 422, where the set M of the selected MC reference frames is maintained in a memory or buffer for filter calculation, as in the following process 500.
[0075] Reference Figures 5A - 5C , an example process 500 for motion - compensated temporal filtering for efficient video coding and particularly for computing MCTF weights is arranged according to at least some implementations of the present disclosure. In the illustrated implementation, process 500 may include one or more operations, functions, or actions shown by one or more of the uniformly numbered operations 502 to 560. By way of non - limiting example, process 500 may be described herein with reference to the example systems or devices 100, 600, 700, and / or 800 of Figure 1 and Figures 6 to 8 and the discussion herein.
[0076] Process 500 may include "obtaining a current frame of a video sequence" 502, and as already described with processes 300 and 400. At this time, for example, the current frame is obtained or received by the block distortion unit 112 in display order, as explained above. This operation also includes any initial pre - processing as described above.
[0077] Process 500 may include "performing motion estimation and motion compensation to match an MC reference block with a current block" 503. This refers to the generation of the MC reference frame as described above and to be used for filtering the current frame. Thus, as described above, the previous current frame to be filtered is used to generate the MC reference frame by applying ME to the previous current frame.
[0078] ME involves performing a search and forming a motion vector (MV) for matching a previous current block with a previous reference block. The MV is provided to the MC unit. As mentioned, motion estimation may use one or alternative multiple block patterns for the best block pattern (shape, size, and / or position on the frame) to be selected by the MC operation, and these blocks may be set without any consideration related to the blocks used for computing weights for filtering. The MC operation may compare each candidate MC reference block with the previous current block and select the MC reference block having the minimum difference from the previous current block, for example, by SAD or other difference calculations and other considerations as needed. The best reference block is selected for the previous current block to form the MC reference block for the new MC reference frame.
[0079] Process 500 may include "obtaining the selected MC reference frames for the current frame" 504, and this refers to obtaining the selected MC reference frames for calculating block distortion, noise, weight calculation, and filtering. In one form, the number of past MC reference frame m_p and future MC reference frame m_f is obtained from the reference frame selection (RFS) unit according to the above process 400, and the image data of frames m_p and m_f is obtained from the RFG unit that performs motion compensation. More precisely, the selected MC reference frames are accessed in a memory or buffer. In one form, either all available reference frames M i are provided to the block distortion unit and the noise unit, and these units only use the selected MC reference frames M, or only the selected MC reference frames M are provided to these units. The MC reference frames may be provided to these units in display order.
[0080] This may also include obtaining the MC reference frames on a per-block basis, where the reference frames and the current frame are divided into the same block pattern (same block size and location), e.g., 8x8 or 16x16 blocks. In one form, the blocks may or may not overlap as needed. In another form, the blocks have a size that fits uniformly within the frame dimension such that no padding is required. In one form, the blocks may be provided in raster order.
[0081] Process 500 may include "generating a noise estimate for each frame or each block of the current frame" 506. The reference frames are also obtained by the above noise unit. The noise is calculated as a frequency. For example block frequency (or noise) calculation:
[0082]
[0083] Among them, HSD and VSD are respectively the sum of the horizontal squared differences and the sum of the vertical squared differences of each pair of two adjacent pixels in the horizontal difference (or simply referred to as diff) block and the vertical difference block. Here, the diff block refers to the pixel grid or surface of the per-pixel subtraction result between the current block and the motion-compensated reference block. Thus, for example, the operations are as follows: (1) Determine the pixel differences between the current block and the pixels in the MC reference block to form a diff block with pixel difference values having pixel image data values at each pixel or element position of the diff block; (2) Then horizontally and vertically subtract those pixel differences at adjacent pixel positions on the diff block; (3) Then square the horizontal difference or vertical difference; and (4) Compare and sum the squared values for each block such that each block comparison has a single HSD or VSD value. The SSD is also determined separately for the entire block. This operation is repeated for each pair of corresponding block pairs on the current frame and the MC reference frame, and then repeated for each MC reference frame being used. The F metric measures the block frequency, where a small F value indicates a high noise level in the block (indicating more noise), and a large F value indicates less noise in the block. Thus, the frequency F in this article can also be considered as noise, noise level, and / or noise estimation.
[0084] Procedure 500 may include "computing the distortion between the MC reference block and the corresponding current block" 508. Here, the block difference (or block distortion) is determined for each selected MC reference block and the current block being filtered. This can be calculated as:
[0085]
[0086] where SSD is the sum of the squared differences between the pixel values at the same pixel positions of the block in the current frame and the block in the motion-compensated reference frame. The variance V of the pixel data in the current block can be used to normalize E. For the current block to be filtered, the result is M E values, each of which comes from each reference frame of the current frame being used and is determined by the reference frame selection (procedure 400( Figure 4 ))). These values can be, for example, the dist_1 to dist_M values provided by the block distortion unit 112.
[0087] Then, the distortion statistics are generated using the distortion E for each reference block at the same block position as the current block. Thus, procedure 500 may include "generating distortion variance (DistVar) and average (DistAvg) for each current block and the corresponding selected reference block" 509. Then, by determining the maximum distortion and the minimum distortion between the block distortions E for a single block position, and then calculating the variance and the average value of a group of block distortions E at a single block position, the distortion statistics for each reference block position are generated, and then these operations are repeated for each block position on the frame.
[0088] The process 500 may include "calculating the dispersion distribution (DD)" 510. To measure how the distortion is distributed, the DD metric is calculated in Equation (7) below.
[0089] Dispersion distribution = (distVar + 1) / (distAvg + 1) (7)
[0090] DD is superior to considering the distortion variance or average separately to normalize the variance. Specifically, since the variance describes the difference from the average, a larger average allows for a larger difference. Thus, the metric DD can be used to normalize the variance according to the average.
[0091] The process 500 may further include "obtaining a reference QP" 512, and obtaining it from the encoder being used. The quantization parameter (QP) used during encoding can be used to set the filter strength for denoising. The higher the QP (to obtain a lower bitrate), the lower the quality of the expected image. In this case, here, a stronger filter weight (with a higher value) can be set to remove a larger amount of noise from the image. Thus, if the actual QP is not used in real time, at least the reference QP can be used. For example, based on the weight w of the block distortion E described herein o modified by the offset applied to E. The QP can be used to determine the offset. Moreover, the QP can be directly used for the attenuation term in the final weight equation. Both of these are described in more detail below.
[0092] Reference Figure 5B , the weight for each reference frame can be calculated on a block basis and based on the block-level distortion value of the current reference block (or in other words, the block distortion E) as well as the statistics of pixel distortion, noise level, and coding parameters.
[0093] Specifically, an offset weight w o (or the weight based on the offset applied to the block distortion) is generated to subsequently determine the final block weight of the MC reference block compared to the current block. The algorithm for the offset weight can be:
[0094]
[0095] where E is the block distortion of the reference block of the reference frame, or in other words, the measurement of the block error after motion compensation, as described above in Example Equation (6). The term min(E) is the minimum block distortion among all reference frames for the same block position. The offset is a value calculated using an algorithm that adapts to the noise level, coding QP, and distortion statistics, which includes the distortion variance as well as the dispersion distribution (DD), as Figure 5B shown and described below. Thus, for at least these reasons, the offset weight w oConsider these statistics to modify E, and thus in turn consider the reference weight bw and the final weight W described below bw , and in turn W r (i,a).
[0096] In one form, the offset can be based on the block-level statistics described above. The algorithm for deriving the "offset" for reference frame weighting can be determined by using the following operations.
[0097] Procedure 500 may include "setting offset O = 1" 514 as an initialization.
[0098] Procedure 500 may include asking "noise level > nsHighThld?" 516, where the noise level calculated as F (Equation 5) for the current block is compared with the experimentally determined high or maximum noise threshold nsHighThld. In one form, for example, nsHighThld is 25.
[0099] If no, and the noise is below the nsHighThld threshold, the block may be very clean with very low noise, such that it may not be necessary to increase the offset by more than one. In this case, the procedure jumps to operation 528.
[0100] If yes, and the noise is above the nsHighThld threshold, such that the block is relatively noisy, it may be desirable to increase the offset. In this case, the procedure proceeds to operation 518 to check the dispersion distribution.
[0101] Procedure 500 may include asking "DD > DDThld?" 518 to check the dispersion distribution. In one form, for example, DDThld is set to 0.5. The DDThld threshold can also be determined experimentally.
[0102] If no, when the DD metric is less than the DDThld threshold, the distortion value in the block can be convergent. This can occur when the motion vectors of multiple (or all) reference frames are accurate enough and the video content is clean enough with very little noise (even if the block noise is high enough above the noise threshold nsHighThld). In this case, the offset can be kept at the minimum value of 1 for all reference blocks, such that blocks with a block distortion E greater than min(E) will have the smallest possible weight, and thus the reference frame or block will have a smaller weight in the temporal filtering. In this case, here too, the procedure jumps to operation 528.
[0103] If so, when the DD metric value is greater than the threshold DDthld, the spread of the distortion value from the mean is wider. In this case, and for content with a high noise level, in order to efficiently reduce noise and improve encoder quality, an offset much greater than 1 is desired, such that the reference block with a block distortion E greater than min(E) (Equation (7) above) will have a greater weight. Process 500 then moves to operations 520 and 522 to check the variance before adjusting the offset.
[0104] First, the low block distortion variance is checked at operation 520. Thus, process 500 may include asking "distvar > varLowThld?" 520. By way of one form, for one example, varLowThld is set to 2. The varLowThold threshold can also be determined experimentally.
[0105] If not, when the distortion variance is very small (which means when the distortion variance has a very small difference from the average distortion), the offset can be kept at the minimum value (e.g., one) to better ensure avoiding errors during the weighted average of the temporal filtering. In this case, the process proceeds to operation 528.
[0106] If so, when the distortion variance distVar is not lower than the low threshold, the high threshold is checked. Thus, process 500 may include asking "distvar > varHighThold?" 522. If so, and distvar is greater than the varHighThold threshold, the variance is considered very high, and the process proceeds to operation 524 to apply an offset with a large upward increment, and it can be the maximum increment. If not, the process proceeds to operation 526 to provide an offset with a relatively small upward increment.
[0107] Specifically, process 500 may include "set O = O + A" 524 and "set O = O + B" 526, where the offset increment A can be set for the maximum offset increment increase (which can be determined experimentally), and the offset increment A can be set relative to other increments B, C, and D in operations 526, 536, and 538, respectively. Increment A may or may not be greater than increment B to adjust the block distortion between the MC reference block and the current block, and thus adjust the resulting weight w o In this case, the above statistics show high noise, dispersion distribution, and variance. When the noise and DD are high, increment B can provide an offset increment, but the variance is not extremely high. As an example, A can be set to 10 and B can be set to 20. Determining any of the offset increments A, B, C, and / or D may involve experimentation.
[0108] Regardless of whether the offset has been adjusted, the process proceeds to operation 528 to consider encoder parameters, and then rechecks the DD and variance thresholds to obtain a smaller offset increment increase than that provided by increments A and B, to provide a very precise offset value, but which can be of the same or different orders of magnitude.
[0109] Process 500 can include querying "QP > HighQPThld?" 528, where encoder settings or parameters are considered, and specifically the quantization parameter (QP) that sets the encoder quantization precision. A larger QP refers to greater quantization and compression (to reduce the actual bitrate), but also reduces the image quality. For content with any noise level, when the encoder QP is considered large, stronger temporal filtering should be used to provide as much coding gain as possible. Therefore, a larger offset should be generated to increase the magnitude of the weights. An example value for HighQPThold can be 32.
[0110] If the query answer is no, and QP is less than the threshold HighQPThld, then no further offset adjustment may be required. In this case, the process continues to apply the offset (operation 540). Thus, if the offset is not adjusted, the offset remains as one at initialization. If operation 524 or 526 made a previous adjustment to add a relatively large offset due to high noise, large distortion variance, and / or large DD, then the large offset is maintained and no smaller high-precision offset increment is applied for QP.
[0111] If yes, and QP is greater than the threshold HighQPThld, then as initially mentioned, regardless of the noise or distortion statistics, it is determined whether further offset increments C or D are needed (operations 536 or 538). Thus, process 500 can first include querying "DD > DDThld2?" 530. As an example, DDThld2 can be the same as or smaller than DDThld, and here can be set to 0.5, and determined experimentally, similar to DDThld above. Otherwise, the operation proceeds similarly to the DD check at operation 518, where no further offset increment is made when DD is less than the threshold DDThld2, but if DD is greater than the threshold DDThld2, then the process proceeds to check the distortion variance.
[0112] Process 500 may include querying "distar > varLowThld?" 532 and "distvar > varHighThold?" 534, where if the distortion change is less than the low threshold distVarThld2, no further offset increment is required and the process proceeds to operation 540. If distVar is greater than the low threshold varLowThld2, the high threshold is checked. Then, if distVar is greater than the high threshold varHighTheld2, an offset increment is applied at operation 536, where process 500 may include "set O = O + C" 536. Otherwise, when sidtVar is less than the high threshold varHighThld2, an offset increment is applied at operation 538, where process 500 may include "set O = O + D" 538. By way of an example, the offset increment C may be 10 and the offset increment D may be 20, which are the same as A and B. In other cases, the offset increments A to D may range from the maximum offset increment to the minimum offset increment, although many other configurations or arrangements may alternatively be used, depending on test parameters, etc.
[0113] To calculate the weights, process 500 may include "calculating the weights (w o )" 540 based on the distortion offsets for each MC block, and this refers to equation (8) repeated here:
[0114]
[0115] The weights w o are referred to as weights considering the offsets or simply offset weights for short, to distinguish this weight from other weights mentioned herein. This includes operation 540 which may include "calculating the block distortion" 542. For an example calculation of the block distortion E, see equation (6) above.
[0116] Process 500 may include "applying the offset" 544, where the offset from operations 514 - 538 is applied in equation (8) to adjust E. By way of an example herein, the offset may be equal to or greater than 1, where 1 indicates less distortion and noise, resulting in a smaller offset weight w o . A larger offset indicates greater noise and / or distortion, where the weight w o will be larger to have a greater gain and remove more noise. In other words, a smaller offset provides a relatively smaller weight to the reference block with distortion = E, while a larger offset provides a relatively larger weight to the reference block. Here, a small weight or a large weight is within the reference block of the current block (or each is associated with the reference block).
[0117] The process 500 may include "generating weights for MC blocks" 546, and this operation 546 may include "calculating weights for equations" 548. Specifically, the final weight equation (9) listed below has a weight part and a decay part. The weight part has an adjusted weight W modified by a plurality of constants s bw , and the decay part is the denominator of a fractional exponent of the Euler number e. As shown in equation 10, the adjusted weight W bw is an offset weight w modified by a reference weight (bw) o (the above equation (8)). The decay term or part takes into account the QP and serves as a sigma weight (sw). The variables bw and sw are generated by using the block distortion E and the block frequency (noise) F to find the predetermined weight factors for establishing bw and sw in the following equations 12 - 17. Details are as follows.
[0118]
[0119] Where:
[0120] W bw = w o x bw (10)
[0121] Where the term ΔI(i) 2 is obtained by:
[0122] ΔI(i) = (I r (i) – I o ) * (1024 / 2 b ) (11)
[0123] Where the variable i is the frame distance from the MC reference block to the current block's frame, the variable a is the number of selected reference frames being used, the variable b is the bit depth being used, I o is the original or current frame pixel value, I r () is the motion - compensated pixel (or sample) value from the MC reference block. Also, the constant s l is the filter strength for the luminance or luma channel, which alternatively may be s c for the chrominance channel, and in one example, s l is 0.4, and s c is 0.55, so the constant s o can take into account the filter strength, depending on whether RA, LD, or another encoder and rendering mode is being used and the hierarchical level of the current frame in the encoder group of the picture, and s r() is the filter strength adjusted for the number of selected MC reference frames being used for the current block. These constants can be determined depending on the layer and experiment in which the current frame to be filtered is located.
[0124] Now calculate bw and sw. Both bw and sw can be initialized to 1.0, and then:
[0125] bw = bw * m_bwDistFactor * m_bwNoiseFactor (12)
[0126] sw = sw * m_swDistFactor * m_swNoiseFactor (13)
[0127] Operation 548 may include "considering distortion block weight" 550, where m_bwDistFactor is a predetermined fixed distortion weight factor value based on the comparison of block distortion E with a distortion threshold. For example:
[0128]
[0129] And where m_swDistFactor is a predetermined fixed distortion weight factor value based on the comparison of E with other distortion thresholds. For example:
[0130]
[0131] Operation 548 may also include "considering noise" 552, where m_bwNoiseFactor is a predetermined fixed noise weight factor value based on the comparison of F with other distortion thresholds. For example:
[0132]
[0133] And where m_swNoiseFactor is a predetermined fixed noise weight factor value based on the comparison of F with other distortion thresholds. For example:
[0134]
[0135] Once the factors in equations (15) to (17) are established, the reference weight bw and the sigma weight sw can be generated, and sw is ready to be input into the final weight equation (9), and bw is ready to be input into equation (10). As described above, this can be achieved by calculating sw and bw in equations (13) and (14).
[0136] Operation 546 may include "calculating the attenuation of the equation" 554, and it may include "obtaining sw" 556 and "considering QP" 558 as described above. Thus, by an example equation:
[0137] σ(QP) = 3 (QP - 10) (18)
[0138] Among them, in addition to the block distortion offset, QP is also considered here. Equation (18) is used to adjust QP so as to weaken filtering when QP is low and strengthen filtering when WP is high.
[0139] Process 500 can then include "generating filtered image data" 560. This is an example equation for the temporal filtering used herein:
[0140]
[0141] where, I n is the filtered pixel value, mp and mf are m_p and m_f respectively, and other variables are as described above.
[0142] The obtained filtered value can then be placed on the filtered frame version of the current frame, and then the filtered frame is provided to the encoder. In one form, although the preprocessings mentioned above can be used, other preprocessings (after MCTF filtering) may not be required, and the denoised and filtered frame can be directly provided to the encoder without further denoising or other modification of image data values related to image quality.
[0143] The obtained filtered value can then be placed on the filtered frame version of the current frame, and then the filtered frame is provided to the encoder. In one form, although the preprocessings mentioned above can be used, other preprocessings (after MCTF filtering) may not be required, and when desired, the denoised and filtered frame can be directly provided to the encoder without further denoising or other modification of image data values related to image quality.
[0144] Experimental results
[0145] To demonstrate the encoder efficiency, experiments were performed to measure the encoder quality gain using the disclosed MCTF method and system. The MCTF as described above operates with a VVC encoder having a GOP16 format (which has a constant quantization parameter (CQP)), and simultaneously uses random access B. The test dataset has 83 clips with various types of content, including natural content (classes A to E), video games (class V), and screen content (classes F, G), which refers to text or work production screens (word processors, spreadsheets, slide presentations, web browsers, etc.). Various resolutions were also tested, including 4K (class A), 1080p (class B), 720p (class E), wide video graphics array (WVGA) (class C), and wide quarter VGA (WQVGA) (class D). Content with varying noise levels was tested, including high noise level content (classes H, I) and clean content (classes F, G, V). Other variations of the content include high motion, video conferencing, rich texture content, etc.
[0146] Two tests were run, including VVC encoding with the above test configuration and without MCTF for comparison or control (or anchor), and VVC encoding with the same test configuration but with the disclosed MCTF method and system and as a preprocessing operation before the encoder. M = 8 reference frames were used.
[0147] For the control or anchor test without MCTF and the test with the disclosed method and system, the peak signal-to-noise ratio (PSNR) was calculated based on the Bjontegaard delta rate (BD rate). Table 1 below shows the resulting BD rate gains, where good quality gains were achieved using the disclosed MCTF method and system. Among all content types, natural content had the highest gain, i.e., up to a 10+% BD rate gain for class B. Class V game content also showed a 3.0+% gain. The quality gains for clean screen content (classes F and G) and very small resolution WQVGA content (class D) were relatively low, which is expected.
[0148] Table 1
[0149]
[0150]
[0151] Although implementations of the example processes 300, 400, and 500 discussed herein may include performing all operations shown in the order shown, the present disclosure is not limited in this regard, and in various examples, implementations of the example processes herein may include only a subset of the shown operations, operations performed in a different order than shown, or additional or fewer operations.
[0152] Additionally, any one or more of the operations discussed herein can be performed in response to instructions provided by one or more computer program products. Such a program product can include a signal-bearing medium providing the instructions that, when executed by, for example, a processor, can provide the functionality described herein. The computer program product can be provided in any form of one or more machine-readable media. Thus, for example, a processor including one or more graphics processing units or processor cores can perform one or more of the blocks of the example processes herein in response to program code and / or instructions or instruction sets transmitted to the processor by one or more machine-readable media. Generally, the machine-readable media can transmit software in the form of program code and / or instructions or instruction sets that can cause any one of the devices and / or systems described herein to implement at least part of the operations discussed herein and / or any part of the devices, systems, or any module or component discussed herein.
[0153] As used in any implementation described herein, the term "module" refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. Software can be embodied as a software package, code, and / or instruction set or instructions, and "hardware" as used in any implementation described herein can include (either alone or in any combination) for example hardwired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, execution unit circuitry, and / or firmware storing instructions executed by the programmable circuitry. Modules can collectively or individually be embodied as circuitry forming part of a larger system (e.g., an integrated circuit (IC), a system on a chip (SoC), etc.).
[0154] As used in any implementation described herein, the term "logic unit" refers to any combination of firmware logic and / or hardware logic configured to provide the functionality described herein. As used in any implementation described herein, "hardware" can include (either alone or in any combination) for example hardwired circuitry, programmable circuitry, state machine circuitry, and / or firmware storing instructions executed by the programmable circuitry. Logic units can collectively or individually be embodied as circuitry forming part of a larger system (e.g., an integrated circuit (IC), a system on a chip (SoC), etc.). For example, a logic unit can be embodied in the logic circuitry of the implementation firmware or hardware for the codec system discussed herein. Those skilled in the art will recognize that operations performed by hardware and / or firmware can alternatively be implemented via software, which can be embodied as a software package, code, and / or instruction set or instructions, and will also recognize that a logic unit can also utilize a portion of software to implement its functionality.
[0155] As used in any implementation described herein, the term "component" may refer to a module or a logical unit, as these terms are described above. Thus, the term "component" may refer to any combination of software logic, firmware logic, and / or hardware logic configured to provide the functionality described herein. For example, those skilled in the art will recognize that operations performed by hardware and / or firmware may alternatively be implemented via software modules, which may be embodied as software encapsulation, code, and / or instruction sets, and will also recognize that a logical unit may also utilize a portion of software to implement its functionality.
[0156] The term "circuit" or "circuitry" as used in any implementation of this document may include or form (alone or in any combination) for example, hardwired circuitry, programmable circuitry such as a computer processor including one or more separate instruction processing cores, state machine circuitry, and / or firmware storing instructions executed by the programmable circuitry. The circuitry may include a processor ("processor circuitry") and / or a controller configured to execute one or more instructions to perform one or more operations described herein. Instructions may be embodied as, for example, an application, software, firmware, etc. configured to cause the circuitry to perform any of the aforementioned operations. Software may be embodied as software encapsulation, code, instructions, instruction sets, and / or data recorded on a computer-readable storage device. Software may be embodied or implemented to include any number of processes, and processes may in turn be embodied or implemented in a hierarchical manner to include any number of threads, etc. Firmware may be embodied as code, instructions, or instruction sets and / or data hardcoded (e.g., non-volatile) in a memory device. The circuits may be embodied collectively or individually as circuitry forming part of a larger system (e.g., an integrated circuit (IC), an application specific integrated circuit (ASIC), a system on a chip (SoC), a desktop computer, a laptop computer, a tablet computer, a server, a smart phone, etc.). Other implementations may be implemented as software executed by a programmable control device. In this case, the term "circuit" or "circuitry" is intended to include a combination of software and hardware, such as a programmable control device or a processor capable of executing software.
[0157] refer to Figure 6, an example video codec system 600 for providing MCTF denoising for video coding and decoding can be arranged according to at least some implementations of the present disclosure. In the illustrated implementation, the system 600 can include an imaging device 601 (e.g., one or more cameras), one or more central and / or graphics processing units or processors 603, a display device 605, one or more memory stores 607, an antenna 650 for wireless transmission, and a processing unit 602 for performing the above operations. The processor 603, the memory store 607, and / or the display device 605 can be capable of communicating with each other via, for example, a bus, a cable, or other access. In various implementations, the display device 605 can be integrated in the system 600 or implemented separately or remotely from the system 600.
[0158] As Figure 6 shown, the processing unit 602 can have a logic circuit system 604, which includes a preprocessing (PP) unit 606 and a separate video encoder unit 608 or has a video decoder unit 610. The preprocessing unit 606 can receive image data for encoding and can have an initial preprocessing (PP) unit 102, a reference frame unit 103 having an ME unit 128 and an MC unit 130, and a temporal filter unit 104 in the system 100 ( Figure 1 ). The temporal filter unit 104 can perform denoising and can have a reference frame selection unit 110, a block distortion unit 112, a noise unit 114, and a reference frame weighting unit 116. The reference frame weighting unit 116 can have a distortion statistics unit 118 and a weight calculation unit 120. The temporal filter 104 can generate a block weight 122 and can have a block attenuation unit 124. The temporal filter can also have a filtering unit 126 including a filter application unit 132. Other preprocessing units can also be provided. All of these units, logics, and / or modules perform at least the tasks described above, as implied by the names of the units, but can also perform additional tasks.
[0159] As will be recognized, Figure 6The modules shown in [Figure] can include various software and / or hardware modules and / or components that can be implemented via software, firmware, hardware, or a combination thereof. For example, a module can be implemented as software via a processing unit 602, or a module can be implemented via a dedicated hardware portion. Additionally, the memory storage 607 shown can be a shared memory for the processing unit 602, e.g., storing or buffering any preprocessed and denoised data, whether stored on any of the aforementioned optional buffers or any of the memories mentioned herein. Further, the system 600 can be implemented in various ways. For example, the system 600 (excluding the display device 605) can be implemented as a single chip or device having a graphics processing unit (GPU), an image signal processor (ISP), a quad-core central processing unit, and / or a memory controller input / output (I / O) module. In other examples, the system 600 (again excluding the display device 605) can be implemented as a chipset or a system-on-a-chip (SoC).
[0160] The processor 603 (or the processor circuitry that forms the processor) can include any suitable implementation, including, for example, a microprocessor, a multi-core processor, an application-specific integrated circuit, a chip, a chipset, a programmable logic device, a graphics card, integrated graphics, a general-purpose graphics processing unit, etc. Additionally, the memory storage 607 can be any type of memory, e.g., volatile memory (e.g., static random access memory (SRAM), dynamic random access memory (DRAM), etc.) or non-volatile memory (e.g., flash memory, etc.). In a non-limiting example, the memory storage 607 can also be implemented via a cache memory.
[0161] Reference Figure 7 , according to the present disclosure and examples, the system 700 can be a media system, although the system 700 is not limited to this context. For example, the system 700 can be incorporated into a personal computer (PC), a laptop computer, an ultra-laptop computer, a tablet computer, a touchpad, a portable computer, a handheld computer, a palmtop computer, a personal digital assistant (PDA), a cellular phone, a combination cellular phone / PDA, a television, a smart device (e.g., a smart phone, a smart tablet computer, or a smart television), a mobile Internet device (MID), a messaging device, a data communication device, etc.
[0162] In various implementations, the system 700 includes a platform 702 communicatively coupled to a display 720. The platform 702 can receive content from a content device such as a content service device 730 or a content delivery device 740 or other similar content sources. A navigation controller 750 including one or more navigation features can be used to interact with, for example, the platform 702 and / or the display 720. Each of these components is described in more detail below.
[0163] In various implementations, platform 702 may include chipset 705, antenna 710, memory 712, storage device 711, graphics subsystem 715, applications 716, and / or radio device 718, and any combination of antenna 710. Chipset 705 may provide intercommunication between processor 714, memory 712, storage device 711, graphics subsystem 715, applications 716, and / or radio device 718. For example, chipset 705 may include a storage device adapter (not shown) capable of providing intercommunication with storage device 711.
[0164] Processor 714 may be implemented as a complex instruction set computer (CISC) or reduced instruction set computer (RISC) processor; an x86 instruction set compatible processor, a multi-core, or any other microprocessor or central processing unit (CPU). In various implementations, processor 714 may be a dual-core processor, a dual-core mobile processor, etc.
[0165] Memory 712 may be implemented as a volatile memory device, such as but not limited to random access memory (RAM), dynamic random access memory (DRAM), or static RAM (SRAM).
[0166] Storage device 711 may be implemented as a non-volatile storage device, such as but not limited to a disk drive, an optical disk drive, a tape drive, an internal storage device, an attached storage device, flash memory, battery-backed SDRAM (synchronous DRAM), and / or a network accessible storage device. In various implementations, for example, when including multiple hard disk drives, storage device 711 may include technologies that enhance protection of storage performance for valuable digital media.
[0167] Graphics subsystem 715 may perform processing of images, such as still or video, for display. Graphics subsystem 715 may be, for example, a graphics processing unit (GPU) or a visual processing unit (VPU). An analog or digital interface may be used to communicatively couple graphics subsystem 715 and display 720. For example, the interface may be any one of a high definition multimedia interface, a DisplayPort, a wireless HDMI, and / or wireless HD compatible technology. Graphics subsystem 715 may be integrated into processor 714 or chipset 705. In some implementations, graphics subsystem 715 may be a separate card communicatively coupled to chipset 705.
[0168] The graphics and / or video processing techniques described herein can be implemented in a variety of hardware architectures. For example, the graphics and / or video functionality can be integrated within a chipset. Alternatively, discrete graphics and / or video processors can be used. As yet another implementation, the graphics and / or video functionality can be provided by a general-purpose processor including a multi-core processor. In other implementations, the functionality can be implemented in consumer electronic devices.
[0169] The radio device 718 can include one or more radio devices capable of transmitting and receiving signals using a variety of suitable wireless communication technologies. Such technologies can involve communication across one or more wireless networks. Example wireless networks include (but are not limited to) wireless local area networks (WLANs), wireless personal area networks (WPANs), wireless metropolitan area networks (WMANs), cellular networks, and satellite networks. When communicating across such networks, the radio device 718 can operate in accordance with one or more applicable standards in any version.
[0170] In various implementations, the display 720 can include any type of television monitor or display. The display 720 can include, for example, a computer display screen, a touchscreen display, a video monitor, a television-like device, and / or a television. The display 720 can be digital and / or analog. In various implementations, the display 720 can be a holographic display. Moreover, the display 720 can be a transparent surface that can receive a visual projection. Such a projection can convey various forms of information, images, and / or objects. For example, such a projection can be a visual overlay for mobile augmented reality (MAR) applications. Under the control of one or more software applications 716, the platform 702 can display a user interface 722 on the display 720.
[0171] In various implementations, the content service device 730 can be hosted by any national, international, and / or independent service and can thus be accessed by the platform 702, for example, via the Internet. The content service device 730 can be coupled to the platform 702 and / or the display 720. The platform 702 and / or the content service device 730 can be coupled to the network 760 to transmit media information to and from the network 760 (e.g., send and / or receive). The content delivery device 740 can also be coupled to the platform 702 and / or the display 720.
[0172] In various implementations, the content service device 730 can include a cable TV box, a personal computer, a network, a telephone, an Internet-enabled device or appliance capable of delivering digital information and / or content, and any other similar device capable of transmitting content unidirectionally or bidirectionally between the content provider and the platform 702 and / or the display 720 via the network 760 or directly. It should be understood that content can be transmitted unidirectionally and / or bidirectionally to and from any one of the components in the system 700 and the content provider. Examples of content can include any media information, including for example video, music, medical, and gaming information, etc.
[0173] The content service device 730 can receive content such as cable TV programs, including media information, digital information, and / or other content. Examples of content providers can include any cable or satellite TV or radio or Internet content provider. The provided examples are not meant to limit the implementations according to the present disclosure in any way.
[0174] In various implementations, the platform 702 can receive a control signal having one or more navigation features from the navigation controller 750. For example, the navigation features of the controller 750 can be used to interact with the user interface 722. In an implementation, the navigation controller 750 can be a pointing device, which can be a computer hardware component (specifically, a human-machine interface device) that allows a user to input spatial (e.g., continuous and multi-dimensional) data into a computer. Many systems such as graphical user interfaces (GUIs) and televisions and monitors allow users to use physical gestures to control and provide data to a computer or a television.
[0175] The movement of the navigation features of the controller 750 can be replicated on a display (e.g., the display 720) by the movement of a pointer, a cursor, a focus ring, or other visual indicators displayed on the display. For example, under the control of the software application 716, the navigation features located on the navigation controller 750 can be mapped to virtual navigation features displayed, for example, on the user interface 722. In an implementation, the controller 750 can not be a separate component, but can be integrated into the platform 702 and / or the display 720. However, the present disclosure is not limited to the elements or contexts shown or described herein.
[0176] In various implementations, a driver (not shown) may include technology that enables a user to immediately turn on and off platform 702. For example, when enabled, the television can utilize a touch of a button to turn on and off after an initial startup. Even when the platform is "off," program logic can allow platform 702 to stream content to a media adapter or other content service device 730 or content delivery device 740. Additionally, for example, chipset 705 may include hardware and / or software support for 7.1 surround sound audio and / or high-definition (7.1) surround sound audio. The driver may include a graphics driver for an integrated graphics platform. In an implementation, the graphics driver may include a fast Peripheral Component Interconnect (PCI) graphics card.
[0177] In various implementations, any one or more of the components shown in system 700 may be integrated. For example, platform 702 and content service device 730 may be integrated, or platform 702 and content delivery device 740 may be integrated, or platform 702, content service device 730, and content delivery device 740 may be integrated. In various implementations, platform 702 and display 720 may be an integrated unit. For example, display 720 and content service device 730 may be integrated, or display 720 and content delivery device 740 may be integrated. These examples are not meant to limit the present disclosure.
[0178] In various implementations, system 700 may be implemented as a wireless system, a wired system, or a combination of both. When implemented as a wireless system, system 700 may include components and interfaces suitable for communicating over a wireless shared medium. For example, one or more antennas, transmitters, receivers, transceivers, amplifiers, filters, control logic, and so on. Examples of wireless shared media may include portions of the wireless spectrum, such as the RF spectrum, etc. When implemented as a wired system, system 700 may include components and interfaces suitable for communicating over a wired communication medium. For example, input / output (I / O) adapters, physical connectors for connecting the I / O adapter to the corresponding wired communication medium, network interface cards (NICs), disk controllers, video controllers, audio controllers, etc. Examples of wired communication media may include wires, cables, metal leads, printed circuit boards (PCBs), backplanes, switch fabrics, semiconductor materials, twisted pairs, coaxial cables, optical fibers, etc.
[0179] Platform 702 may establish one or more logical or physical channels to transmit information. The information may include media information and control information. Media information may refer to any data representing content for a user. Examples of content may include, for example, data from a voice conversation, a video conference, a streaming video, an email message, a voicemail message, alphanumeric symbols, graphics, images, video, text, etc. Data from a voice conversation may be, for example, voice information, silence periods, background noise, comfort noise, tones, etc. Control information may refer to any data representing commands, instructions, or control words for an automated system. For example, control information may be used to route media information through the system or to instruct a node to process media information in a predetermined manner. However, the implementation is not limited to Figure 7 the elements or contexts shown or described
[0180] Reference Figure 8 , the small form factor device 800 is an example of a physical style or form factor in which variations of the system 600 or 700 may be embodied. By this method, the device 800 may be implemented as a mobile computing device with wireless capabilities. For example, a mobile computing device may refer to any device having a processing system and a mobile power source or power supply (e.g., one or more batteries).
[0181] As described above, examples of mobile computing devices may include digital still cameras, digital video cameras, mobile devices with camera or video capabilities (e.g., imaging phones), webcams, personal computers (PCs), laptop computers, ultra-laptop computers, tablet computers, touchpads, portable computers, handheld computers, palmtop computers, personal digital assistants (PDAs), cellular phones, combination cellular phone / PDAs, televisions, smart devices (e.g., smartphones, smart tablet computers, or smart TVs), mobile Internet devices (MIDs), messaging devices, data communication devices, etc.
[0182] Examples of mobile computing devices may also include computers arranged to be worn by an individual, such as, for example, wrist computers, finger computers, ring computers, glasses computers, wallet computers, armband computers, shoe computers, clothing computers, and other wearable computers. In various implementations, for example, a mobile computing device may be implemented as a smartphone capable of executing computer applications as well as voice communication and / or data communication. Although some implementations may be described by way of example using a mobile computing device implemented as a smartphone, it will be understood that other implementations may also be implemented using other wireless mobile computing devices. The implementation is not limited in this context.
[0183] As Figure 8As shown, the device 800 may include a housing having a front portion 801 and a rear portion 802. The device 800 includes a display 804, an input / output (I / O) device 806, and an integrated antenna 808. The device 800 may also include a navigation feature 812. The I / O device 806 may include any suitable I / O device for typing information into the mobile computing device. Examples of the I / O device 806 may include an alphanumeric keyboard, a numeric keypad, a touchpad, input keys, buttons, switches, a microphone, a speaker, a voice recognition device, and software, etc. Information may also be input into the device 800 through the microphone 814, or may be digitized by a voice recognition device. As shown, the device 800 may include a camera 805 (e.g., including at least one lens, an aperture, and an imaging sensor) and a flash 810 integrated into the rear portion 802 (or elsewhere) of the device 800. Implementations are not limited in this context.
[0184] The various forms of the devices and processes described herein may be implemented using hardware elements, software elements, or a combination of both. Examples of hardware elements may include a processor, a microprocessor, a circuit, circuit elements (e.g., transistors, resistors, capacitors, inductors, etc.), an integrated circuit, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a digital signal processor (DSP), a field programmable gate array (FPGA), logic gates, registers, semiconductor devices, chips, microchips, chip sets, etc. Examples of software may include software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, processes, software interfaces, application program interfaces (API), instruction sets, computing code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. Determining whether to implement using hardware elements and / or software elements may vary according to any number of factors, such as the desired computing rate, power level, heat capacity limit, processing cycle budget, input data rate, output data rate, memory resources, data bus speed, and other design or performance constraints.
[0185] One or more of the above aspects may be implemented by representative instructions stored on a machine-readable medium, which represent various logics within a processor and, when read by the machine, cause the machine to fabricate the logics to execute the techniques described herein. Such representations (referred to as “IP cores”) may be stored on a tangible machine-readable medium and supplied to various customers or manufacturing facilities to be loaded into the manufacturing machines of the actual fabrication logics or processors.
[0186] Although the features described herein have been described with reference to various implementations, the description is not intended to be construed in a limiting sense. Accordingly, various modifications of the implementations described herein, as well as other implementations that are apparent to those skilled in the art to which this disclosure pertains, are considered to be within the spirit and scope of this disclosure.
[0187] The following examples relate to additional implementations.
[0188] By Example 1, a computer-implemented method for video encoding and decoding includes: obtaining image data of frames of a video sequence; determining a plurality of reference frames of a current frame of the video sequence, wherein each of the reference frames has at least one motion-compensated (MC) block of image data; generating weights that take into account noise, distortion variance, and dispersion distribution between at least one same MC block position of the plurality of reference frames and a current block of the current frame; and generating denoised and filtered image data, including applying one of the weights to the image data of the at least one MC block.
[0189] By Example 2, the subject matter of Example 1, wherein the weights take into account quantization parameters of an encoder arranged to receive the denoised and filtered image data.
[0190] By Example 3, the subject matter of Example 1 or 2, wherein the dispersion distribution is the distortion variance divided by the average distortion between the same MC block position on the plurality of reference frames and the current block, where the average distortion is the average of the distortions of the plurality of reference frames.
[0191] By Example 4, the subject matter of any one of Examples 1 to 3, wherein generating the weights includes: using block distortion, which is calculated by using both the sum of squared differences (SSD) between the current block and the MC block and the variance of the pixel image data in the current block.
[0192] By Example 5, the subject matter of Example 4, wherein generating the weights includes: selecting a predetermined weight factor value depending on the magnitude of the block distortion.
[0193] By Example 6, the subject matter of Example 4, wherein generating the weights includes: modifying the block distortion by an offset depending on a comparison of each of the noise, distortion variance, and dispersion distribution associated with the current block and the MC block with a threshold.
[0194] By Example 7, the subject matter of Example 4, wherein generating the weights includes: considering a weight block portion and an attenuation block portion, and wherein both the block distortion and the noise are considered in both the weight block portion and the attenuation block portion.
[0195] By Example 8, according to the subject matter of any one of Examples 1 to 7, wherein determining the reference frame includes considering the following: (1) the coding parameters of an encoder for receiving the denoised and filtered image data; (2) the proximity of a scene change to the current frame; and (3) the correlation between the image data on the current frame and the image data on one of the reference frames in the reference frames.
[0196] By Example 9, a computer-implemented system includes: a memory for storing image data of frames of a video sequence; and a processor circuitry communicatively coupled to the memory and arranged to operate by: determining a plurality of reference frames for a current frame of the video sequence, wherein each of the reference frames has at least one motion-compensated (MC) block of image data; generating weights that consider noise, distortion variance, and dispersion distribution between the same MC block positions of the plurality of reference frames and a current block of the current frame; and generating denoised and filtered image data including applying one of the weights to the image data of the MC block.
[0197] By Example 10, according to the subject matter of Example 9, wherein the determining includes: selecting the reference frames for the current frame at least in part depending on an encoding mode that is associated with a reference-frame dependency structure of an encoder for receiving the denoised and filtered image data.
[0198] By Example 11, according to the subject matter of Example 10, wherein the encoding mode is low latency or random access.
[0199] By Example 12, according to the subject matter of any one of Examples 9 to 11, wherein the determining includes: selecting the reference frames for the current frame at least in part depending on whether the current frame is within an available number of consecutive reference frames from a scene start or a scene end, wherein the number includes zero.
[0200] By Example 13, according to the subject matter of any one of Examples 9 to 12, wherein the determining includes: selecting the reference frames for the current frame at least in part depending on the correlation between the image data at the same pixel positions on the current frame and the image data of one of the reference frames in the reference frames.
[0201] By Example 14, according to the subject matter of Example 13, wherein the determining includes: selecting the reference frames for the current frame at least in part depending on comparing a correlation value with a threshold.
[0202] By Example 15, at least one non-transitory article includes: at least one computer-readable medium having instructions stored thereon that, when executed, cause a computing device to operate by: obtaining image data of frames of a video sequence; determining a plurality of reference frames for a current frame of the video sequence, wherein each reference frame of the plurality of reference frames has at least one motion-compensated (MC) block of image data; generating weights that take into account noise, distortion variance, and dispersion distribution between the same MC block positions of the plurality of reference frames and a current block of the current frame; and generating denoised and filtered image data, including applying one of the weights to the image data of the MC block.
[0203] By Example 16, the subject matter of Example 15, wherein even when an equal number of available reference frames are available before and after the current frame and when the available reference frames are closer to the current frame than the closest scene change in the video sequence, the number of determined reference frames before the current frame and the number of determined reference frames after the current frame are different.
[0204] By Example 17, the subject matter of Example 15 or 16, wherein generating the weights includes: selecting a value of a predetermined weight factor at least partially depending on a calculation of noise between the MC block and the current block.
[0205] By Example 18, the subject matter of any one of Examples 15 to 17, wherein generating the weights includes: selecting a value of a predetermined weight factor at least partially depending on a calculation of distortion between the MC block and the current block.
[0206] By Example 19, the subject matter of any one of Examples 15 to 18, wherein generating the weights includes: taking into account an encoder quantization parameter, a difference between the image data of the MC block data and the image data of the current block, and a block distortion, wherein the block distortion takes into account a sum of squared differences between the MC block and the current block and a variance of the image data of the current block.
[0207] By Example 20, the subject matter of Example 19, wherein the block distortion is modified by an offset at least partially depending on the noise, the distortion variance, and the dispersion distribution, wherein the dispersion distribution is the distortion variance divided by a distortion average between the same MC block positions of the plurality of reference frames and the current block on the plurality of reference frames, and the distortion average is an average of the distortions of the plurality of reference frames.
[0208] In another example, at least one machine-readable medium can include multiple instructions that, when executed on a computing device, cause the computing device to perform the method according to any one of the above examples.
[0209] In yet another example, an apparatus can include units for performing the method according to any one of the above examples.
[0210] The above examples can include specific combinations of features. However, the above examples are not limited in this regard, and in various implementations, the above examples can include only a subset of such features, such features in a different order, different combinations of such features, and / or additional features other than those explicitly listed. For example, all features described with respect to any example method herein can be implemented with respect to any example apparatus, example system, and / or example article of manufacture, and vice versa.
[0211] In another example, at least one machine-readable medium can include multiple instructions that, when executed on a computing device, cause the computing device to perform the method according to any one of the above examples.
[0212] In yet another example, an apparatus can include units for performing the method according to any one of the above examples.
[0213] The above examples can include specific combinations of features. However, the above examples are not limited in this regard, and in various implementations, the above examples can include only a subset of such features, such features in a different order, different combinations of such features, and / or additional features other than those explicitly listed. For example, all features described with respect to any example method herein can be implemented with respect to any example apparatus, example system, and / or example article of manufacture, and vice versa.
Claims
1. A computer-implemented method for video encoding and decoding, comprising: obtaining image data of frames of a video sequence; determining a plurality of reference frames for a current frame of the video sequence, wherein each of the reference frames has at least one motion compensated (MC) block of image data; generating weights that take into account noise, distortion variance, and dispersion distribution between at least one same MC block position of the plurality of reference frames and a current block of the current frame; and Generating denoised filtered image data comprises applying one of the weights to the image data of the at least one MC block.
2. The method according to claim 1, wherein: The weights take into account quantization parameters of an encoder arranged to receive the de-noised filtered image data.
3. The method according to claim 1 or 2, wherein: The dispersion distribution is the distortion variance divided by the average distortion between the same MC block position on the multiple reference frames and the current block, wherein the average distortion is an average of the distortions of the multiple reference frames.
4. The method according to claim 1 or 2, wherein: Generating the weight includes using a block distortion calculated by using a sum of squared differences (SSD) between the current block and an MC block, and a variance of pixel image data in the current block.
5. The method according to claim 4, wherein: Generating the weights includes selecting a predetermined weight factor value depending on the magnitude of the block distortion.
6. The method according to claim 4, wherein: Generating the weights includes modifying the block distortion by an offset depending on a comparison of each of noise, distortion variance, and dispersion distribution associated with the current block and the MC block with a threshold.
7. The method according to claim 4, wherein: Generating the weights comprises considering a weight block portion and an attenuation block portion, and wherein both the block distortion and noise are considered in both the weight block portion and the attenuation block portion.
8. The method according to claim 1 or 2, wherein: Determining the reference frame includes considering: (1) encoding parameters of an encoder used to receive the denoised filtered image data; (2) the proximity of the scene change to the current frame; and (3) a correlation between the image data on the current frame and the image data on one of the reference frames.
9. A computer-implemented system comprising: A memory for storing image data of frames of a video sequence; as well as processor circuitry communicatively coupled to the memory and arranged to operate by: determining a plurality of reference frames for a current frame of the video sequence, wherein each of the reference frames has at least one motion compensated (MC) block of image data; generating weights that take into account noise, distortion variance, and dispersion distribution between the same MC block positions of the multiple reference frames and the current block of the current frame; and Generating denoised filtered image data comprises applying one of the weights to the image data of the MC block.
10. The system according to claim 9, wherein: The determining includes selecting a reference frame for the current frame depending at least in part on a coding mode associated with a reference frame dependency structure of an encoder receiving the de-noised filtered image data.
11. The system according to claim 10, wherein: The encoding modes are low latency or random access.
12. The system according to any one of claims 9 to 11, wherein: The determining includes selecting a reference frame for the current frame based at least in part on whether the current frame is within an available number of consecutive reference frames from a scene start or a scene end, wherein the number includes zero.
13. The system according to any one of claims 9 to 11, wherein: The determining includes selecting a reference frame for the current frame based at least in part on a correlation of image data at a same pixel location on the current frame with image data of one of the reference frames.
14. The system according to claim 13, wherein: The determining includes selecting a reference frame for the current frame based at least in part on comparing a correlation value to a threshold value.
15. At least one non-transitory article of manufacture having at least one computer-readable medium having instructions stored thereon that, when executed, cause a computing device to: obtaining image data of frames of a video sequence; Determine a plurality of reference frames of a current frame of the video sequence, wherein: Each reference frame of the plurality of reference frames has at least one motion compensated (MC) block of image data; generating weights that take into account noise, distortion variance, and dispersion distribution between the same MC block positions of the multiple reference frames and the current block of the current frame; as well as Generating denoised filtered image data comprises applying one of the weights to the image data of the MC block.
16. The article of claim 15, wherein: Even if an equal number of available reference frames are available before and after the current frame, and when the available reference frames are closer to the current frame than the closest scene change in the video sequence, the number of reference frames determined before the current frame and the number of reference frames determined after the current frame are different.
17. The article according to claim 15 or 16, wherein Generating the weights includes selecting a predetermined weight factor value depending at least in part on a calculation of noise between the MC block and the current block.
18. The article according to claim 15 or 16, wherein Generating the weights includes selecting a predetermined weight factor value depending at least in part on a calculation of a distortion between the MC block and the current block.
19. The article of claim 15, wherein: Generating the weights includes considering an encoder quantization parameter, a difference between image data of the MC block data and image data of the current block, and block distortion, wherein the block distortion considers the sum of squared differences between the MC block and the current block and the variance of the image data of the current block.
20. The article of claim 19, wherein: The block distortion is modified by an offset depending at least in part on the noise, the distortion variance, and the dispersion distribution, wherein the dispersion distribution is the distortion variance divided by the average of the distortions between the same MC block position on the multiple reference frames and the current block, wherein the distortion average is the average of the distortions of the multiple reference frames.