Side-window bilateral filtering for video encoding and decoding
The side window bilateral filtering technique addresses the challenge of maintaining edge sharpness and reducing noise in video encoding/decoding, enhancing compression performance by adaptively adjusting the encoding/decoding strategy.
Patent Information
- Application Number
- JP2023579005
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-22
- Filing Date
- 2022-06-21
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-06-21
AI Technical Summary
Existing video encoding and decoding technologies face challenges in maintaining edge sharpness while reducing noise through bilateral filtering, leading to suboptimal compression performance.
Implementing a side window filtering technique in bilateral filtering to adaptively adjust the encoding/decoding strategy, incorporating a side window bilateral filtering (SWBIF) process to improve video compression performance by maintaining edge sharpness and reducing noise.
Enhances encoding/decoding efficiency by effectively reducing noise while preserving edge sharpness, thereby improving video compression performance.
Smart Images

Figure 0007821822000030 
Figure 0007821822000031 
Figure 0007821822000032
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure is based on and claims priority to U.S. Provisional Application No. 63 / 213,268, filed June 22, 2021, the entire contents of which are incorporated herein by reference. [Technical Field]
[0002] This disclosure relates to video encoding / decoding and compression, and more particularly, to video processing systems and methods for bilateral filtering with side windows in video encoding / decoding. [Background technology]
[0003] Digital video is supported by a variety of electronic devices, such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video teleconferencing devices, and video streaming devices. The electronic devices transmit, receive, or communicate digital video data over communication networks and / or store the digital video data in storage devices. Due to the limited bandwidth capacity of communication networks and the limited memory resources of storage devices, video coding and decoding may be used to compress video data according to one or more video coding and decoding standards before communicating or storing the video data. For example, video coding and decoding standards include Versatile Video Coding and Decoding (VVC), Joint Search and Test Model (JEM), High Efficiency Video Coding and Decoding (HEVC / H.265), Advanced Video Coding and Decoding (AVC / H.264), Moving Picture Experts Committee (MPEG) coding and decoding, etc. Video coding and decoding typically utilizes prediction methods (e.g., inter-prediction, intra-prediction, etc.) that exploit the redundancy inherent in video data. Video encoding and decoding aims to compress video data into a format with a lower bit rate while avoiding or minimizing degradation of video quality. Summary of the Invention
[0004] An embodiment of the present disclosure provides a video processing method for bilateral filtering with a side window. The video processing method may include receiving, by one or more processors, a video frame of a video for in-loop filtering. The video processing method may include selecting, by the one or more processors, a bilateral filtering window for in-loop filtering of a pixel of interest in the video frame from a group of candidate filtering windows including multiple side filtering windows and one full filtering window. The video processing method may further include filtering, by the one or more processors, the pixel of interest in the video frame with the selected bilateral filtering window.
[0005] An embodiment of the present disclosure further provides a video processing device for bilateral filtering with a side window. The video processing device may include a memory and one or more processors. The memory may be configured to store video. The one or more processors may be configured to receive video frames of video for in-loop filtering. The one or more processors may be further configured to select, for a pixel of interest in the video frame, a bilateral filtering window for in-loop filtering of the pixel of interest from a group of candidate filtering windows including multiple side filtering windows and one full filtering window. The one or more processors may be further configured to filter the pixel of interest in the video frame with the selected bilateral filtering window.
[0006] Embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform a video processing method for bilateral filtering with a side window. The video processing method may include receiving a video frame of video for in-loop filtering. The video processing method may further include selecting, for a pixel of interest in the video frame, a bilateral filtering window for in-loop filtering of the pixel of interest from a group of candidate filtering windows including multiple side filtering windows and one full filtering window. The video processing method may further include filtering the pixel of interest in the video frame with the selected bilateral filtering window. The filtered video frame is stored in the non-transitory computer-readable storage medium.
[0007] It should be noted that the foregoing general description and the following detailed description are merely examples and are not intended to limit the scope of the present disclosure. [Brief explanation of the drawings]
[0008] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0009] [Figure 1] FIG. 1 is a block diagram illustrating an example system for encoding and decoding video blocks in accordance with some embodiments of the present disclosure. [Figure 2] FIG. 1 is a block diagram illustrating an exemplary video encoder according to some embodiments of the present disclosure. [Figure 3] FIG. 2 is a block diagram illustrating an exemplary video decoder according to some embodiments of the present disclosure. [Figure 4]1 is a graphical representation illustrating how a frame is recursively divided into multiple video blocks of different sizes and shapes, according to some embodiments of the present disclosure. [Figure 5] 1 is an illustration of typical edge types in a video frame. [Figure 6] 1 is a graphical representation illustrating an exemplary side filtering window according to some embodiments of the present disclosure. [Figure 7] FIG. 1 is a block diagram illustrating an example process for bilateral filtering with side filtering windows according to some embodiments of the present disclosure. [Figure 8] 1 is a flow diagram of an example method for bilateral filtering with side filtering windows according to some embodiments of the present disclosure. [Figure 9] 9 is a flow diagram of a first exemplary embodiment of a method for bilateral filtering with side filtering windows in FIG. 8 according to some embodiments of the present disclosure. [Figure 10] 9 is a flow diagram of a second exemplary embodiment of a method for bilateral filtering with side filtering windows in FIG. 8 according to some embodiments of the present disclosure. [Figure 11] 9 is a flow diagram of a third exemplary embodiment of a method for bilateral filtering with side filtering windows in FIG. 8 according to some embodiments of the present disclosure. [Figure 12] 10 is a flow diagram of an example method for selecting a filtered block for a video block according to some embodiments of the present disclosure. [Figure 13] FIG. 1 is a block diagram illustrating a computing environment coupled to a user interface according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0010] Reference will now be made in detail to specific embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous non-limiting specific details are set forth in order to provide an understanding of the subject matter described herein. However, it will be apparent to those skilled in the art that various alternative embodiments may be employed and the subject matter may be practiced without these specific details without departing from the scope of the claims. For example, it will be apparent to those skilled in the art that the subject matter described herein may be implemented in multiple types of electronic devices having digital video capabilities.
[0011] It should be noted that terms such as "first," "second," etc. used in the specification and claims of the present disclosure and in the accompanying drawings are used to distinguish between objects and are not used to describe any specific order or sequence. It should be noted that data used in this manner can be substituted under appropriate conditions, thereby enabling the implementation of the embodiments of the present disclosure described herein in an order other than that shown in the accompanying drawings or described in the present disclosure.
[0012] Currently, the VVC standard and the 3rd Generation Audio-Video Coding / Decoding Standard (AVS3) may have one or more in-loop filtering modules, including a deblocking filter (DBF), a sample adaptive offset (SAO) filter, and an adaptive loop filter (ALF). During the development of the VVC standard, a bilateral filter is first proposed to refine the reconstruction block after the inverse transform. Then, the application of the bilateral filter is expanded to become part of the in-loop filtering, allowing it to be used in conjunction with the SAO. For example, the output of the combination of the bilateral filter and the SAO filter can be expressed as follows: I OUT =clip3(I C +ΔI BIF +ΔI SAO ) (1)
[0013] In the above formula (1), I OUT denotes the output of the combination of the bilateral filter and the SAO filter. Cdenotes the brightness of the central sample. ΔI BIF denotes the difference value produced by the bilateral filter using samples from the deblocking filter as input. SAO denotes the offset value produced by the SAO filter using samples from the deblocking filter as input. clip3(·) denotes the clipping function that ensures the output is in the range [minValue,maxValue] and is defined as follows: clip3(x)=min(max(minValue,x),maxValue) (2)
[0014] In Exploratory Experiment (EE) 2 for improved compression beyond VVC capabilities, bilateral filters are considered to be one of the promising tools for further improving compression performance. In some exemplary applications, the pixel filtered by the bilateral filter is located at the center of the filtering window, and such filtering operation may lead to smoothing of edges that are expected to maintain sharpness during filtering. To avoid undesirable smoothing of edges, side window filtering techniques can be adopted to remove noise while maintaining edge sharpness or other types of desirable signals.
[0015] In accordance with the present disclosure, a side window filtering technique is applied to video encoding / decoding to improve encoding / decoding efficiency. To achieve the purpose of the bilateral filter in video encoding / decoding and make the filtered image as close as possible to the original image, the side window selection is carefully considered. According to the present disclosure, the encoding / decoding strategy of side window filtering can be adaptively adjusted to improve compression performance. For example, the quality and statistical properties of video blocks in a video frame vary from block to block, and therefore, a block adaptive mechanism of side window filtering needs to be adopted to improve video compression performance.
[0016] Consistent with the present disclosure, a video processing system and method are disclosed herein to improve the encoding / decoding efficiency of bilateral filtering. For example, the system and method disclosed herein can incorporate a side window filtering mechanism into bilateral filtering, thereby applying a side window bilateral filtering (SWBIF) process to improve the encoding / decoding efficiency of bilateral filtering.
[0017] 1 is a block diagram illustrating an exemplary system 10 for concurrently encoding and decoding video blocks in accordance with some embodiments of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data to be subsequently decoded by a target device 14. Source device 12 and target device 14 may include any of a variety of electronic devices, including desktop or laptop computers, tablet computers, smartphones, set-top boxes, digital televisions, cameras, displays, digital media players, video game consoles, video streaming devices, etc. In some embodiments, source device 12 and target device 14 are equipped with wireless communication capabilities.
[0018] In some embodiments, target device 14 may receive the encoded video data to be decoded via link 16. Link 16 may include any type of communication medium or device capable of transferring encoded video data from source device 12 to target device 14. In one embodiment, link 16 may include a communication medium that enables source device 12 to transmit encoded video data directly to target device 14 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to target device 14. The communication medium may include any wireless communication medium or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from source device 12 to target device 14.
[0019] In some other embodiments, the encoded video data may be transmitted from output interface 22 to storage device 32. The encoded video data in storage device 32 may then be accessed by target device 14 via input interface 28. Storage device 32 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, a Blu-ray disc, a digital versatile disc (DVD), a compact disc read-only memory (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In further embodiments, storage device 32 may correspond to a file server or other intermediate storage device that may store the encoded video data generated by source device 12. Target device 14 may access the stored video data from storage device 32 via streaming or download. A file server may be any type of computer capable of storing encoded video data and transmitting the encoded video data to target device 14. Exemplary file servers include a web server (e.g., for a website), a file transfer protocol (FTP) server, a network-attached storage (NAS) device, or a local disk drive. Target device 14 may access the encoded video data over any standard data connection, including a wireless channel (e.g., a Wi-Fi (Wireless Fidelity) connection), a wired connection (e.g., a DSL (Digital Subscriber Line), a cable modem, etc.), or a combination thereof, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 32 may be a streaming transmission, a download transmission, or a combination of both.
[0020] 1 , source device 12 includes video source 18, video encoder 20, and output interface 22. Video source 18 may include sources such as a video capture device such as a video camera, a video archive containing previously captured video, a video feed interface receiving video data from a video content provider, and / or a computer graphics system generating computer graphics data as source video, or a combination of these sources. As an example, if video source 18 is a video camera in a security surveillance system, source device 12 and target device 14 may include a camera phone or video phone. However, embodiments described in this disclosure may be applicable to video encoding and decoding generally, and may be applicable to wireless and / or wired applications.
[0021] The captured, pre-captured, or computer-generated video may be encoded by video encoder 20. The encoded video data may be transmitted directly to target device 14 via output interface 22 of source device 12. The encoded video data may also (or alternatively) be stored in storage device 32 for later access and decoding and / or playback by target device 14 or another device. Output interface 22 may further include a modem and / or a transmitter.
[0022] Target device 14 includes an input interface 28, a video decoder 30, and a display device 34. Input interface 28 may include a receiver and / or a modem to receive encoded video data over link 16. The encoded video data communicated over link 16 or provided to storage device 32 may include various syntax elements generated by video encoder 20 that are used by video decoder 30 in decoding the video data. Such syntax elements may be included in the encoded video data transmitted over a communications medium, stored on a storage medium, or stored on a file server.
[0023] In some embodiments, target device 14 may include a display device 34, which may be an integrated display device and an external display device configured to communicate with target device 14. Display device 34 displays the decoded video data to a user and may include any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.
[0024] Video encoder 20 and video decoder 30 may operate in accordance with proprietary or industry standards, such as, for example, VVC, HEVC, MPEG-4 Part 10, AVC, or extensions of those standards. It should be understood that this disclosure is not limited to a particular video encoding / decoding standard and may be applicable to other video encoding / decoding standards. It is contemplated that video encoder 20 of source device 12 may generally be configured to encode video data in accordance with any of these current or future standards. Similarly, it is contemplated that video decoder 30 of target device 14 may generally be configured to decode video data in accordance with any of these current or future standards.
[0025] Video encoder 20 and video decoder 30 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. If implemented partially in software, the electronic device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware by one or more processors to perform the video encoding / decoding operations disclosed in this disclosure. Video encoder 20 and video decoder 30 may each be included in one or more encoders or decoders, or either may be integrated within the respective device as part of a combined encoder / decoder (CODEC).
[0026] 2 is a block diagram illustrating an example video encoder 20 according to some embodiments described herein. The video encoder 20 can perform intra-predictive and inter-predictive coding of video blocks within video frames. Intra-predictive coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or image. Inter-predictive coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or images of a video sequence. Note that the term "frame" may be used synonymously with the terms "image" or "picture" in the field of video encoding and decoding.
[0027] As shown in FIG. 2, video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a partitioning unit 45, an intra-prediction processing unit 46, and an intra-block copy (BC) unit 48. In some embodiments, video encoder 20 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62 for video block reconstruction. An in-loop filter 63, such as a deblocking filter, may be disposed between adder 62 and DPB 64 to filter block boundaries and remove block artifacts from the reconstructed video data. In addition to the deblocking filter filtering the output of adder 62, another in-loop filter, such as an SAO filter and / or an adaptive in-loop filter (ALF), may also be employed. In some embodiments, the in-loop filter may be omitted, and the decoded video block may be provided directly to DPB 64 by summer 62. Video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided into one or more of the illustrated fixed or programmable hardware units.
[0028] Video data memory 40 may store video data to be encoded by components of video encoder 20. The video data in video data memory 40 may be obtained, for example, from video source 18 shown in FIG. 1. DPB 64 is a buffer that stores reference video data (e.g., reference frames or images) employed in encoding video data by video encoder 20 (e.g., in intra-predictive coding mode or inter-predictive coding mode). Video data memory 40 and DPB 64 may be formed by any of a variety of memory devices. In various embodiments, video data memory 40 may be on-chip with other components of video encoder 20 or off-chip relative to those components.
[0029] As shown in FIG. 2, after receiving video data, a partitioning unit 45 in the prediction processing unit 41 partitions the video data into video blocks. This partitioning may include dividing the video frame into slices, tiles (e.g., sets of video blocks), or other larger coding units (CUs) according to a predetermined partitioning structure, such as a quadtree (QT) structure associated with the video data. A video frame may be, or may be considered to be, a two-dimensional array or matrix of samples having sample values. The samples in the array may be referred to as pixels or picture elements. The plurality of samples in the horizontal and vertical directions (or axes) of the array or image defines the size and / or resolution of the video frame. The video frame may be partitioned into a plurality of video blocks, for example, using QT partitioning. A video block has smaller dimensions than a video frame, but may also be, or may be considered to be, a two-dimensional array or matrix of samples having sample values. The plurality of samples in the horizontal and vertical directions (or axes) of a video block defines the size of the video block. A video block may be further divided into one or more block partitions or sub-block partitions (which may again form blocks), e.g., by iteratively employing QT partitioning, binary tree (BT) partitioning, ternary tree (TT) partitioning, or any combination thereof. Note that, as used herein, the term "block" or "video block" may refer to a portion of a frame or image, particularly a rectangular (square or non-square) portion. For example, with reference to HEVC and VVC, a block or video block may be or correspond to a coding tree unit (CTU), CU, prediction unit (PU), or transform unit (TU), and / or a corresponding block such as a coding tree block (CTB), coding block (CB), prediction block (PB), or transform block (TB). Additionally or alternatively, a block or video block may be or correspond to a sub-block such as a CTB, CB, PB, TB, etc.
[0030] Prediction processing unit 41 may select one of multiple possible predictive coding modes, such as one of multiple intra-predictive coding modes or one of multiple inter-predictive coding modes, for the current video block based on the error result (e.g., coding rate and distortion level). Prediction processing unit 41 provides the resulting intra- or inter-predictively coded block (e.g., a predictive block) to summer 50 to generate a residual block and to summer 62 to reconstruct a coded block for later use as part of a reference frame. Prediction processing unit 41 also provides syntax elements such as motion vectors, intra-mode indicators, partition information, and other such syntax information to entropy coding unit 56.
[0031] To select an appropriate intra-prediction coding mode for a current video block, intra-prediction processing unit 46 within prediction processing unit 41 performs intra-prediction coding of the current video block relative to one or more neighboring blocks in the same frame as the current block being coded to provide spatial prediction. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 perform inter-prediction coding of the current video block relative to one or more predictive blocks in one or more reference frames to provide temporal prediction. Video encoder 20 may, for example, perform multiple coding passes to select an appropriate coding mode for each block of video data.
[0032] In some embodiments, motion estimation unit 42 determines the inter-prediction mode for a current video frame by generating a motion vector, which indicates the displacement of a video block in the current video frame relative to a predictive block in a reference frame according to a predetermined pattern within the sequence of video frames. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. The motion vector may, for example, indicate the displacement of a video block in a current video frame or image relative to a predictive block in a reference frame. The predetermined pattern may designate a video frame in the sequence as a P-frame or a B-frame. Intra BC unit 48 may determine vectors, e.g., block vectors, for intra BC coding in a manner similar to how motion estimation unit 42 determines motion vectors for inter prediction, and may utilize motion estimation unit 42 to determine the block vectors.
[0033] A prediction block for a video block may be or may correspond to a block of a reference frame or reference block that is deemed to closely match the video block being encoded in terms of pixel differences, which may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some embodiments, video encoder 20 may calculate values for sub-integer pixel positions of the reference frame stored in DPB 64. For example, video encoder 20 may interpolate values for quarter-pixel positions, eighth-pixel positions, and other fractional pixel positions of the reference frame. Thus, motion estimation unit 42 may perform motion search for whole pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.
[0034] Motion estimation unit 42 calculates a motion vector for a video block in an inter-predictively coded frame by comparing the position of the video block with the position of a predictive block of a reference frame selected from a first reference frame list (List 0) or a second reference frame list (List 1), each of which identifies one or more reference frames stored in DPB 64. Motion estimation unit 42 sends the calculated motion vector to motion compensation unit 44 and then to entropy coding unit 56.
[0035] The motion compensation performed by motion compensation unit 44 may include fetching or generating a predictive block based on the motion vector determined by motion estimation unit 42. Upon receiving the motion vector for the current video block, motion compensation unit 44 may identify the predictive block to which the motion vector points in one of the reference frame lists, retrieve the predictive block from DPB 64, and forward the predictive block to summer 50. Summer 50 then forms a residual block of pixel difference values by subtracting pixel values of the predictive block provided by motion compensation unit 44 from pixel values of the current video block being coded. The pixel difference values forming the residual block may include a luma difference component, a chroma difference component, or both. Motion compensation unit 44 may also generate syntax elements associated with the video blocks of the video frame for use by video decoder 30 in decoding the video blocks of the video frame. The syntax elements may include, for example, a syntax element defining the motion vector used to identify the predictive block, any flags indicating a prediction mode, or any other syntax information described herein. It should be noted that motion estimation unit 42 and motion compensation unit 44 may be integrated together and are shown separately in FIG. 2 for conceptual purposes.
[0036] In some embodiments, the intra BC unit 48 may generate vectors and fetch predictive blocks in a manner similar to that described above in connection with the motion estimation unit 42 and the motion compensation unit 44, except that the predictive block is within the same frame as the current block being coded, and the vectors are referred to as block vectors, as opposed to motion vectors. In particular, the intra BC unit 48 may determine the intra prediction mode to be used to code the current block. In some implementations, the intra BC unit 48 may code the current block with various intra prediction modes, e.g., during separate coding passes, and test their performance through rate-distortion analysis. The intra BC unit 48 may then select and use an appropriate intra prediction mode from the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 may calculate rate-distortion values for the various tested intra prediction modes through rate-distortion analysis, and select the intra prediction mode with the best rate-distortion characteristics from the tested modes as the appropriate intra prediction mode to use. Rate-distortion analysis generally determines the amount of distortion (or error) between a coded block and the original uncoded block that was coded to generate the coded block, and the bitrate (i.e., number of bits) used to generate the coded block. Intra BC unit 48 can calculate ratios from the distortions and rates of various coded blocks to determine which intra-prediction mode exhibits the optimal rate-distortion value for the block.
[0037] In other implementations, intra BC unit 48 may employ, in whole or in part, motion estimation unit 42 and motion compensation unit 44 to perform such functions for intra BC prediction in accordance with embodiments described herein. In either case, for intra block copying, the predictive block may be a block that is deemed to closely match the block being coded in terms of pixel differences, which may be determined by SAD, SSD, or other difference metrics, and identification of the predictive block may include calculating values for sub-integer pixel positions.
[0038] Regardless of whether the predictive block is from the same frame via intra prediction or a different frame via inter prediction, video encoder 20 may form a residual block by subtracting pixel values of the predictive block from pixel values of the current video block being coded to form pixel difference values. The pixel difference values forming the residual block may include both luma and chroma component differences.
[0039] Intra-prediction processing unit 46 may intra-predict the current video block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44 or intra-block copy prediction performed by intra BC unit 48, as described above. In particular, intra-prediction processing unit 46 may determine the intra-prediction mode to be employed to encode the current block. For example, intra-prediction processing unit 46 may encode the current block in various intra-prediction modes, e.g., during separate encoding passes, and intra-prediction processing unit 46 (or, in some embodiments, a mode selection unit) may select an appropriate intra-prediction mode to use from the tested intra-prediction modes. Intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to entropy coding unit 56. Entropy coding unit 56 may encode the information indicating the selected intra-prediction mode into the bitstream.
[0040] After prediction processing unit 41 determines a predictive block for the current video block via either inter-prediction or intra-prediction, adder 50 forms a residual block by subtracting the predictive block from the current video block. The residual video data in the residual block, which may be included in one or more TUs, is provided to transform processing unit 52. Transform processing unit 52 employs a transform, such as a discrete cosine transform (DCT) or a conceptually similar transform, to convert the residual video data into transform coefficients.
[0041] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54, which quantizes the transform coefficients to further reduce the bit rate. The quantization process may also reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be varied by adjusting a quantization parameter. In some embodiments, quantization unit 54 may then perform a scan of a matrix containing the quantized transform coefficients. Alternatively, entropy coding unit 56 may perform the scan.
[0042] Following quantization, entropy coding unit 56 employs an entropy coding technique, such as context-adaptive variable length coding (CAVLC), context-adaptive binary arithmetic coding (CABAC), syntactic context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding method or technique, to encode the quantized transform coefficients into a video bitstream. The encoded bitstream may then be transmitted to video decoder 30 shown in FIG. 1 or archived in storage device 32 shown in FIG. 1 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 may also employ entropy coding techniques to encode motion vectors and other syntax elements for the current video frame being coded.
[0043] Inverse quantization unit 58 and inverse transform processing unit 60 may apply inverse quantization and inverse transform, respectively, to reconstruct residual blocks in the pixel domain for generating reference blocks for predicting other video blocks. The reconstructed residual blocks may be generated. As described above, motion compensation unit 44 may generate motion-compensated prediction blocks from one or more reference blocks of frames stored in DPB 64. Motion compensation unit 44 may also apply one or more interpolation filters to the prediction blocks to calculate sub-integer pixel values used for motion estimation.
[0044] Adder 62 adds the reconstructed residual block to the motion compensated prediction block generated by motion compensation unit 44 to generate a reference block for storage in DPB 64. The reference block may then be used by intra BC unit 48, motion estimation unit 42, and motion compensation unit 44 as a prediction block for inter predicting another video block in a subsequent video frame.
[0045] 3 is a block diagram illustrating an exemplary video decoder 30 according to some embodiments of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 may perform a decoding process that is generally inverse to the encoding process described above for the video encoder 20 in connection with FIG. 2. For example, the motion compensation unit 82 may generate prediction data based on a motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 may generate prediction data based on an intra prediction mode indicator received from the entropy decoding unit 80.
[0046] In some implementations, units of video decoder 30 may be tasked with performing embodiments of the present disclosure. Also, in some implementations, embodiments of the present disclosure may be divided among one or more of the units of video decoder 30. For example, intra BC unit 85 may perform embodiments of the present disclosure alone or in combination with other units of video decoder 30, such as motion compensation unit 82, intra prediction unit 84, and entropy decoding unit 80. In some implementations, video decoder 30 may not include intra BC unit 85, and the functionality of intra BC unit 85 may be performed by other components of prediction processing unit 81, such as motion compensation unit 82.
[0047] Video data memory 79 may store video data, such as an encoded video bitstream, for decoding by other components of video decoder 30. The video data stored in video data memory 79 may be obtained, for example, from storage device 32, from a local video source such as a camera, via wired or wireless network communication of video data, or by accessing a physical data storage medium (e.g., a flash drive or hard disk). Video data memory 79 may include a coded picture buffer (CPB) that stores coded video data from the coded video bitstream. A data buffer 92 of video decoder 30 stores reference video data used by video decoder 30 in decoding video data (e.g., in intra- or inter-prediction coding modes). Video data memory 79 and DPB 92 may be formed by any of a variety of memory devices, such as dynamic random access memory (DRAM), including synchronous dynamic random access memory (SDRAM), magnetoresistive random access memory (MRAM), resistive random access memory (RRAM), or other types of memory devices. For ease of explanation, video data memory 79 and DPB 92 are shown as two separate components of video decoder 30 in FIG. 3 . However, it will be apparent to those skilled in the art that video data memory 79 and DPB 92 may be provided by the same memory device or by separate memory devices. In some embodiments, video data memory 79 may be on-chip with other components of video decoder 30 or may be off-chip with respect to those components.
[0048] During the decoding process, video decoder 30 receives an encoded video bitstream representing video blocks and associated syntax elements of encoded video frames. Video decoder 30 may receive the syntax elements at the video frame level and / or the video block level. Entropy decoding unit 80 of video decoder 30 employs entropy decoding techniques to decode the bitstream to obtain quantized coefficients, motion vectors or intra-prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then forwards the motion vectors or intra-prediction mode indicators and other syntax elements to prediction processing unit 81.
[0049] If the video frame is coded as an intra-predictively coded (e.g., I) frame, or for intra-coded predictive blocks within other types of frames, intra-prediction unit 84 of prediction processing unit 81 may generate predictive data for video blocks of the current video frame based on the signaled intra-prediction mode and reference data from previously decoded blocks of the current frame.
[0050] If a video frame is coded as an inter-predictive (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 generates one or more predictive blocks of a video block of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the predictive blocks may be generated from one reference frame of the reference frame lists. Video decoder 30 may employ a default construction technique to construct the reference frame lists, e.g., List 0 and List 1, based on the reference frames stored in DPB 92.
[0051] In some embodiments, when a video block is encoded according to the intra BC mode described herein, intra BC unit 85 of prediction processing unit 81 generates a predictive block of the current video block based on the block vectors and other syntax elements received from entropy decoding unit 80. The predictive block may be within the same reconstructed region of the image as the current video block processed by video encoder 20.
[0052] Motion compensation unit 82 and / or intra BC unit 85 parse the motion vectors and other syntax elements to determine prediction information for video blocks of the current video frame, and then generate a predictive block for the current video block being decoded using the prediction information. For example, motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra-prediction or inter-prediction) used to encode the video blocks of the video frame, the inter-prediction frame type (e.g., B or P), configuration information for one or more of the frame's reference frame list, the motion vector for each inter-predictively coded video block of the frame, the inter-prediction state for each inter-predictively coded video block of the frame, and other information for decoding video blocks in the current video frame.
[0053] Similarly, intra BC unit 85 uses some of the received syntax elements, such as flags, to determine that the current video block is predicted in intra BC mode, configuration information about which video blocks of the frame are in the reconstruction domain and should be stored in DPB 92, block vectors for each intra BC predicted video block of the frame, intra BC prediction states for each intra BC predicted video block of the frame, and other information for decoding video blocks in the current video frame.
[0054] Motion compensation unit 82 may also perform interpolation using an interpolation filter, such as that employed by video encoder 20 during encoding of the video block, to calculate interpolated values for sub-integer pixels of the reference block. In this case, motion compensation unit 82 may determine the interpolation filter employed by video encoder 20 from the received syntax element and generate the predictive block using the interpolation filter.
[0055] Inverse quantization unit 86 employs the same quantization parameters calculated by video encoder 20 for each video block in a video frame to inverse quantize the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 80 to determine the degree of quantization. Inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients in order to reconstruct the residual block in the pixel domain.
[0056] After motion compensation unit 82 or intra BC unit 85 generates a predictive block for the current video block based on the vectors and other syntax elements, summer 90 reconstructs a decoded video block for the current video block by adding a residual block from inverse transform processing unit 88 and the corresponding predictive block generated by motion compensation unit 82 and intra BC unit 85. The decoded video block may also be referred to as a reconstructed block of the current video block. An in-loop filter 91, such as a deblocking filter, an SAO filter, and / or an ALF, may be disposed between summer 90 and DPB 92 to further process the decoded video block. In some embodiments, in-loop filter 91 may be omitted, and the decoded video block may be provided directly to DPB 92 by summer 90. The decoded video block in a given frame is then stored in DPB 92, which stores reference frames for subsequent motion compensation of the next video block. DPB 92, or a memory device separate from DPB 92, may store the decoded video for later display on a display device, such as display device 34 of FIG. 1.
[0057] In a typical video encoding / decoding process (e.g., including a video encoding process and a video decoding process), a video sequence typically includes an ordered set of frames or images. Each frame may include three sample arrays, denoted SL, SCb, and SCr. SL is a two-dimensional array of luma samples. SCb is a two-dimensional array of Cb chroma samples. SCr is a two-dimensional array of Cr chroma samples. In another example, a frame may be monochrome and therefore include only a two-dimensional array of luma samples.
[0058] As shown in FIG. 4A, video encoder 20 (or more specifically, division unit 45) generates a coded representation of a frame by first dividing the frame into a set of CTUs. A video frame may include an integer number of CTUs arranged consecutively in raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by video encoder 20 in a sequence parameter set so that all CTUs in a video sequence have the same size, which may be one of 128×128, 64×64, 32×32, and 16×16. However, it should be noted that CTUs in this disclosure are not necessarily limited to a particular size. As shown in FIG. 4B, each CTU may include one CTB for luma samples, two corresponding coding tree blocks for chroma samples, and syntax elements for encoding and decoding the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of coding blocks of pixels, including inter or intra prediction, intra prediction mode, motion vectors, and other parameters, and how the video sequence is reconstructed in video decoder 30. For monochrome images or images with three distinct color planes, a CTU may include a single coding tree block and syntax elements for encoding and decoding samples of the coding tree block. A coding tree block may be an N×N block of samples.
[0059] To achieve superior performance, video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree block of a CTU to partition the CTU into smaller CUs. As shown in FIG. 4C , a 64×64 CTU 400 is first partitioned into four smaller CUs, each with a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each partitioned into four 16×16 CUs by block size. Two 16×16 CUs, 430 and 440, are each further partitioned into four 8×8 CUs by block size. FIG. 4D shows a quad tree data structure illustrating the final result of the partitioning process of CTU 400 as shown in FIG. 4C , where each leaf node of the quad tree corresponds to one CU of a respective size ranging from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B , each CU may include a CB for luma samples, two corresponding coding blocks for chroma samples of the same size frame, and syntax elements for encoding and decoding the samples of the coding block. In monochrome images or images with three separate color planes, a CU may include a single coding block and syntax structures for encoding and decoding the samples of the coding block. Note that the quadtree partitioning shown in FIGS. 4C and 4D is for illustrative purposes only; a CTU can be divided into CUs to accommodate various local characteristics based on quadtree, ternary, or binary tree partitioning. In a multi-type tree structure, a CTU is partitioned by a quadtree structure, and each quadtree leaf CU can be further partitioned by binary and ternary tree structures. As shown in FIG. 4E , there are multiple possible partition types for a coding block with width W and height H: quad partition, vertical 2-partition, horizontal 2-partition, vertical 3-partition, vertically extended 3-partition, horizontally extended 3-partition, and horizontally extended 3-partition.
[0060] In some embodiments, video encoder 20 may further divide the coding blocks of a CU into one or more M×N PBs. A PB may include rectangular (square or non-square) blocks of samples to which the same prediction, inter or intra prediction, is applied. A PU of a CU may include a PB of luma samples, two corresponding PBs of chroma samples, and syntax elements for predicting the PBs. In monochrome images or images with three separate color planes, a PU may include a single PB and syntax structures for predicting the PB. Video encoder 20 may generate predicted luma blocks, Cb blocks, and Cr blocks for the luma PB, Cb PB, and Cr PB of each PU of a CU.
[0061] Video encoder 20 may employ intra prediction or inter prediction to generate the predictive block of a PU. If video encoder 20 employs intra prediction to generate the predictive block of a PU, it may generate the predictive block of the PU based on decoded samples of a frame associated with the PU. If video encoder 20 employs inter prediction to generate the predictive block of a PU, it may generate the predictive block of the PU based on decoded samples of one or more frames other than the frame associated with the PU.
[0062] After video encoder 20 generates the predicted luma block, Cb block, and Cr block of one or more PUs of a CU, video encoder 20 may generate a luma residual block of the CU by subtracting the predicted luma block of the CU from the original luma coding block such that each sample in the luma residual block of the CU indicates a difference between a luma sample in one of the predicted luma blocks of the CU and a corresponding sample in the original luma coding block of the CU. Similarly, video encoder 20 may generate a Cb residual block and a Cr residual block of the CU such that each sample in the Cb residual block of the CU indicates a difference between a Cb sample in one of the predicted Cb blocks of the CU and a corresponding sample in the original Cb coding block of the CU, and each sample in the Cr residual block of the CU indicates a difference between a Cr sample in one of the predicted Cr blocks of the CU and a corresponding sample in the original Cr coding block of the CU.
[0063] Furthermore, as shown in FIG. 4C , video encoder 20 may employ quadtree partitioning to decompose the luma, Cb, and Cr residual blocks of a CU into one or more luma, Cb, and Cr transform blocks, respectively. A transform block may include a rectangular (square or non-square) block of samples to which the same transform is applied. A TU of a CU may include a transform block of luma samples, two corresponding transform blocks of chroma samples, and syntax elements for transforming the transform block samples. Thus, each TU of a CU may be associated with a luma transform block, a Cb transform block, and a Cr transform block. In some embodiments, the luma transform block associated with a TU may be a sub-block of the luma residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In monochrome images or images with three separate color planes, a TU may include a single transform block and syntax structures for transforming the samples of the transform block.
[0064] Video encoder 20 may apply one or more transforms to a luma transform block of a TU to generate a luma coefficient block of the TU. The coefficient block may be a two-dimensional array of transform coefficients. The transform coefficients may be scalar quantities. Video encoder 20 may apply one or more transforms to a Cb transform block of the TU to generate a Cb coefficient block of the TU. Video encoder 20 may apply one or more transforms to a Cr transform block of the TU to generate a Cr coefficient block of the TU.
[0065] After generating a coefficient block (e.g., a luma coefficient block, a Cb coefficient block, or a Cr coefficient block), video encoder 20 may quantize the coefficient block. Quantization generally refers to a process of quantizing transform coefficients to potentially reduce the amount of data for representing the transform coefficients and provide further compression. After quantizing the coefficient block, video encoder 20 may apply an entropy coding technique to encode syntax elements indicating the quantized transform coefficients. For example, video encoder 20 may perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, video encoder 20 may output a bitstream including a sequence of bits forming a representation of the encoded frame and associated data, which may be either stored in storage device 32 or transmitted to target device 14.
[0066] After receiving the bitstream generated by video encoder 20, video decoder 30 may parse the bitstream to obtain syntax elements from the bitstream. Video decoder 30 may reconstruct frames of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing video data is generally the reverse of the encoding process performed by video encoder 20. For example, video decoder 30 may perform an inverse transform on coefficient blocks associated with TUs of a current CU to reconstruct residual blocks associated with the TUs of the current CU. Video decoder 30 also reconstructs coding blocks of the current CU by adding samples of predictive blocks of PUs of the current CU to corresponding samples of transform blocks of the TUs of the current CU. After reconstructing the coding blocks of each CU of the frame, video decoder 30 may reconstruct the frame.
[0067] As mentioned above, video coding and decoding mainly realizes video compression through two modes: intra-frame prediction (or intra-prediction) and inter-frame prediction (or inter-prediction). Note that intra-block copy (IBC) can be considered as intra-frame prediction or a third mode. Between the two modes, inter-frame prediction contributes to coding and decoding efficiency more than intra-frame prediction because it employs a motion vector to predict a current video block from a reference video block.
[0068] However, as video data collection technology improves day by day, the video block size that preserves the details of the video data becomes finer, and the amount of data required to represent the motion vector of the current frame also increases substantially. One way to solve this problem is to benefit from the fact that a group of adjacent CUs in both the spatial and temporal domains not only have similar video data for prediction, but also similar motion vectors between these adjacent CUs. Therefore, the motion information of spatially adjacent CUs and / or temporally co-located CUs can be used as an approximation of the motion information (e.g., motion vector) of the current CU by exploring their spatial and temporal correlations, also called the "motion vector predictor (MVP)" of the current CU.
[0069] Instead of encoding the actual motion vector of the current CU into the video bitstream (the actual motion vector is determined by motion estimation unit 42 as described above in connection with FIG. 2), the motion vector predictor of the current CU is subtracted from the actual motion vector of the current CU to generate a motion vector difference (MVD) for the current CU. This eliminates the need to encode the motion vector determined by motion estimation unit 42 for each CU of a frame into the video bitstream, significantly reducing the amount of data required to represent motion information in the video bitstream.
[0070] Similar to the process of selecting a predictive block in a reference frame during inter-frame prediction of a coding block, both video encoder 20 and video decoder 30 can use a set of rules to build a motion vector candidate list (also called a "merge list") for the current CU using potential candidate motion vectors associated with spatially adjacent CUs and / or temporally co-located CUs of the current CU, and then select one member from the motion vector candidate list as the motion vector predictor for the current CU. This eliminates the need to transmit the motion vector candidate list itself from video encoder 20 to video decoder 30; the index of a motion vector predictor selected from the motion vector candidate list is sufficient for video encoder 20 and video decoder 30 to encode and decode the current CU using the same motion vector predictor in the motion vector candidate list. Therefore, only the index of the selected motion vector predictor needs to be transmitted from video encoder 20 to video decoder 30.
[0071] This specification provides a brief description of bilateral filtering. For a filter kernel for bilateral filtering, the contribution of each sample in a filtering window depends not only on the spatial distance between the samples but also on the luminance difference between the samples. For example, a sample at position (i,j) can be filtered using a neighboring sample at position (k,l). The weight ω(i,j,k,l) assigned to sample (k,l) for filtering sample (i,j) can be expressed as follows:
number
[0072] In the above formula (3), I(i,j) and I(k,l) represent the luminance values of samples (i,j) and (k,l), respectively. d denotes the spatial intensity parameter. r denotes the brightness intensity parameter. The strength of the bilateral filter is σ d and σr The output sample of the bilateral filter for position (i,j) may be a weighted average of the samples within the filtering window. For example,
number
[0073] In the above formula (4), S out(i,j ) denotes the output sample at position (i,j). S(k,l) denotes the sample at position (k,l) in the filtering window, where 1≦k≦K and 1≦l≦L (e.g., S(k,l) with 1≦k≦K and 1≦l≦L denotes a pixel in the filtering window). ω(i,j,k,l) denotes the weight of S(k,l). The product KL denotes the total number of samples in the filtering window.
[0074] In some embodiments, the picture parameter set RBSP (Raw Byte Sequence Payload) syntax, slice header syntax, and coding tree unit syntax of the bilateral filter may be realized by the following Tables 1, 2, and 3, respectively. [Table 1] [Table 2] [Table 3]
[0075] This specification also provides a brief description of side window filtering (SWF) with reference to FIGS. 5 and 6, where FIG. 5 illustrates representative edge types in a video frame and FIG. 6 illustrates a representative side filtering window according to some embodiments of the present disclosure. For image filtering, it is desirable to remove noise while preserving edges and other signal details in the image. In some application scenarios, the image filtering method may be based on a linear approximation, which assumes that the image is piecewise linear. The pixel to be filtered may then be centered in a local filtering window, and the corresponding filtered pixel may be calculated as a weighted average of neighboring pixels on the local filtering window associated with the pixel to be filtered. For some edges, e.g., step edges, ramp edges, or roof edges (as shown in FIG. 5), the sharpness of the edge can be preserved after filtering by limiting the filtering window for each edge to one side of the respective edge.
[0076] For example, for pixel "a" on the step edge 510 shown in FIG. 5, a side filtering window 512 may be limited to the left of pixel "a" and employed to filter pixel "a." For pixel "b" on the step edge 510, a side filtering window 514 may be limited to the right of pixel "b" and employed to filter pixel "b." Similarly, for pixel "c" on the ramp edge 520, a side filtering window 522 limited to one side of pixel "c" may be employed to filter pixel "c." For pixel "d" on the ramp edge 520, a side filtering window 524 limited to one side of pixel "d" may be employed to filter pixel "d." Also, for pixel "e" on the roof edge 530, a side filtering window 532 limited to one side of pixel "e" may be employed to filter pixel "e." For pixel "f" on the roof edge 530, a side filtering window 534 limited to one side of pixel "f" may be employed to filter pixel "f."
[0077] 6 illustrates various side filtering windows according to some embodiments of the present disclosure. For example, part (a) of FIG. 6 illustrates a continuous side filtering window 610 that can be described by three parameters (r, θ, ρ). r indicates the radius of the side filtering window, θ indicates the angle between the window and the horizon, and ρ indicates the position of the pixel of interest (x, y). A fixed-pattern side filtering window may be a special case of a continuous filtering window.
[0078] In practice, as shown in parts (b) to (d) of Figure 6, at least eight fixed-pattern side filtering windows can simplify the filtering calculation. For example, part (b) of Figure 6 shows the left side filtering window (L) and the right side filtering window (R) of the pixel of interest (x, y). Part (c) of Figure 6 shows the upper side filtering window (U) and the lower side filtering window (D) of the pixel of interest (x, y). Part (d) of Figure 6 shows the northeast filtering window (NE), the southeast filtering window (SE), the southwest filtering window (SW), and the northwest filtering window (NW) of the pixel of interest (x, y), respectively.
[0079] When applying a filter kernel with eight side filtering windows to each pixel on an edge, as shown in parts (b) to (d) of FIG. 6, eight possible filtered outputs can be obtained. One of the eight possible filtered outputs (and, accordingly, one of the eight side filtering windows) can be selected as the output filtered pixel to minimize the distance between the input original pixel at the edge and the output filtered pixel, thereby preserving the sharpness of the edge. That is, the output filtered pixel may be the same as the input original pixel at the edge or may be as close as possible to the input original pixel, so the filtered output with the smallest distance to the input luminance may be selected as the output filtered pixel. By adjusting the corresponding filtering window and renormalizing the corresponding filter weights, any filter kernel can be applied to the side window filtering scheme disclosed herein, and the type of filter kernel is not limited herein.
[0080] Consistent with this disclosure, the side window filtering scheme disclosed herein can improve the encoding / decoding performance of in-loop filtering beyond the VVC and AVS3 standards. In some embodiments, bilateral filtering can work in conjunction with an SAO filter to improve encoding / decoding efficiency. As an example, a bilateral filtering technique is employed to explain the spirit of the side window filtering scheme in this disclosure. The side window filtering scheme disclosed herein is believed to be applicable to any filtering module in modern video encoding / decoding techniques and is not limited to bilateral filters.
[0081] 7 is a block diagram illustrating an example process of bilateral filtering with side filtering windows according to some embodiments of the present disclosure. In some embodiments, the example process of FIG. 7 may be performed by in-loop filter 63 of video encoder 20 or in-loop filter 91 of video decoder 30. In some embodiments, the example process of FIG. 7 may be performed by a video processor on the encoder side or the decoder side (e.g., processor 1320 shown in FIG. 13). For illustrative purposes only, the example process of FIG. 7 is described below with respect to a video processor.
[0082] First, a video processor may receive a video frame of video for in-loop filtering. The video frame may include multiple video blocks, each including multiple pixels of interest to be processed. For each pixel of interest in a video block in the video frame, the video processor may perform bilateral filtering with side windows 700 on the pixel of interest.
[0083] For example, for each pixel of interest, the video processor may perform a bilateral filtering window determination operation 702 to select a bilateral filtering window for the pixel of interest from a group of candidate filtering windows. The group of candidate filtering windows may include multiple side filtering windows and one full filtering window. The multiple side filtering windows may include, for example, an L window, an R window, a U window, a D window, an NW window, an SE window, a SW window, or an SE window for the pixel of interest, as shown in FIG. 6. The full filtering window may be, for example, a symmetric filtering window with the pixel of interest at the center of the window.
[0084] The video processor may then perform a bilateral filtering operation 704 to filter a pixel of interest of a video block in the video frame with the selected bilateral filtering window. For example, the video processor may apply equation (3) above to derive a filtered pixel of the pixel of interest based on the selected bilateral filtering window.
[0085] Hereinafter, this specification provides several exemplary embodiments of bilateral filtering with side windows 700 for deriving respective filtered pixels corresponding to each pixel of interest in a video frame. In the bilateral filtering disclosed herein, filtering of each pixel of interest may be a process of weighting all neighboring pixels within a bilateral filtering window. The bilateral filtering window may be one of the full filtering window or side filtering windows disclosed herein. As described in more detail below, one or more criteria may be employed to determine which filtering window is optimal for the pixel of interest.
[0086] In a first exemplary embodiment of bilateral filtering with side windows 700, the video processor may filter a pixel of interest with multiple side filtering windows to obtain multiple filtered pixel values, respectively. For example, for each side filtering window, the video processor may apply equation (3) above to derive a corresponding filtered pixel value for the pixel of interest based on the side filtering window. As a result, multiple filtered pixel values may be generated for the multiple side filtering windows, respectively.
[0087] The video processor may then calculate a plurality of differences corresponding to the plurality of filtered pixel values, each difference being a difference between a corresponding filtered pixel value and an original pixel value of the pixel of interest. The video processor may identify, from the plurality of filtered pixel values, the filtered pixel value associated with the smallest difference among the plurality of differences as the optimal filtered pixel value. The video processor may select a side filtering window corresponding to the optimal filtered pixel value as a bilateral filtering window for the pixel of interest. The video processor may output a filtered pixel for the pixel of interest using the selected bilateral filtering window.
[0088] For example, a first exemplary embodiment of bilateral filtering with side windows 700 can be described as follows. JPEG0007821822000006.jpg159167
[0089] Consistent with this disclosure, the first exemplary embodiment of bilateral filtering with side windows 700 follows the criterion that the optimal side filtering window results in an output filtered pixel that is closest to the original pixel. In the first exemplary embodiment, only the side filtering window is selected, and the full filtering window is not considered.
[0090] A second exemplary embodiment of bilateral filtering with side windows 700 can be formed by modifying the first exemplary embodiment in one or more aspects. In some embodiments, the first exemplary embodiment can be modified to incorporate a full filtering window into the set of candidate filtering windows. That is, the set of candidate filtering windows in the second exemplary embodiment can be modified to S={F, L, R, U, D, NW, NE, SW, SE} to include the full filtering window; JPEG0007821822000007.jpg78 represents the full filtering window. Accordingly, the side filtering window or the full filtering window can be selected for bilateral filtering.
[0091] In some embodiments, an additional judgment condition may be provided to improve the reliability of the selection of the side filtering window in the second exemplary embodiment. For example, the additional judgment condition may be added after step 2 of the first exemplary embodiment. As described in more detail below, if (a) the candidate filtering window corresponding to the optimal filtered pixel value in step 2 of the second exemplary embodiment is a side filtering window, and (b) the first difference between the original pixel value before filtering and the full-window filtered pixel value generated by the full filtering window is significantly greater than the second difference between the original pixel value and the optimal filtered pixel value generated by the side filtering window (e.g., first difference - second difference ≥ predetermined threshold th1), the side filtering window corresponding to the optimal filtered pixel value may be selected as the bilateral filtering window. Otherwise, the full filtering window may be selected as the bilateral filtering window.
[0092] Specifically, similar to the first exemplary embodiment, in the second exemplary embodiment, the video processor may filter the pixel of interest with multiple side filtering windows to obtain multiple filtered pixel values. For example, for each side filtering window, the video processor may apply Equation (3) above to derive a corresponding filtered pixel value of the pixel of interest based on the side filtering window. As a result, multiple filtered pixel values can be generated for the multiple side filtering windows, respectively.
[0093] The video processor may then calculate a plurality of differences corresponding to the plurality of filtered pixel values, each difference being a difference between a corresponding filtered pixel value and the original pixel value of the pixel of interest. From the plurality of filtered pixel values, the video processor may identify the filtered pixel value corresponding to the smallest difference among the plurality of differences as the optimal filtered pixel value.
[0094] The video processor may also filter the pixel of interest through a full filtering window to obtain a full-window filtered pixel value. The video processor may calculate a first difference between the full-window filtered pixel value of the pixel of interest and the original pixel value. The video processor may also calculate a second difference between the optimal filtered pixel value of the pixel of interest and the original pixel value. The video processor may determine whether the first difference and the second difference satisfy a predetermined condition. In some embodiments, the predetermined condition indicates that a third difference between the first difference and the second difference is less than a predetermined threshold th1. For example, if the third difference is less than the predetermined threshold th1, the predetermined condition is satisfied (e.g., third difference = first difference - second difference, and third difference < predetermined threshold th1).
[0095] If the first difference and the second difference satisfy a predetermined condition, the video processor may select the full filtering window as the bilateral filtering window. Otherwise, the video processor may select a side filtering window corresponding to an optimal filtered pixel value as the bilateral filtering window for the pixel of interest. The video processor may output a pixel filtered for the pixel of interest using the selected bilateral filtering window.
[0096] For example, the second exemplary embodiment can be described as follows. JPEG0007821822000008.jpg243169JPEG0007821822000009.jpg50166
[0097] A third exemplary embodiment of bilateral filtering with side windows 700 may include determining a bilateral filtering window for each pixel of interest based on an optimal pixel similarity value. For example, as described in more detail below, the average difference between the pixel of interest and the pixels in the bilateral filtering window may be the smallest of a set of average differences associated with a set of candidate filtering windows.
[0098] Specifically, the video processor may determine a set of pixel similarity values for each of the candidate filtering windows. For example, the pixel similarity value for each of the candidate filtering windows is between a pixel of interest and an adjacent pixel within the candidate filtering window for the pixel of interest. The pixel similarity value for each of the candidate filtering windows may be an average difference value determined between an original pixel value of the pixel of interest and an original pixel value of an adjacent pixel within the candidate filtering window for the pixel of interest. As a result, a set of average difference values is determined for each of the candidate filtering windows and used as the set of pixel similarity values. The video processor may determine an optimal pixel similarity value from the set of pixel similarity values. For example, the video processor may determine the optimal pixel similarity value as the smallest average difference value among the set of average difference values.
[0099] The video processor may select the candidate filtering window corresponding to the optimal pixel similarity value from the group of candidate filtering windows as the bilateral filtering window. Specifically, if the candidate filtering window corresponding to the optimal pixel similarity value is a side filtering window, the video processor may select the bilateral filtering window based on a comparison of the optimal pixel similarity value and the pixel similarity value determined for the full filtering window. For example, if the difference between the pixel similarity value of the full filtering window and the optimal pixel similarity value is less than a predetermined threshold, the video processor may select the bilateral filtering window based on a comparison of the optimal pixel similarity value and the pixel similarity value determined for the full filtering window. If the pixel similarity value of the full filtering window is less than or equal to JPEG0007821822000010.jpg714, the video processor may select the full filtering window as the bilateral filtering window. However, if the difference between the pixel similarity value of the full filtering window and the optimal pixel similarity value is less than or equal to a predetermined threshold If JPEG0007821822000011.jpg714 or greater, the video processor may select the side filtering window corresponding to the best pixel similarity value as the bilateral filtering window.
[0100] For example, the third exemplary embodiment can be described as follows. JPEG0007821822000012.jpg248168JPEG0007821822000013.jpg24164
[0101] In some application scenarios, direct application of side window filtering may fail for some video blocks due to diversity and non-stationarity in the video content. For example, only full window filtering may be applied to some video blocks, while side window filtering may be applied to other video blocks. To provide the flexibility to apply either full window filtering or side window filtering at the block level, the video processor may also perform filtered block selection 706 to select an appropriate output filtered block for each video block. For example, a first potential filtered block of a video block may be a full window filtered block in which all filtered pixels are generated by only the full filtering window. A second potential filtered block of a video block may be a side window filtered block in which all pixels are generated by the side window bilateral filtering 700 described above. The second potential filtered block may be any other appropriate filtered block and is not limited to the filtered block generated by the side window bilateral filtering 700. The output filtered block of the video block may be determined to be one of the first and second potential filtered blocks.
[0102] In some embodiments, filtered block selection 706 may be performed at different levels of granularity. Each video block may be a transform unit (TU), a coding unit (CU), or a coding tree unit (CTU). The video processor may flag each video block to indicate whether the video block is full-window filtered or side-window filtered. For example, the video processor may flag each video block to indicate whether the output filtered block of the video block is a full-window filtered block or a side-window filtered block. Filtered block selection 706 is also described in more detail below with reference to FIG. 12.
[0103] For example, since bilateral filtering may be performed at the TU level, the filtered block selection 706 may be performed at the TU level. Table 4 below provides syntax elements for filtered block selection at the TU level. The element "side_window_flag" may be signaled to each TU to indicate whether a full window filtered block or a side window filtered block is adopted as the output filtered block. For example, a first value of side_window_flag may indicate that a full window filtered block is adopted as the output filtered block. A second value of side_window_flag may indicate that a side window filtered block is adopted as the output filtered block. On the encoder side, rate-distortion optimization can be applied to determine whether a full window filtered block or a side window filtered block is adopted as the output filtered block. [Table 4] JPEG0007821822000015.jpg249170JPEG0007821822000016.jpg169165
[0104] In another embodiment, the filtered block selection 706 may be performed at the CU level. The side_window_flag may be signaled at the CU level. Table 5 below provides syntax elements for filtered block selection at the CU level. [Table 5] JPEG0007821822000018.jpg215167JPEG0007821822000019.jpg210170
[0105] In yet another embodiment, since selecting a filtered block at the TU or CU level may cause overhead, to avoid the overhead, the filtered block selection 706 may be performed at the CTU level. Table 6 below provides syntax elements for selecting a filtered block at the CTU level. If bilateral_filter_ctb_flag is true, side_window_flag is signaled to indicate whether a side window filtered block is employed. [Table 6]
[0106] In some embodiments, the video processor may also perform block-level filtering window selection 708. For example, the video processor may adaptively select a bilateral filtering window for each video block. For example, on the encoder side, the video processor may determine that each pixel of the video block can be filtered by a full filtering window. Alternatively, the video processor may determine that each pixel of the video block can be filtered by a particular side filtering window. A filter window index is signaled to indicate which filtering window is to be used for the video block. On the decoder side, the filtering window index can be parsed and bilateral filtering can be performed using the filtering window signaled by the filtering window index.
[0107] Consistent with this disclosure, block-level filtering window selection 708 may also be performed at different levels of granularity. Table 7 below provides syntax elements for filtering window selection at the TU level. For each TU, filter_window_index may be signaled to indicate the filtering window employed for the TU. [Table 7] JPEG0007821822000022.jpg249169JPEG0007821822000023.jpg167165
[0108] Table 8 below provides the syntax elements for filtering window selection at the CU level: filter_window_index is signaled to indicate which filtering window is adopted for the CU. [Table 8] JPEG0007821822000025.jpg227167JPEG0007821822000026.jpg226164JPEG0007821822000027.jpg85165
[0109] Table 9 below provides syntax elements for filtering window selection at the CTU level. If bilateral filtering is available for the CTU, filter_window_index may be signaled to indicate which filtering window is to be employed for the CTU. [Table 9]
[0110] Consistent with the present disclosure, a fixed-pattern side filtering window may be described by a filter window index, as illustrated by the following Table 10. When a fixed-pattern side filtering window is applied to video encoding / decoding, the syntax element "filter_window_index" may be employed to indicate the filtering window type of the fixed-pattern side filtering window. [Table 10]
[0111] 8 is a flow diagram of an example method 800 of bilateral filtering with side filtering windows according to some embodiments of the present disclosure. Method 800 may be performed by a video processor associated with video encoder 20 or video decoder 30 and may include steps 802-804, which are described below. Some steps may be optional in implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 8.
[0112] In step 802, a video processor may receive video frames of a video for in-loop filtering.
[0113] In step 804, for a pixel of interest in the video frame, the video processor may select a bilateral filtering window from a group of candidate filtering windows for in-loop filtering of the pixel of interest. The group of candidate filtering windows may include multiple side filtering windows and one full filtering window.
[0114] In step 806, the video processor may filter the pixel of interest of the video frame with the selected bilateral filtering window.
[0115] 9 is a flow diagram of a method 900, which is a first exemplary embodiment of the method 800 in FIG. 8 for bilateral filtering with side filtering windows, according to some embodiments of the present disclosure. Method 900 may be implemented by a video processor associated with video encoder 20 or video decoder 30 and may include steps 902-912, described below. Specifically, steps 904-910 of method 800 may be performed as an exemplary embodiment of step 804 of method 800. Some steps may be optional for implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 9.
[0116] In step 902, a video processor may receive video frames of a video for in-loop filtering.
[0117] In step 904, the video processor may filter a pixel of interest in the video frame with multiple side filtering windows to obtain multiple filtered pixel values, respectively. The implementation of step 904 is described above in relation to Figure 7. For example, step 904 may be performed according to step 1 in the above first exemplary embodiment.
[0118] In step 906, the video processor may calculate the difference between each filtered pixel value and the original pixel value of the pixel of interest to generate a plurality of differences for the plurality of filtered pixel values, respectively.
[0119] In step 908, the video processor may identify, from the plurality of filtered pixel values, the filtered pixel value associated with the smallest difference among the plurality of differences as the optimal filtered pixel value.
[0120] The implementation of steps 906 and 908 is described above in relation to Figure 7. For example, steps 906 and 908 may be performed according to step 2 in the first exemplary embodiment above.
[0121] In step 910, the video processor may select as the bilateral filtering window the side filtering window corresponding to the best filtered pixel value from the plurality of filtered pixel values.
[0122] In step 912, the video processor may filter the pixel of interest of the video frame with the selected bilateral filtering window.
[0123] The implementation of steps 910 and 912 is described above in relation to Figure 7. For example, steps 910 and 912 may be performed according to the return step in the first exemplary embodiment above.
[0124] 10 is a flow diagram of a method 1000, which is a second exemplary embodiment of the method 800 for bilateral filtering with side filtering windows in FIG. 8 , according to some embodiments of the present disclosure. The method 1000 may be implemented by a video processor associated with the video encoder 20 or the video decoder 30 and may include steps 1002-1018 described below. Specifically, steps 1004-1016 of the method 1000 may be performed as an exemplary embodiment of step 804 of the method 800. Some steps may be optional for implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 10 .
[0125] In step 1002, a video processor may receive video frames of a video for in-loop filtering.
[0126] In step 1004, the video processor may filter a pixel of interest in the video frame with multiple side filtering windows to obtain multiple filtered pixel values, respectively. The implementation of step 1004 is described above in relation to Figure 7. For example, step 1004 may be performed according to step 1 in the above second exemplary embodiment.
[0127] In step 1006, the video processor may calculate the difference between each filtered pixel value and the original pixel value of the pixel of interest to generate a plurality of differences for the plurality of filtered pixel values, respectively.
[0128] In step 1008, the video processor may identify, from the plurality of filtered pixel values, the filtered pixel value associated with the smallest difference among the plurality of differences as the optimal filtered pixel value.
[0129] The implementation of steps 1006 and 1008 is described above in relation to Figure 7. For example, steps 1006 and 1008 may be performed according to step 2 in the second exemplary embodiment above.
[0130] In step 1010, the video processor may calculate (a) a first difference between the full window filtered pixel value of the pixel of interest and the original pixel value, and (b) a second difference between the optimal filtered pixel value of the pixel of interest and the original pixel value.
[0131] In step 1012, the video processor may determine whether the first difference and the second difference satisfy a predetermined condition. If the first difference and the second difference satisfy the predetermined condition, method 1000 may proceed to step 1014. Otherwise, method 1000 may proceed to step 1016. In step 1014, the video processor may select the full filtering window as the bilateral filtering window.
[0132] In step 1016, the video processor may select the side filtering window corresponding to the best filtered pixel value as the bilateral filtering window.
[0133] The implementation of steps 1010-1016 is described above in relation to Figure 7. For example, steps 1010-1016 may be performed according to step 3 in the above second exemplary embodiment.
[0134] In step 1018, the video processor may filter the pixel of interest of the video frame with the selected bilateral filtering window.
[0135] 11 is a flow diagram of a method 1100, which is a third exemplary embodiment of method 800 for bilateral filtering with side filtering windows in FIG. 8, according to some embodiments of the present disclosure. Method 1100 may be implemented by a video processor associated with video encoder 20 or video decoder 30 and may include steps 1102-1116 described below. Specifically, steps 1104-1114 of method 1100 may be performed as an exemplary embodiment of step 804 of method 800. Some steps may be optional for implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 11.
[0136] In step 1102, a video processor may receive a video frame of a video for in-loop filtering.
[0137] In step 1104, the video processor may determine, for a pixel of interest in the video frame, a pixel similarity value for each candidate filtering window between the pixel of interest and neighboring pixels within the candidate filtering window of the pixel of interest. Consequently, a set of pixel similarity values is determined for each set of candidate filtering windows. The set of candidate filtering windows may include multiple side filtering windows and one full filtering window. The implementation of step 1104 is described above in connection with FIG. 7 . For example, step 1104 may be performed according to step 1 in the third exemplary embodiment.
[0138] In step 1106, the video processor may determine a candidate filtering window corresponding to an optimal pixel similarity value from the group of pixel similarity values. The implementation of step 1106 is described above in relation to Figure 7. For example, step 1106 may be performed according to step 2 in the above third exemplary embodiment.
[0139] In step 1108, the video processor may determine whether the candidate filtering window corresponding to the best pixel similarity value is a side filtering window. If the candidate filtering window corresponding to the best pixel similarity value is a side filtering window, method 1100 may proceed to step 1110. Otherwise, method 1100 may proceed to step 1112.
[0140] In step 1110, the video processor may determine whether the difference between the pixel similarity values of the full filtering window and the optimal pixel similarity values is less than or equal to a predetermined threshold. If the difference between the pixel similarity values of the full filtering window and the optimal pixel similarity values is less than or equal to the predetermined threshold, method 1100 may proceed to step 1112. Otherwise, method 1100 may proceed to step 1114.
[0141] In step 1112, the video processor may select the full filtering window as the bilateral filtering window.
[0142] In step 1114, the video processor may select the side filtering window corresponding to the best pixel similarity value as the bilateral filtering window.
[0143] The implementation of steps 1108-1114 is described above in relation to Figure 7. For example, steps 1108-1114 may be performed according to step 3 in the above third exemplary embodiment.
[0144] In step 1116, the video processor may filter the pixel of interest of the video frame with the selected bilateral filtering window.
[0145] 12 is a flow diagram of an example method 1200 for selecting a filtered block for a video block according to some embodiments of the present disclosure. Method 1200 may be performed by a video processor associated with video encoder 20 or video decoder 30 and may include steps 1202-1212, described below. Some steps may be optional in implementing the disclosure provided herein. Furthermore, some steps may be performed simultaneously or in a different order than that shown in FIG. 12.
[0146] In step 1202, the video processor may perform bilateral filtering with a side window for each pixel of interest from a plurality of pixels of interest of the video block to generate a corresponding first filtered pixel, thereby generating a plurality of first filtered pixels for the plurality of pixels of interest of the video block. For example, the video processor may perform method 800, 900, 1000, or 1100 for each pixel of interest of the video block to filter the pixel of interest, thereby generating a corresponding first filtered pixel for the pixel of interest.
[0147] In step 1204, the video processor may generate a side window filtered block including a plurality of the first filtered pixels.
[0148] In step 1206, the video processor may perform bilateral filtering for each pixel of interest from the plurality of pixels of interest with the pixel of interest's respective full filtering window to generate a corresponding second filtered pixel, thereby generating a plurality of second filtered pixels for the plurality of pixels of interest of the video block.
[0149] In step 1208, the video processor may generate a full window filtered block including a plurality of second filtered pixels.
[0150] In step 1210, the video processor may select a side window filtered block or a full window filtered block as the optimal filtered block of the video block. For example, the video processor may select a side window filtered block or a full window filtered block as the optimal filtered block based on rate-distortion optimization.
[0151] In step 1212, the video processor may flag the video block to indicate whether the video block is full-window filtered or side-window filtered based on the optimal filtered block of the video block. For example, if the optimal filtered block is a side-window filtered block, the video processor may flag the video block to be side-window filtered. If the optimal filtered block is a full-window filtered block, the video processor may flag the video block to be full-window filtered.
[0152] 13 illustrates a computing environment 1310 coupled to a user interface 1350, according to some embodiments of the present disclosure. The computing environment 1310 may be part of a data processing server. The computing environment 1310 may include a processor 1320, a memory 1330, and an input / output (I / O) interface 1340.
[0153] The processor 1320 typically controls the overall operation of the computing environment 1310, such as operations related to display, data acquisition, data communication, and image processing. The processor 1320 may include one or more processors for executing instructions to perform all or a portion of the steps of the methods described above. The processor 1320 may also include one or more modules that facilitate interaction between the processor 1320 and other components. The processor 1320 may be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphics processing unit (GPU), etc.
[0154] Memory 1330 is configured to store various types of data to support the operation of computing environment 1310. Memory 1330 may include predetermined software 1332. Examples of such data include instructions for any applications or methods operating on computing environment 1310, video data sets, image data, etc. Memory 1330 may be implemented by employing any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk, or a combination thereof.
[0155] The I / O interface 1340 provides an interface between the processor 1320 and a peripheral interface module, such as a keyboard, a click wheel, buttons, etc. The buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. The I / O interface 1340 may be coupled to an encoder and a decoder.
[0156] In some embodiments, a non-transitory computer-readable storage medium is also provided, including a plurality of programs executable by the processor 1320 in the computing environment 1310 to perform the methods described above, such as contained in the memory 1330. Alternatively, the non-transitory computer-readable storage medium may store a bitstream or datastream, the bitstream or datastream including coded video information (e.g., video information including one or more syntax elements) generated by, for example, an encoder (e.g., video encoder 20 in FIG. 2 ) using the coding method described above, which is employed by a decoder (e.g., video decoder 30 in FIG. 3 ) to decode the video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0157] In some embodiments, a computing device is also provided that includes one or more processors (e.g., processor 1320) and a non-transitory computer-readable storage medium or memory 1330 having stored thereon a plurality of programs executable by the one or more processors, the one or more processors being configured to perform the above-described methods when executing the plurality of programs.
[0158] In some embodiments, a computer program product is also provided that includes a plurality of programs executable by the processor 1320 in the computing environment 1310 to perform the methods described above, such as contained in memory 1330. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0159] In some embodiments, the computing environment 1310 may be implemented with one or more ASICs, DSPs, digital signal processors (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0160] The description of the present disclosure has been presented for purposes of illustration and is not intended to be exhaustive or limiting of the disclosure. Many modifications, variations and alternative embodiments will become apparent to one skilled in the art having the benefit of the teachings presented in the foregoing descriptions and the associated drawings.
[0161] Unless otherwise specified, the order of steps in the method according to the present disclosure is intended as an example only, and the steps in the method according to the present disclosure are not limited to the order specifically described above, and may be changed according to actual conditions. In addition, at least one step of the steps in the method according to the present disclosure may be adjusted, combined, or deleted according to actual needs.
[0162] The examples have been chosen and described to explain the principles of the disclosure and to enable others skilled in the art to understand the disclosure in various embodiments and to best utilize the underlying principles and various implementations with various modifications suitable for the particular applications intended. Therefore, it is to be understood that the scope of the disclosure should not be limited to the specific examples of the disclosed embodiments, and that modifications and other embodiments are intended to be included within the scope of the disclosure.
Claims
1. receiving, by one or more processors, video frames of a video for in-loop filtering; selecting, by the one or more processors, a bilateral filtering window for in-loop filtering of a pixel of interest of the video frame from a group of candidate filtering windows including a plurality of side filtering windows and one full filtering window; filtering the pixel of interest of the video frame with the selected bilateral filtering window by the one or more processors; A video processing method with bilateral filtering, comprising:
2. Selecting a bilateral filtering window for the pixel of interest includes: filtering the pixel of interest with the plurality of side filtering windows to obtain a plurality of filtered pixel values, respectively; selecting a side filtering window corresponding to an optimal filtered pixel value among the plurality of filtered pixel values as the bilateral filtering window; The video processing method of claim 1 further comprising:
3. Selecting a side filtering window corresponding to the optimal filtered pixel value comprises: calculating a plurality of differences corresponding to the plurality of filtered pixel values, each difference being a difference between a corresponding filtered pixel value of the pixel of interest and an original pixel value; identifying, from the plurality of filtered pixel values, a filtered pixel value associated with a smallest difference among the plurality of differences as the optimal filtered pixel value; The video processing method of claim 2 further comprising:
4. Selecting a bilateral filtering window for the pixel of interest includes: filtering the pixel of interest with the full filtering window to obtain a full window filtered pixel value; Calculating a first difference between the full-window filtered pixel value of the pixel of interest and the original pixel value; calculating a second difference between the optimal filtered pixel value and the original pixel value of the pixel of interest; selecting the full filtering window as the bilateral filtering window if the first difference and the second difference satisfy a predetermined condition; The video processing method of claim 2 further comprising:
5. 5. The video processing method of claim 4, wherein the predetermined condition is that a third difference between the first difference and the second difference is lower than a predetermined threshold.
6. Selecting a bilateral filtering window for the pixel of interest includes: determining pixel similarity values for a set of candidate filtering windows, each pixel similarity value being a pixel similarity value between the pixel of interest and a neighboring pixel of the pixel of interest that is within the candidate filtering window; selecting, from the group of candidate filtering windows, a candidate filtering window corresponding to an optimal pixel similarity value from the group of pixel similarity values as the bilateral filtering window; The video processing method of claim 1 further comprising:
7. 7. The video processing method of claim 6, wherein the pixel similarity value for each candidate filtering window is an average value of differences determined between the pixel value of the pixel of interest and the pixel values of its neighboring pixels that are within the candidate filtering window.
8. Selecting a bilateral filtering window for the pixel of interest includes: selecting the bilateral filtering window based on a comparison of the optimal pixel similarity value and a pixel similarity value determined for the full filtering window, if the candidate filtering window corresponding to the optimal pixel similarity value is a side filtering window; The video processing method of claim 6 further comprising:
9. Selecting the bilateral filtering window based on a comparison of the optimal pixel similarity value and a pixel similarity value determined for the full filtering window comprises: selecting the full filtering window as the bilateral filtering window if a difference between the pixel similarity value of the full filtering window and the optimal pixel similarity value is less than or equal to a predetermined threshold. The video processing method of claim 8 further comprising:
10. the video frame includes a plurality of video blocks; The video processing method includes: flagging each video block to indicate whether the video block is full-window filtered or side-window filtered; The video processing method of claim 1 further comprising:
11. The video processing method of claim 10 , wherein each video block is a transform unit (TU), a coding unit (CU), or a coding tree unit (CTU).
12. 2. The video processing method of claim 1, wherein each side filtering window has a filtering window type indicated by a filtering window index.
13. one or more processors; a memory configured to store instructions executable by one or more processors; Including, A video processing apparatus for bilateral filtering, wherein the one or more processors are configured to perform the method of any one of claims 1 to 12 upon execution of the instructions.
14. A computer program having instructions that, when executed by one or more processors, cause the one or more processors to perform a method according to any one of claims 1 to 12.
15. A method for storing a bitstream, comprising: generating a bitstream by executing a video processing method according to any one of claims 1 to 12; storing the generated bitstream on a storage medium.
Citation Information
Patent Citations
Video coding and decoding based on image refinement
JP2014534744A
Filtering and edge encoding
JP2016146655A
Method for deblocking block of video sample
JP2019071632A
Non-local bilateral filter
US20190082176A1
Dynamic image encoding / decoding device
WO2009110559A1