Video processing method and related video processing device for disabling sample adaptive offset filtering across virtual boundaries in a reconstructed frame

By disabling SAO filtering on the virtual boundaries of reconstructed frames in omnidirectional video coding, the visual quality and coding efficiency problems caused by image content discontinuity in omnidirectional video coding are solved, achieving higher visual quality and coding efficiency.

CN114731432BActive Publication Date: 2025-12-05MEDIATEK INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080080035.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-25
Filing Date
2020-12-31
Publication Date
2025-12-05
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

In omnidirectional video coding, the discontinuity of image content at the edges of faces in projection-based frames leads to poor visual quality and reduced coding efficiency. Existing loop filters such as SAO filters may exacerbate this problem.

Method used

Sample Adaptive Offset (SAO) filtering is disabled on the virtual boundaries of the reconstructed frame. The SAO filter is instructed by the control circuit to keep the current pixel value unchanged at the virtual boundaries. An innovative SAO filter design is adopted to avoid filtering operations across virtual boundaries.

Benefits of technology

It improves the visual quality and coding efficiency of reconstructed frames, reduces image discontinuity issues caused by virtual boundaries, and enhances the overall performance of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114731432B_ABST
    Figure CN114731432B_ABST
Patent Text Reader

Abstract

A method of video processing includes receiving a reconstructed frame and applying in-loop filtering to the reconstructed frame by at least one in-loop filter. The step of applying in-loop filtering includes performing a sample adaptive offset (SAO) filtering operation. The step of performing the SAO filtering operation includes preserving a value of a current pixel included in the reconstructed frame by preventing the SAO filtering operation of the current pixel from crossing a virtual boundary defined in the reconstructed frame.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing

[0002] This application claims priority to U.S. Provisional Application No. 62 / 956,680, filed January 3, 2020, the entire contents of which are incorporated herein by reference. [Technical Field]

[0003] The present invention relates to processing reconstructed frames generated during video encoding or video decoding, and more specifically, to a video processing method and associated video processing apparatus for disabling sample adaptive offset (SAO) filtering across virtual boundaries in reconstructed frames. [Background Technology]

[0004] Virtual reality (VR) with head-mounted displays (HMDs) is associated with a variety of applications. The ability to display wide-field-of-view content to users can be used to provide an immersive visual experience. The real-world environment must be captured in all directions to produce omnidirectional image content corresponding to a sphere. With advancements in camera devices and HMDs, the delivery of VR content may soon become a bottleneck due to the high bitrate required to render such 360-degree image content. When the resolution of omnidirectional video is 4K or higher, data compression / encoding is crucial for reducing the bitrate.

[0005] Omnidirectional video data compression / encoding can be achieved using conventional video codec standards that typically employ block-based encoding / decoding techniques to leverage spatial and temporal redundancy. For example, the basic approach is to divide the source frame into multiple blocks (or codec units), perform intra-frame prediction / inter-frame prediction on each block, transform the residuals of each block, and then quantize and entropy encode them. Furthermore, reconstructed frames are generated to provide reference pixel data for encoding and decoding subsequent blocks. For some video codec standards, loop filters can be used to enhance the image quality of the reconstructed frames. For example, video encoders use Sample Adaptive Shift (SAO) filters to minimize the average sample distortion in regions. Video decoders are used to perform the inverse operations of the video encoding operations performed by the video encoder. Therefore, video decoders also have loop filters for enhancing the image quality of the reconstructed frames. For example, video decoders also use SAO filters to reduce distortion.

[0006] Generally, the omnidirectional video content corresponding to a sphere is converted into an image sequence, where each image is a projection-based frame. The 360-degree image content is represented by one or more projection surfaces arranged in a 360-degree virtual reality (360VR) projection layout. The projection-based frame sequence is then encoded into a bitstream for transmission. However, projection-based frames may have discontinuities in image content at facial edges (i.e., facial boundaries). Applying loop filtering (e.g., SAO filtering) to these discontinuous facial edges can lead to poor visual quality and reduced coding efficiency. [Summary of the Invention]

[0007] One object of the claimed invention is to provide a video processing method and related video processing apparatus that disables sample adaptive offset (SAO) filtering on virtual boundaries in reconstructed frames.

[0008] According to a first aspect of the present invention, an exemplary video processing method is disclosed. The exemplary video processing method includes: receiving a reconstructed frame, and applying loop filtering to the reconstructed frame through at least one loop filter, including performing a Sample Adaptive Offset (SAO) filtering operation. The step of performing the SAO filtering operation includes: maintaining the value of the current pixel unchanged by preventing the SAO filtering operation of the current pixel included in the reconstructed frame from crossing a virtual boundary defined in the reconstructed frame.

[0009] According to a second aspect of the present invention, an exemplary video processing apparatus is disclosed. The exemplary video processing apparatus includes an encoding circuit arranged to receive a video frame and encode the video frame to generate a portion of a bitstream. During encoding of the video frame, the encoding circuit generates a reconstructed frame and applies a loop filter to the reconstructed frame. The loop filter includes a Sample Adaptive Offset (SAO) filtering operation performed at a SAO filter included in the encoding circuit. The SAO filter is arranged to maintain the value of the current pixel unchanged by preventing the SAO filtering operation of the current pixel included in the reconstructed frame from being applied across a virtual boundary defined in the reconstructed frame.

[0010] According to a third aspect of the present invention, an exemplary video processing apparatus is disclosed. The exemplary video processing apparatus includes a decoding circuit arranged to receive a bitstream and decode a portion of the bitstream, wherein the portion of the bitstream includes encoded information of video frames. During decoding of the portion of the bitstream, the decoding circuit is configured to generate a reconstructed frame and apply a loop filter to the reconstructed frame. The loop filter includes a Sample Adaptive Offset (SAO) filtering operation performed at a SAO filter included in the decoding circuit. The SAO filter is arranged to maintain the value of the current pixel unchanged by preventing the application of the SAO filtering operation of the current pixel included in the reconstructed frame across a virtual boundary defined in the reconstructed frame.

[0011] These and other objects of the invention will undoubtedly become apparent to those skilled in the art after reading the following detailed description of the preferred embodiments shown in the various accompanying drawings. [Attached Image Description]

[0012] Figure 1 This is a diagram illustrating a 360-degree virtual reality (360VR) system according to an embodiment of the present invention.

[0013] Figure 2 This is a diagram illustrating a video encoder according to an embodiment of the present invention.

[0014] Figure 3 This is a diagram illustrating a video decoder according to an embodiment of the present invention.

[0015] Figure 4 This is a diagram illustrating a cube-based projection according to an embodiment of the present invention.

[0016] Figure 5 This is a diagram illustrating the horizontal style (SAO EO class == 0) used for pixel classification in EO mode.

[0017] Figure 6 This diagram illustrates the vertical style (SAO EO class == 1) used for pixel classification in EO mode.

[0018] Figure 7 This diagram illustrates the 135-degree diagonal pattern (SAO EO class == 2) used for pixel classification in EO mode.

[0019] Figure 8 This is a diagram illustrating the 45-degree diagonal style (SAO EO class == 3) used for pixel classification in EO mode.

[0020] Figure 9 This is a flowchart illustrating a video processing method according to an embodiment of the present invention.

Detailed Implementation Methods

[0021] Certain terms used in the following description and claims refer to specific components. As those skilled in the art will understand, electronic device manufacturers may use different names to refer to a component. This document is not intended to distinguish between components with different names but the same function. In the following description and claims, the terms "comprising" and "including" are used in an open-ended manner and should therefore be interpreted as meaning "including but not limited to...". Furthermore, the term "coupled" is intended to indicate an indirect or direct electrical connection. Thus, if one device is coupled to another device, the connection can be through a direct electrical connection or through an indirect electrical connection via other devices and connections.

[0022] Virtual boundaries can be defined by different applications or needs. Taking 360-degree video as an example, the layout of a specific projection format may have one or more discontinuous boundaries between adjacent projection planes encapsulated in a projection-based frame, where these discontinuous boundaries can be defined as virtual boundaries. Applying loop filtering (e.g., SAO filtering) to these discontinuous boundaries (e.g., virtual boundaries) can lead to poor visual quality and reduced coding efficiency. To address this issue, this invention proposes an innovative SAO filter design that allows disabling SAO filtering across discontinuous boundaries (e.g., virtual boundaries) in edge offset mode. Further details of the proposed SAO filter design are described with reference to the accompanying drawings.

[0023] To better understand the technical features of the proposed SAO filter design, it is assumed that a video encoder using the proposed SAO filter design is configured to encode projection-based frames into a bitstream, and a video decoder uses the proposed SAO filter design to decode the bitstream to generate decoded projection-based frames. However, this is for illustrative purposes only and does not imply limitation of the invention. In practice, any video processing apparatus that uses the proposed SAO filter design to process reconstructed frames with one or more virtual boundaries (which can be either projection-based frames with one or more virtual boundaries or non-projection-based frames with one or more virtual boundaries) falls within the scope of this invention.

[0024] Figure 1This diagram illustrates a 360-degree virtual reality (360VR) system according to an embodiment of the present invention. The 360VR system 100 includes two video processing devices (e.g., source electronics 102 and destination electronics 104). Source electronics 102 includes a video capture device 112, a conversion circuit 114, and a video encoder 116. For example, the video capture device 112 may be a set of cameras for providing omnidirectional content (e.g., multiple images covering the entire environment) S_IN corresponding to a sphere. The conversion circuit 114 is coupled between the video capture device 112 and the video encoder 116. The conversion circuit 114 generates a projection-based frame IMG with a 360-degree virtual reality (360VR) projection layout L_VR based on the omnidirectional content S_IN. For example, the projection-based frame IMG may be one frame included in a sequence of projection-based frames generated from the conversion circuit 114. The video encoder 116 is designed to encode / compress the projection-based frame IMG to generate a portion of a bitstream BS, and output the bitstream BS to the destination electronics 104 via a transmission device 103. For example, a projection-based frame sequence can be encoded into a bitstream BS, such that a portion of the bitstream BS transmits the encoded information of the projection-based frame IMG. Furthermore, the transmission device 103 can be a wired / wireless communication link or a storage medium.

[0025] The intended electronic device 104 may be a head-mounted display (HMD) device. For example... Figure 1 As shown, the target electronic device 104 includes a video decoder 122, a graphics rendering circuit 124, and a display screen 126. The video decoder 122 is designed to receive a bitstream BS from a transmission device 103 (e.g., a wired / wireless communication link or storage medium) and decode the received bitstream BS. For example, the video decoder 122 generates a sequence of decoded frames by decoding the received bitstream BS, where a decoded frame IMG' is a frame included in the sequence of decoded frames. That is, since a portion of the bitstream BS transmits encoded information of a projection-based frame IMG, the video decoder 122 decodes a portion of the received bitstream BS to generate a decoded frame IMG', which is the result of decoding the encoded information of the projection-based frame IMG. In this embodiment, the projection-based frame IMG to be encoded by the video encoder 116 has a 360VR projection format that supports the projection layout. Therefore, after the video decoder 122 decodes the bitstream BS, the decoded frame IMG' has the same 360VR projection format and the same projection layout. A graphics rendering circuit 124 is coupled between the video decoder 122 and the display screen 126. The graphics rendering circuit 124 renders and outputs image data based on the decoded frame IMG' and displays it on the display screen 126. For example, a viewport area associated with a portion of the 360-degree content carried by the decoded frame IMG' can be displayed on the display screen 126 through the graphics rendering circuit 124.

[0026] This invention proposes techniques at the encoding / decoding tool level to overcome the negative impacts introduced by image content discontinuities (i.e., discontinuous surface edges) caused by the packing of projection surfaces. In other words, video encoder 116 can use the proposed encoding / decoding tool to encode projection-based frame IMGs, and the corresponding video decoder 122 can also use the proposed encoding / decoding tool to generate decoded frame IMGs. For example, video encoder 116 uses the proposed SAO filter for loop filtering, and video decoder 122 also uses the proposed SAO filter for loop filtering.

[0027] Figure 2 This is a diagram illustrating a video encoder according to an embodiment of the present invention. Figure 1 The video encoder 116 shown can be used Figure 2 The video encoder 200 shown is used to implement this. The video encoder 200 includes a control circuit 202 and an encoding circuit 204. It should be noted that... Figure 2 The video encoder architecture shown is for illustrative purposes only and is not intended to limit the invention. For example, the architecture of the encoding circuit 204 can vary depending on the encoding / decoding standard. The encoding circuit 204 encodes video frames (e.g., projection-based frame IMGs, which have 360-degree content represented by projection surfaces arranged in a 360VR projection layout L_VR) to generate a portion of the bitstream BS. Figure 2 As shown, the encoding circuit 204 includes a residual calculation circuit 211, a transform circuit (denoted by "T") 212, a quantization circuit (denoted by "Q") 213, an entropy encoding circuit (e.g., a variable length encoder) 214, an inverse quantization circuit (denoted by "IQ") 215, an inverse transform circuit (denoted by "IT") 216, a reconstruction circuit 217, at least one loop filter 218, a reference frame buffer 219, an inter-frame prediction circuit 220 (which includes a motion estimation circuit (denoted by "ME") 221 and a motion compensation circuit (denoted by "MC") 222), an intra-frame prediction circuit (denoted by "IP") 223, and an intra-frame / inter-frame mode selection switch 224. The loop filter 218 may include a deblocking filter, a sample adaptive offset (SAO) filter, and / or an adaptive loop filter (ALF).

[0028] It is worth noting that the reconstructed frame IMG_R generated by the reconstruction circuit 217 is stored in the reference frame buffer 219 as a reference frame after being processed by the loop filter 218. The reconstructed frame IMG_R can be viewed as a decoded version of the projection-based frame IMG. Therefore, the reconstructed frame IMG_R also has 360-degree image content represented by projection surfaces arranged in the same 360VR projection layout L_VR. In this embodiment, the reconstructed frame IMG_R is received by the loop filter 218, and the SAO filter 226 (represented by "SAO") is coupled between the reconstruction circuit 217 and the reference buffer 219. That is, the loop filtering applied to the reconstructed frame IMG_R includes SAO filtering.

[0029] The main difference between encoding circuit 204 and typical encoding circuits is that SAO filter 226 can be instructed by control circuit 202 to enable the proposed function, namely, preventing SAO filtering from being applied across virtual boundaries (e.g., discontinuous boundaries generated by projection plane packing). For example, control circuit 202 generates a control signal C1 to enable SAO filter 226. Furthermore, control circuit 202 can be further configured to set one or more syntax elements (SEs) associated with enabling / disabling the proposed function at SAO filter 226, wherein syntax elements SEs are signaled to the video decoder via a bitstream BS generated from entropy encoding circuit 214.

[0030] Figure 3 This is a diagram illustrating a video decoder according to an embodiment of the present invention. It can be used... Figure 3 The video decoder 300 shown implements this. Figure 1 The video decoder 122 is shown. The video decoder 300 can be connected to the video encoder (e.g., via a transmission device such as a wired / wireless communication link or storage medium). Figure 1 The video encoder 100 shown or Figure 2 The video encoder 200 shown communicates with the video decoder 300. In this embodiment, the video decoder 300 receives a bitstream BS and decodes a portion of the received bitstream BS to generate a decoded frame IMG'. (Refer to...) Figure 3 The video decoder 300 includes a decoding circuit 320 and a control circuit 330. It should be noted that... Figure 3The video decoder architecture shown is for illustrative purposes only and is not intended to limit the invention. For example, the architecture of the decoding circuit 320 may vary depending on the encoding / decoding standard. The decoding circuit 320 includes an entropy decoding circuit (e.g., a variable-length decoder) 302, an inverse quantization circuit (referred to as "IQ") 304, an inverse transform circuit (referred to as "IT") 306, a reconstruction circuit 308, an inter-frame prediction circuit 312 (which includes a motion vector calculation circuit (referred to as "MV calculation") 310 and a motion compensation circuit (referred to as "MC") 313), an intra-frame prediction circuit (referred to as "IP") 314, an intra / inter-frame mode selection switch 316, at least one loop filter (e.g., a deblocking filter, a SAO filter, and / or an ALF) 318, and a reference frame buffer 319. In this embodiment, the projection-based frame IMG encoded by the video encoder 100 has 360-degree content represented by a projection surface arranged in a 360VR projection layout L_VR. Therefore, after the video decoder 300 decodes the bitstream BS, the decoded frame IMG' also has 360-degree image content represented by projection surfaces arranged in the same 360VR projection layout L_VR. The reconstructed frame IMG_R' generated by the reconstruction circuit 308 is stored in the reference frame buffer 319 as a reference frame and is also used as the decoded frame IMG' after being processed by one or more loop filters 318. Therefore, the reconstructed frame IMG_R' also has 360-degree image content represented by projection surfaces arranged in the same 360VR projection layout L_VR. In this embodiment, the reconstructed frame IMG_R' is received by the loop filter 318, and the SAO filter 322 (represented by "SAO") is coupled between the reconstruction circuit 308 and the reference buffer 319. That is, the loop filtering applied to the reconstructed frame IMG_R' includes SAO filtering.

[0031] The main difference between decoding circuit 320 and typical decoding circuits is that SAO filter 322 can be instructed by control circuit 330 to enable the proposed function that prevents SAO filtering across virtual boundaries (e.g., discontinuous boundaries generated by projection surface stacking). For example, control circuit 330 generates a control signal C1' to enable the proposed function at SAO filter 322. Furthermore, entropy decoding circuit 302 is further used to process the bitstream BS to obtain one or more syntax elements SE, which are related to enabling / disabling the proposed function at SAO filter 322. Therefore, control circuit 330 of video decoder 300 can refer to the parsed syntax element SE to determine whether to enable the proposed function at SAO filter 322.

[0032] In this invention, the 360VR projection layout L_VR can be any available projection layout. For example, the 360VR projection layout L_VR can be a cube-based projection layout. In practice, the proposed encoding / decoding tools at SAO filters 226 / 322 can be used to process 360VR frames with projection surfaces packed in other projection layouts.

[0033] Figure 4 This diagram illustrates a cube-based projection according to an embodiment of the present invention. The 360-degree content on the sphere 400 is projected onto the six faces of a cube 401 in three-dimensional (3D) space, including the top, bottom, left, front, right, and back faces. Specifically, the image content of the north polar region of the sphere 400 is projected onto the top face of the cube 401, the image content of the south polar region of the sphere 400 is projected onto the bottom face of the cube 401, and the image content of the equatorial region of the sphere 400 is projected onto the left, front, right, and back faces of the cube 401.

[0034] The square projection surfaces encapsulated in the cube-based projection layout are derived from the six faces of cube 401. For example, the square projection surface on the two-dimensional (2D) plane (labeled "Top") is derived from the top face of cube 401 in 3D space, the square projection surface on the 2D plane (labeled "Back") is derived from the back face of cube 401 in 3D space, the square projection surface on the 2D plane (labeled "Bottom") is derived from the bottom face of cube 401 in 3D space, the square projection surface on the 2D plane (labeled "Right") is derived from the right face of cube 401 in 3D space, the square projection surface on the 2D plane (labeled "Front") is derived from the front face of cube 401 in 3D space, and the square projection surface on the 2D plane (labeled "Left") is derived from the left face of cube 401 in 3D space.

[0035] When 360VR projection layout L_VR is Figure 4When the cubemap projection (CMP) layout 402 is set up as shown, the square projection faces "top," "back," "bottom," "right," "front," and "left" are encapsulated within the CMP layout 402 corresponding to the unfolded cube. However, the projection-based frame IMG to be encoded needs to be rectangular. If the CMP layout 402 is used directly to create the projection-based frame IMG, the projection-based frame IMG must be filled with dummy areas (e.g., black, gray, or white areas) to form a rectangular frame for encoding. Alternatively, the projection-based frame IMG can have projected image data arranged in a compact projection layout to avoid using dummy areas (e.g., black, gray, or white areas). Figure 4 As shown, the square projection surfaces "top", "back", and "bottom" are rotated and then encapsulated in a compact CMP layout 404. Therefore, the square projection surfaces "top", "back", "bottom", "right", "front", and "left" are arranged in a compact CMP layout 404, i.e., a 3x2 layout. This improves encoding and decoding efficiency.

[0036] However, according to the compact CMP layout 404, the packaging of square projection surfaces may result in discontinuous edges of image content between adjacent square projection surfaces. For example... Figure 4As shown, a projection-based frame IMG with a compact CMP layout 404 has a top sub-frame (which is a 3x1 face row consisting of square projection planes "right", "front", and "left") and a bottom sub-frame (which is another 3x1 face row consisting of square projection planes "bottom", "back", and "top"). There are image content discontinuities between the top and bottom sub-frames. Specifically, the face edge S1 of the square projection plane "right" is connected to the face edge S6 of the square projection plane "bottom", the face edge S2 of the square projection plane "front" is connected to the face edge S5 of the square projection plane "back", and the face edge S3 of the square projection plane "left" is connected to the face edge S4 of the square projection plane "top". There are image content discontinuities between face edges S1 and S6, between face edges S2 and S5, and between face edges S3 and S4. Therefore, the discontinuity boundaries between the top and bottom subframes include discontinuous edges between projection planes "Right" and "Bottom", discontinuous edges between projection planes "Front" and "Back", and discontinuous edges between projection planes "Left" and "Top". The image quality around the discontinuity boundaries between the top and bottom subframes of a projection-based reconstructed frame (e.g., IMG_R or IMG_R') will be degraded by a typical SAO filter. Since the pixels on opposite sides of the discontinuity boundary between the top and bottom subframes are not truly adjacent pixels, this filter applies typical SAO filtering to pixels located at the bottom boundary of the top subframe and pixels located at the top boundary of the bottom subframe. In one embodiment of the invention, Figure 4 The discontinuity boundary between the top and bottom subframes shown can be defined as a virtual boundary. That is, for the SAO filtering of the reconstructed frame IMG_R / IMG_R', the discontinuity boundary generated by the projection plane packing is aligned with a virtual boundary defined in the reconstructed frame IMG_R / IMG_R'.

[0037] You can send a syntax element (e.g., `sps_loop_filter_across_virtual_boundaries_disabled_present_flag` in Table 1) to specify that in-loop filtering is disabled on virtual boundaries. If the flag disabling loop filters across virtual boundaries is true (e.g., `sps_loop_filter_across_virtual_boundaries_disabled_present_flag == 1`), the number of vertical / horizontal virtual boundaries (e.g., `sps_num_ver_virtual_boundaries` and `sps_num_ver_virtual_boundaries` in Table 1) is also sent. The x-position of the vertical virtual boundary (e.g., `sps_virtual_boundaries_pos_x[i]` in Table 1) is sent when the number of vertical virtual boundaries is greater than 0, and the y-position of the horizontal virtual boundary (e.g., `sps_virtual_boundaries_pos_y[i]` in Table 1) is sent when the number of horizontal virtual boundaries is greater than 0. The number of bits can be reduced by sending signal positions in predefined units, such as in units of 8, using luma positions. Additionally, syntactic elements can be sent in the image header, allowing the position of vertical / horizontal virtual boundaries, the number of vertical / horizontal virtual boundaries, and the flag for disabling loop filters across virtual boundaries to vary across different images.

[0038]

[0039] Table 1

[0040] Each of SAO filters 226 and 322 can be a sample-based SAO filter operating on a coding tree unit (CTU). A CTU consists of a coding tree block (CTB) of three color components. That is, a CTU has one luma CTB and two chroma CTBs. The luma CTB consists of luma (Y) samples. One chroma CTB consists of chroma (Cb) samples, and the other chroma CTB consists of chroma (Cr) samples. In other words, SAO filtering can use blocks (e.g., CTBs) as basic processing units, where pixels in a block can be luma or chroma samples. In one exemplary design, SAO filters 226 / 322 can be implemented by dedicated hardware for performing SAO filtering on pixels in a block. In another exemplary design, SAO filters 226 / 322 can be implemented by a general-purpose processor that executes program code to perform SAO filtering on pixels in a block. However, these are for illustrative purposes only and are not intended to limit the invention.

[0041] Each of SAO filters 226 and 322 supports three different filter modes: not applied mode (SAO type == 0), band offset (BO) mode (SAO type == 1), and edge offset (EO) mode (SAO type == 2). In the not applied mode, the pixel remains unchanged. In BO mode, the SAO filtering of the current pixel in the block depends on the intensity of the current pixel. That is, pixels in a block are classified into multiple bands based on their pixel intensity, and an offset value is added to one or more bands. In EO mode, the SAO filtering of the current pixel in the block depends on the relationship between the current pixel and its neighboring pixels. Furthermore, EO mode has four directional modes. Figure 5 This is a diagram illustrating the horizontal style (SAO EO class == 0) used for pixel classification in EO mode. For example... Figure 6 This diagram illustrates the vertical style (SAO EO class == 1) used for pixel classification in EO mode. For example... Figure 7 This is a diagram illustrating the 135-degree diagonal pattern (SAO EO class == 2) used for pixel classification in EO mode. For example... Figure 8 This diagram illustrates the 45-degree diagonal pattern (SAO EO class == 3) used for pixel classification in EO mode. Figure 5-8 In each image, the pixel value of the current pixel is labeled "c", and the pixel values ​​of adjacent pixels are labeled "a" and "b" respectively. For each directional pattern, pixels in the same block are classified into different edge types based on the relationship between the current pixel and its neighbors. For example, the current pixel can be classified into one of five edge types according to the following classification rules, where the edge type can be monotonic, minimum, maximum, flat segment with downward slope, and flat segment with upward slope.

[0042]

[0043] Table 2

[0044] In this embodiment, at the encoder-side SAO filter 226, the offset values ​​for each different edge type (edgeIdx = 0, 1, 2, 3, and 4) of a given EO class are appropriately calculated according to rate-distortion optimization (RDO) and explicitly signaled to the decoder-side SAO filter 322 to effectively reduce sample distortion. Meanwhile, classification of each pixel is performed at both the encoder-side SAO filter 226 and the decoder-side SAO filter 322 to significantly conserve side information. For example, information about the EO class selected for a block and information about the offset values ​​selected for the edge types of the EO class can be signaled from the video encoder 200 to the video decoder 300. Specifically, the SAO parameters encoded in the bitstream BS can include SAO type information and offset information for the block that is SAO filtered using the EO mode. The SAO type information includes a syntactic element indicating that the SAO type is an EO mode and another syntactic element indicating the selected EO class. The offset information includes syntax elements indicating the offset values ​​for different edge types (edgeIdx = 0, 1, 2, 3, and 4) of the selected EO class. Therefore, the video decoder 300 obtains the SAO parameters of the block via the decoded bitstream BS. For EO mode SAO processing of the same block in reconstructed frames IMG_R and IMG_R', the behavior of the decoder-side SAO filter 322 is similar to that of the encoder-side SAO filter 226. For example, after the current pixel in the current block of a reconstructed projection-based frame IMG_R' is classified as one of the edge types of the selected EO class indicated by the SAO type information obtained from the decoded bitstream BS, the SAO filter 322 can add the offset value of the edge type to which the pixel is classified, where the offset value of the edge type is indicated by the offset information obtained from the decoded bitstream BS.

[0045] However, at least one neighboring pixel and the current pixel used to determine the selected orientation style of the edge type may be located at a virtual boundary defined in the reconstructed frame (e.g., using a virtual boundary defined in the reconstructed frame). Figure 4 The reconstructed compact CMP layout 404 shows the opposite sides of the discontinuous boundary between the top and bottom subframes of the projected frame IMG_R / IMG_R'. This invention proposes that for each SAO filter 226 and 322, filtering can be disabled when the EO mode is applied to the current pixel.

[0046] Figure 9This is a flowchart illustrating a video processing method according to an embodiment of the present invention. Each of the video encoders 116, 200 and the video decoders 122, 300 can employ the video processing method. In step 902, processing to generate a reconstructed frame is performed. For example, encoding circuit 204 receives a video frame (e.g., a projection-based frame IMG having 360-degree content represented by a projection plane arranged in a 360VR projection layout L_VR) and encodes the video frame to generate a portion of a bitstream (e.g., a portion of a bitstream BS including encoded information of the projection-based frame IMG). While the video frame is being encoded, encoding circuit 204 generates a reconstructed frame IMG_R at reconstruction circuit 217. As another example, decoding circuit 320 receives a bitstream (e.g., a bitstream BS generated from encoding circuit 204) and decodes a portion of the bitstream, wherein a portion of the bitstream includes encoded information of the video frame (e.g., encoded information of the projection-based frame IMG).

[0047] In step 904, a loop filter is applied to the reconstructed frame, wherein the loop filter includes performing a SAO filter operation. For example, the SAO filter operation including steps 906, 908, 910, and 912 is performed at SAO filter 226 of the encoding circuit 204. As another example, the SAO filter operation including steps 906, 908, 910, and 912 is performed at SAO filter 322 of the decoding circuit 320.

[0048] In step 906, the SAO filtering operation checks several conditions for the current pixel in the current block of the reconstructed frame IMG_R / IMG_R'. In this embodiment, the conditions may include whether the current pixel has applied EO mode, whether the flag disabling loop filters across virtual boundaries is true, and whether filtering is applied across a virtual boundary defined in the reconstructed frame IMG_R / IMG_R'. For example, the first condition is satisfied when the SAO type of the current block is EO mode; the second condition is satisfied when the syntax element sps_loop_filter_across_virtual_boundaries_disabled_present_flag is equal to 1; and the third condition is satisfied when the current pixel of the selected SAO EO class (i.e., the selected direction style) and at least one adjacent pixel are located on different sides of a virtual boundary in EO mode.

[0049] If at least one of the conditions checked in step 906 is not met (step 908), the process proceeds to step 912. In step 912, SAO filters 226 / 322 apply SAO filtering to the current pixel in a typical manner, regardless of the virtual boundary defined in the reconstructed frame IMG_R / IMG_R'. That is, depending on the SAO type selected for the current block's SAO filtering, SAO filtering is applied to the current pixel according to the parameters specified by the EO mode, BO mode, or no-application mode.

[0050] If all conditions checked in step 906 are met (step 908), the process continues to step 910. In step 910, SAO filters 226 / 322 maintain the value of the current pixel by preventing the SAO filtering operation of the current pixel from being applied to the virtual boundary defined in the reconstructed frame. Since SAO filtering is disabled across the virtual boundary in EO mode, visual quality and / or encoding / decoding efficiency around the virtual boundary can be improved. To better understand the technical features of the proposed SAO filter design, several scenarios for disabling SAO filtering of the current pixel in EO mode are provided below.

[0051] In the first scenario, the reconstructed frame IMG_R / IMG_R' has a virtual boundary, which is a vertical virtual boundary. The direction style selected by the current pixel and its neighboring pixels in EO mode is not... Figure 6 The vertical pattern shown indicates that the current pixel is located to the left of and close to the vertical virtual boundary, and at least one adjacent pixel is located to the right of the vertical virtual boundary. SAO filtering of the current pixel can be disabled when the "Disable Loop Filter Across Virtual Boundaries" flag is true. For example, the x-position xS of the current pixel in the current CTB of width nCtbSw. i The value of `VirtualBoundariesPosX[n]` minus one, equal to the x-position of one of the vertical virtual boundaries, disables filtering for the current pixel, where `i` is selected from {0,...,nCtbSw-1}. The variable `VirtualBoundariesNumVer` specifies the number of vertical virtual boundaries in the reconstructed frame `IMG_R / IMG_R'`. The variable `cIdx` specifies the color component index of the CTB; for Y, the color component index is 0; for Cb, it is 1; and for Cr, it is 2. A pair of variables (rx, ry) specifies the CTB position. An example of semantic design is as follows:

[0052] For any value in n = 0..VirtualBoundariesNumVer-1, VirtualBoundariesDisabledFlag equals 1, xS iIt equals ((VirtualBoundariesPosX[n] / scaleWidth)-1), and SaoEoClass[cIdx][rx][ry] is not equal to 1.

[0053] In the second case, the reconstructed frame IMG_R / IMG_R' has a virtual boundary, which is a vertical virtual boundary. The orientation style selected by the current pixel and its neighboring pixels in EO mode is not... Figure 6 The vertical pattern shown indicates that the current pixel is located to the right of and close to the vertical virtual boundary, and at least one adjacent pixel is located to the left of the vertical virtual boundary. SAO filtering of the current pixel can be disabled when the "Disable Loop Filter Across Virtual Boundaries" flag is true. For example, at position xS of the current pixel in the current CTB of width nCtbSw. i The x-position of one of the vertical virtual boundaries, VirtualBoundariesPosX[n], can disable filtering for the current pixel, where i is selected from {0,...,nCtbSw-1}. An example of semantic design is as follows:

[0054] For any value in n = 0..VirtualBoundariesNumVer-1, VirtualBoundariesDisabledFlag is equal to 1, xS i It equals (VirtualBoundariesPosX[n] / scaleWidth), and SaoEoClass[cIdx][rx][ry] is not equal to 1.

[0055] In the third case, the reconstructed frame IMG_R / IMG_R' has a virtual boundary, which is a horizontal virtual boundary. The orientation style selected by the current pixel and its neighboring pixels in EO mode is not... Figure 5 The horizontal pattern shown indicates that the current pixel is above and close to the horizontal virtual boundary, and at least one of its neighboring pixels is below the horizontal virtual boundary. SAO filtering for the current pixel can be disabled when the "Disable Loop Filter Across Virtual Boundaries" flag is true. For example, the y-position yS of the current pixel in the current CTB of height nCtbSh. j The value of `VirtualBoundariesPosY[n]` minus one, equal to the y-position of one of the horizontal virtual boundaries, disables filtering for the current pixel, where `j` is selected from {0,...,nCtbSh-1}. The variable `VirtualBoundariesNumHor` specifies the number of horizontal virtual boundaries in the reconstructed frame `IMG_R / IMG_R'`. An example of semantic design is as follows:

[0056] For any value in n = 0..VirtualBoundariesNumHor–1, VirtualBoundariesDisabledFlag equals 1, yS j It equals ((VirtualBoundariesPosY[n] / scaleHeight)-1), and SaoEoClass[cIdx][rx][ry] is not equal to 0.

[0057] In the fourth scenario, the reconstructed frame IMG_R / IMG_R' has a virtual boundary, which is a horizontal virtual boundary. The orientation style selected by the current pixel and its neighboring pixels in EO mode is not... Figure 5 The horizontal pattern shown indicates that the current pixel is below and close to the horizontal virtual boundary, and at least one of its adjacent pixels is above the horizontal virtual boundary. SAO filtering for the current pixel can be disabled when the "Disable Loop Filter Across Virtual Boundaries" flag is true. For example, the y-position yS of the current pixel in the current CTB of height nCtbSh. j The y-position of one of the horizontal virtual boundaries, VirtualBoundariesPosY[n], can disable filtering for the current pixel, where j is selected from {0,...,nCtbSh-1}. An example of semantic design is as follows:

[0058] For any value in n = 0..VirtualBoundariesNumHor–1, VirtualBoundariesDisabledFlag equals 1, yS j It equals (VirtualBoundariesPosY[n] / scaleHeight), and SaoEoClass[cIdx][rx][ry] is not equal to 0.

[0059] The above conditions apply to both the luma and chroma components. That is, the current pixel can be a luma sample when the SAO filter 226 / 322 processes the luma CTB of the CTU, and a chroma sample when the SAO filter 226 / 322 processes the chroma CTB of the CTU. As mentioned above, the position of the virtual boundary is represented by the luma position. For SAO filtering of the luma component, it is not necessary to scale / transform the position of the virtual boundary to check the above conditions to indicate whether the filtering process is disabled on the virtual boundary in the luma CTB. Therefore, scaleWidth and scaleHeight can be omitted in the above semantic design. Alternatively, for SAO filtering of the luma component, the position can be converted / scaled by 1 (i.e., scaleWidth = 1 and scaleHeight = 1) to check the above conditions to indicate whether the filtering process is disabled on the virtual boundary in the luma CTB.

[0060] For SAO filtering of the chroma component, the position of the virtual boundary represented by the signal at the luma position can be scaled / converted to the chroma position, where scaleWidth ≠ 1 and scaleHeight ≠ 1. The scaled / converted position can be used to check these conditions to indicate whether the filtering process is disabled at the virtual boundary in the chroma CTB.

[0061] As described above, when all the conditions checked in step 906 are met, SAO filters 226 / 322 prevent the application of SAO filtering operations for the current pixel included in the reconstructed frame across the virtual boundary defined in the reconstructed frame. In a first exemplary design that disables SAO filtering for the current pixel, SAO filters 226 / 322 may add a zero offset to the value of the current pixel, thereby keeping the value of the current pixel unchanged in EO mode.

[0062] In a second exemplary design that disables SAO filtering for the current pixel, the SAO filters 226 / 322 can intentionally set the edge type of the current pixel using a monotonic type (edgeIdx = 0). Typically, the offset of the monotonic type is zero. Therefore, the value of the current pixel remains unchanged in EO mode.

[0063] In the third exemplary design where SAO filtering is disabled for the current pixel, SAO filters 226 / 322 can directly skip the SAO filtering operation for the current pixel. Since no arithmetic operation is performed to add an offset value to the current pixel's value, the current pixel's value remains unchanged in EO mode.

[0064] The above description is presented to enable those skilled in the art to practice the invention provided in the context of a particular application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments. Therefore, the invention is not intended to be limited to the specific embodiments shown and described, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein. In the above detailed description, various specific details have been set forth in order to provide a thorough understanding of the invention. However, those skilled in the art will understand that the invention can be practiced.

[0065] The embodiments of the invention described above can be implemented in various hardware, software code, or combinations thereof. For example, one embodiment of the invention may be circuitry integrated into a video compression chip or program code integrated into video compression software to perform the processes described herein. Embodiments of the invention may also be program code to be executed on a digital signal processor (DSP) to perform the processes described herein. The invention may also relate to numerous functions performed by a computer processor, digital signal processor, microprocessor, or field-programmable gate array (FPGA). These processors may be configured to perform specific tasks according to the invention by executing machine-readable software code or firmware code that defines the specific methods embodied in the invention. The software code or firmware code may be developed in different programming languages ​​and in different formats or styles. The software code may also be compiled for different target platforms. However, different code formats, styles, and languages ​​of the software code, as well as other ways of configuring the code to perform the tasks according to the invention, will not depart from the spirit and scope of the invention.

[0066] The invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described examples are to be considered illustrative rather than limiting in all respects. Therefore, the scope of the invention is indicated by the appended claims rather than by the foregoing description. All variations within the equivalent meaning and scope of the claims should be included within their scope.

[0067] Those skilled in the art will readily observe that many modifications and alterations can be made to the apparatus and method while retaining the teachings of the present invention. Therefore, the foregoing disclosure should be construed as being limited only by the scope and limits of the appended claims.

Claims

1. A video processing method, comprising: Receive reconstruction frame; as well as Applying loop filtering to the reconstructed frame using at least one loop filter includes: Perform sample adaptive offset filtering operation, including: When the filtering mode of the sample adaptive offset filtering operation of the current pixel is edge offset mode, the flag of disabling the loop filter across the virtual boundary is true, and the current pixel and at least one neighboring pixel used to filter the current pixel are located on different sides of the virtual boundary, the value of the current pixel is kept unchanged by preventing the application of the sample adaptive offset filtering operation of the current pixel included in the reconstruction frame across the virtual boundary defined in the reconstruction frame; otherwise, the sample adaptive offset is applied to the current pixel. Specifically, by setting the edge type of the current pixel to a monotonic type, the application of the sample adaptive offset filtering operation of the current pixel included in the reconstruction frame across the virtual boundary defined in the reconstruction frame is prevented. The edge type of the current pixel in the edge offset mode is classified as the monotonic type only if the value (c) of the current pixel is less than the value (a, b) of one neighboring pixel used to filter the current pixel and greater than the value (a, b) of another neighboring pixel used to filter the current pixel, or if the value (c) of the current pixel is equal to the values ​​(a, b) of two neighboring pixels used to filter the current pixel, and the offset of the monotonic type is zero.

2. The video processing method as described in claim 1, wherein, The reconstructed frame is a projection-based frame that includes multiple projection surfaces encapsulated in a projection layout of a 360-degree virtual reality (360VR) projection. The 360-degree image content of the sphere is mapped onto these multiple projection surfaces, and the virtual boundary is aligned with the discontinuous boundary of the image content generated by the encapsulation of these multiple projection surfaces in the projection-based frame.

3. The video processing method as described in claim 1, characterized in that, The virtual boundary is a vertical virtual boundary; the selected direction style of the current pixel and multiple adjacent pixels in the edge offset mode is not a vertical style; the current pixel is located to the left of the vertical virtual boundary; one of the multiple adjacent pixels is located to the right of the vertical virtual boundary.

4. The video processing method as described in claim 1, characterized in that, The virtual boundary is a vertical virtual boundary; the selected direction style of the current pixel and multiple adjacent pixels in the edge offset mode is not a vertical style; the current pixel is located to the right of the vertical virtual boundary; one of the multiple adjacent pixels is located to the left of the vertical virtual boundary.

5. The video processing method as described in claim 1, characterized in that, The virtual boundary is a horizontal virtual boundary; the selected direction style of the current pixel and multiple adjacent pixels in the edge offset mode is not a horizontal style; the current pixel is above the horizontal virtual boundary; one of the multiple adjacent pixels is below the horizontal virtual boundary.

6. The video processing method as described in claim 1, characterized in that, The virtual boundary is a horizontal virtual boundary; the selected direction style of the current pixel and multiple adjacent pixels in the edge offset mode is not a horizontal style; the current pixel is located below the horizontal virtual boundary; one of the multiple adjacent pixels is located above the horizontal virtual boundary.

7. A video processing apparatus, comprising: Encoding circuitry is used to receive video frames and encode them to generate a portion of a bitstream. Specifically, when encoding the video frame, the encoding circuit generates a reconstructed frame and applies loop filtering to the reconstructed frame. The loop filtering includes a sample adaptive offset filtering operation performed at a sample adaptive offset filter included in the encoding circuit. When the filtering mode of the sample adaptive offset filtering operation of the current pixel is edge offset mode, the flag for disabling the loop filter across the virtual boundary is true, and the current pixel and at least one adjacent pixel used to filter the current pixel are located on different sides of the virtual boundary, the sample adaptive offset filter is arranged to keep the value of the current pixel unchanged by preventing the application of the sample adaptive offset filtering operation of the current pixel included in the reconstructed frame across the virtual boundary defined in the reconstructed frame; otherwise, a sample adaptive offset is applied to the current pixel. Specifically, by setting the edge type of the current pixel to a monotonic type, the application of the sample adaptive offset filtering operation of the current pixel included in the reconstruction frame across the virtual boundary defined in the reconstruction frame is prevented. The edge type of the current pixel in the edge offset mode is classified as the monotonic type only if the value (c) of the current pixel is less than the value (a, b) of one neighboring pixel used to filter the current pixel and greater than the value (a, b) of another neighboring pixel used to filter the current pixel, or if the value (c) of the current pixel is equal to the values ​​(a, b) of two neighboring pixels used to filter the current pixel, and the offset of the monotonic type is zero.

8. A video processing apparatus, comprising: A decoding circuit is used to receive a bitstream and decode a portion of the bitstream, wherein the portion of the bitstream includes the encoded information of video frames; When decoding this portion of the bitstream, the decoding circuit generates a reconstructed frame and applies a loop filter to the reconstructed frame. The loop filter includes a sample adaptive offset filtering operation performed at a sample adaptive offset filter included in the decoding circuit. When the filtering mode of the sample adaptive offset filtering operation for the current pixel is edge offset mode, the flag disabling the loop filter across the virtual boundary is true, and the current pixel and at least one adjacent pixel used to filter the current pixel are located on different sides of the virtual boundary, the sample adaptive offset filter is arranged to keep the value of the current pixel unchanged by preventing the application of the sample adaptive offset filtering operation for the current pixel included in the reconstructed frame to cross the virtual boundary defined in the reconstructed frame; otherwise, a sample adaptive offset is applied to the current pixel. Specifically, by setting the edge type of the current pixel to a monotonic type, the application of the sample adaptive offset filtering operation of the current pixel included in the reconstruction frame across the virtual boundary defined in the reconstruction frame is prevented. The edge type of the current pixel in the edge offset mode is classified as the monotonic type only if the value (c) of the current pixel is less than the value (a, b) of one neighboring pixel used to filter the current pixel and greater than the value (a, b) of another neighboring pixel used to filter the current pixel, or if the value (c) of the current pixel is equal to the values ​​(a, b) of two neighboring pixels used to filter the current pixel, and the offset of the monotonic type is zero.

Citation Information

Patent Citations

  • Video encoding method and apparatus with in-loop filtering process not applied to reconstructed blocks located at image content discontinuity edge and associated video decoding method and apparatus

    CN109219958A

  • 360-degree video coding using face continuities

    WO2018191224A1

  • Handling face discontinuities in 360-degree video coding

    WO2019060443A1