A video decoding method

By introducing protection bands and constrained projection surface sizes in the decoding process of 360-degree all-round video, the problem of difficult to compress and decode high bit rate 360-degree video in the prior art is solved, and the effect of reducing bandwidth and improving picture quality is achieved.

CN114651271BActive Publication Date: 2025-05-27MEDIATEK INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080073248.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-21
Filing Date
2020-10-28
Publication Date
2025-05-27
Estimated Expiration
2040-10-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively compress and decode high bit rate 360-degree all-around video, especially in chroma subsampling technology, and an innovative design is needed to impose constraints on projection-based frames to reduce bandwidth and maintain picture quality.

Method used

A video decoding method is proposed to generate frames with a constraint protection belt size, a constraint projection surface size, and/or a constraint picture size by decoding a part of the bit stream. The method includes packing at least one protective tape in the projection layout and mapping a portion of the spherical 360-degree content to at least one projection surface by projection, ensuring that the size of each protective tape is an even number of brightness samples.

Benefits of technology

By introducing protection bands and constrained projection surface size, the bandwidth is effectively reduced, seam artifacts are reduced, picture quality is improved, and efficient decoding of high-bit rate 360-degree video is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114651271B_ABST
    Figure CN114651271B_ABST
Patent Text Reader

Abstract

A video decoding method includes: decoding a part of a bitstream to generate a decoded frame. The decoded frame is a projection-based frame, including at least one projection face and at least one guard band packed in a projection layout. At least a part of a spherical 360-degree content is mapped to the at least one projection face via projection. The decoded frame is in a 4:2:0 chrominance format or a 4:2:2 chrominance format, and a guard band size of each of the at least one guard band is equal to an even number of luminance samples.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference

[0002] This application claims the following priority: U.S. Provisional Patent Application No. 62 / 926,600, filed on October 28, 2019, and U.S. Provisional Patent Application No. 63 / 042,610, filed on June 23, 2020. The entire contents of the related applications (including U.S. Provisional Patent Application No. 62 / 926,600 and U.S. Provisional Patent Application No. 63 / 042,610) are incorporated herein by reference. Technical Field

[0003] The present invention relates to video processing, and in particular, to a video decoding method for decoding a portion of a bitstream to generate frames based on projections and having a constrained post-guard band size, a constrained post-projection plane size, and / or a constrained post-picture size. Background Art

[0004] Virtual reality (VR) with head-mounted displays (HMDs) is associated with a variety of applications. The ability to display a wide field of view of viewing content to a user can be used to provide an immersive visual experience. A real-world environment must be captured from all directions to form an omnidirectional video corresponding to a viewing sphere. With the development and progress of camera devices and HMDs, the transmission of VR content may quickly become a bottleneck due to the high bitrates required to represent such 360-degree content. When the resolution of the omnidirectional video is 4K or higher, data compression / encoding is crucial for reducing the bitrate.

[0005] Generally, an omnidirectional video corresponding to a sphere is converted into a frame having 360-degree image content, which is represented by one or more projection planes arranged in a 360-degree virtual reality (360VR) projection layout, and then the resulting frame is encoded into a bitstream for transmission. Chroma subsampling is a video coding technique that uses less resolution for chroma information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance. In other words, chroma subsampling is a type of compression that reduces the color information in a signal, thus reducing the bandwidth without significantly affecting the picture quality. For example, a video encoder can encode a projection-based frame into a bitstream for transmission in a 4:2:0 chroma format or a 4:2:2 chroma format. Therefore, when chroma subsampling is performed with less resolution for chroma information, an innovative design is needed to impose constraints on projection-based frames. Summary of the Invention

[0006] One object of the claimed invention is to provide a video decoding method for decoding a portion of a bitstream to generate a frame based on projection and having a constrained post-padding size, a constrained post-projection plane size, and / or a constrained picture size.

[0007] According to a first aspect of the present invention, an exemplary video decoding method is disclosed. The exemplary video decoding method includes: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including at least one projection plane and at least one padding packed in a projection layout, and at least a portion of a spherical 360-degree content is mapped to the at least one projection plane via projection. The decoded frame is in a 4:2:0 chroma format or a 4:2:2 chroma format, and a padding size of each of the at least one padding is equal to an even number of luminance samples.

[0008] According to a second aspect of the present invention, an exemplary video decoding method is disclosed. The exemplary video decoding method includes: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including a plurality of projection planes packed in a projection layout, and at least a portion of a spherical 360-degree content is mapped to the projection planes via projection. The decoded frame is in a 4:2:0 chroma format or a 4:2:2 chroma format, and a width of each of the projection planes is equal to an even number of luminance samples.

[0009] According to a third aspect of the present invention, an exemplary video decoding method is disclosed. The exemplary video decoding method includes: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including a plurality of projection planes packed in a projection layout, and at least a portion of a spherical 360-degree content is mapped to the projection planes via projection. The decoded frame is in a 4:2:0 chroma format, and a height of each of the projection planes is equal to an even number of luminance samples.

[0010] According to a fourth aspect of the present invention, an exemplary video decoding method is disclosed. The exemplary video decoding method includes: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including a plurality of projection planes packed in a projection layout having M projection plane columns and N projection plane rows, where M and N are positive integers, and at least a portion of a spherical 360-degree content is mapped to the projection planes via projection. For the decoded frame, a picture width excluding padding samples is an integer multiple of M, and a picture height excluding padding samples is an integer multiple of N.

[0011] According to a fifth aspect of the present invention, an exemplary video decoding method is disclosed. The exemplary video decoding method includes: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including a plurality of projection planes packed in a hemispherical cube map projection layout, and a portion of a spherical 360-degree content is mapped to the projection planes via hemispherical cube map projection. With respect to the decoded frame, one of the picture width excluding guard band samples and the picture height excluding guard band samples is equal to an integer multiple of 6.

[0012] These and other objects of the present invention will no doubt become apparent to those skilled in the art after reading the following detailed description of the preferred embodiments illustrated in the various drawings and figures. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 As a diagram, a 360-degree virtual reality (360VR) system is illustrated according to an embodiment of the present invention.

[0014] Figure 2 As a diagram, a cube-based projection is illustrated according to an embodiment of the present invention.

[0015] Figure 3 As a diagram, another cube-based projection is illustrated according to an embodiment of the present invention.

[0016] Figures 4-6 As a diagram, a specification of a guard band packed in a regular cube map projection or a hemispherical cube map projection is illustrated according to an embodiment of the present invention.

[0017] Figures 7-9 As a diagram, another specification of a guard band packed in a regular cube map projection or a hemispherical cube map projection is illustrated according to an embodiment of the present invention.

[0018] Figure 10 As a diagram, an example of a 1x6 layout without a guard band is illustrated.

[0019] Figure 11 As a diagram, an example of a 2x3 layout without a guard band is illustrated.

[0020] Figure 12 As a diagram, an example of a 3x2 layout without a guard band is illustrated.

[0021] Figure 13 As a diagram, an example of a 6x1 layout without a guard band is illustrated. DETAILED DESCRIPTION

[0022] Certain terms are used throughout the following description and the claims to refer to particular components. As one skilled in the art will recognize, electronic device manufacturers may refer to a component by different names. This document does not intend to distinguish between components that have different names but perform the same function. In the following description and the claims, the terms "comprising" and "including" are used in an open-ended fashion and should be interpreted to mean "including, but not limited to...". The term "coupled" is also intended to mean an indirect or direct electrical connection. Thus, if a device is coupled to another device, that connection may be through a direct electrical connection or through other devices and connections by an indirect electrical connection.

[0023] Figure 1 For illustration, a 360-degree virtual reality (360VR) system is depicted in accordance with an embodiment of the present invention. The 360VR system 100 includes two video processing devices (e.g., a source electronic device 102 and a destination electronic device 104). The source electronic device 102 includes a video capture device 112, a conversion circuit 114, and a video encoder 116. For example, the video capture device 112 may be a set of cameras to provide an omnidirectional image content (e.g., multiple images covering an entire scene) S_IN corresponding to a sphere. The conversion circuit 114 is coupled between the video capture device 112 and the video encoder 116. The conversion circuit 114 generates a projection-based frame IMG having a 360-degree virtual reality (360VR) projection layout L_VR based on the omnidirectional image content S_IN. For example, the projection-based frame IMG may be one of a series of projection-based frames generated by the conversion circuit 114. The video encoder 116 is an encoding circuit for encoding / compressing the projection-based frames IMG to generate a part of a bitstream BS. In addition, the video encoder 116 outputs the bitstream BS to the destination electronic device 104 through a transmission means 103. For example, the series of projection-based frames may be encoded into the bitstream BS, and the transmission means 103 may be a wired / wireless communication link or a storage medium.

[0024] The destination electronic device 104 may be a head-mounted display (HMD). As Figure 1As shown, the destination electronic device 104 includes a video decoder 122, a graphic rendering circuit 124, and a display screen 126. The video decoder 122 is a decoding circuit for receiving a bitstream BS from a transmission medium 103 (e.g., a wired / wireless communication link or a storage medium), and decoding a part of the received bitstream BS to generate a decoded frame IMG'. For example, the video decoder 122 generates a series of decoded frames by decoding the received bitstream BS, and the decoded frame IMG' is one of the frames included in the series of decoded frames. In this environment, the projection-based frame IMG to be encoded on the encoder side is in a 360VR projection format with a projection layout. Therefore, after the bitstream BS is decoded on the decoder side, the decoded frame IMG' has the same 360VR projection format and the same projection layout. The graphic rendering circuit 124 is coupled between the video decoder 122 and the display screen 126. The graphic rendering circuit 124 generates and displays an output image data on the display screen 126 according to the decoded frame IMG'. For example, a viewport area associated with a part of the 360-degree image content carried by the decoded frame IMG' can be displayed on the display screen 126 via the graphic rendering circuit 124.

[0025] As described above, the conversion circuit 114 generates the projection-based frame IMG according to the 360VR projection layout L_VR and the omnidirectional image content S_IN. In this environment, the 360VR projection layout L_VR can be selected from a group consisting of: a cube-based projection layout with / without a guard band, a triangle-based projection layout with / without a guard band, a segmented sphere projection layout with / without a guard band, a rotated sphere projection layout with / without a guard band, a viewport-dependent projection layout with / without a guard band, an equi-rectangular projection layout with / without a guard band, an equi-angular cubemap projection layout with / without a guard band, and an equatorial cylindrical projection layout with / without a guard band. For example, the 360VR projection layout L_VR can be set by a general cube map projection layout with / without a guard band or a hemisphere cube map projection layout with / without a guard band.

[0026] Consider an example where the 360VR projection layout L_VR is a cube-based projection layout. Thus, at least a portion (i.e., part or all) of a spherical 360-degree content is mapped to a projection surface via a cube-based projection, and projection surfaces derived from different faces of a three-dimensional object (such as a cube or a hemispherical cube) are packed into a two-dimensional cube-based projection layout, and the projection layout is implemented based on the projection frame IMG / decoded frame IMG'.

[0027] In one embodiment, a cube-based projection with six square projection surfaces representing the entire 360°x180° omnidirectional video (i.e., all of a spherical 360-degree content) can be implemented. Regarding the conversion circuit 114 of the source electronic device 102, a cube-based projection is implemented to generate square projection surfaces of a cube in a three-dimensional (3D) space. Figure 2 For illustration, a cube-based projection is shown according to an embodiment of the present invention. The entire 360-degree content of the sphere 200 is projected onto six square faces of a cube 201, including an upper face (labeled "up"), a lower face (labeled "down"), a left face (labeled "left"), a front face (labeled "front"), a right face (labeled "right"), and a rear face (labeled "rear"). As Figure 2 shown, an image content of a northern end region of the sphere 200 is projected onto the upper face "up", an image content of a southern end region of the sphere 200 is projected onto the lower face "down", and an image content of an equatorial region of the sphere 200 is projected onto the left face "left", the front face "front", the right face "right", and the rear face "rear".

[0028] Forward transformation can be used to transform from 3D space to a 2D plane. Thus, the upper face "up", the lower face "down", the left face "left", the front face "front", the right face "right", and the rear face "rear" of the cube 201 in 3D space are transformed into an upper face (labeled "2"), a lower face (labeled "3"), a left face (labeled "5"), a front face (labeled "0"), a right face (labeled "4"), and a rear face (labeled "1") of a 2D plane.

[0029] Inverse transformation can be used to transform from a 2D plane to 3D space. Thus, an upper face (labeled "2"), a lower face (labeled "3"), a left face (labeled "5"), a front face (labeled "0"), a right face (labeled "4"), and a rear face (labeled "1") of a 2D plane are transformed into the upper face "up", the lower face "down", the left face "left", the front face "front", the right face "right", and the rear face "rear" of the cube 201 in 3D space.

[0030] The inverse transformation can be used by the conversion circuit 114 of the source electronic device 102 to generate top "2", bottom "3", left "5", front "0", right "4", and back "1" above. The top "2", bottom "3", left "5", front "0", right "4", and back "1" on the 2D plane are packed to form the projection-based frame IMG encoded by the video encoder 116.

[0031] The video decoder 122 receives the bitstream BS from the transmission medium 103 and decodes a part of the received bitstream BS to generate a decoded frame IMG', which has the same projection layout L_VR as that used on the encoder side. Regarding the graphics generation circuit 124 of the destination electronic device, the forward transformation can be used to transform from 3D space to 2D plane to determine the pixel values of the pixels in any of top "up", bottom "down", left "left", front "front", right "right", and back "back". Alternatively, the inverse transformation can be used to transform from 2D plane to 3D space to remap the sample positions of a projection-based frame onto a sphere.

[0032] As described above, the top "2", bottom "3", left "5", front "0", right "4", and back "1" are packed to form the projection-based frame IMG. For example, the conversion circuit 114 can select a packing type such that the projection-based frame IMG can have the projected image data arranged as in the cube-based projection layout 202, where the projection-based frame IMG with the cube-based projection layout 202 has a picture height cmpPicHeight and a picture width cmpPicWidth, and each square face packed in the cube-based projection layout 202 has a face height faceHeight and a face width faceWidth. In another example, the conversion circuit 114 can select another packing type such that the projection-based frame IMG can have the projected image data different from that of the cube-based projection layout 202 and arranged as in the cube-based projection layout 204, where the projection-based frame IMG with the cube-based projection layout 204 has a picture height cmpPicHeight and a picture width cmpPicWidth, and each square face packed in the cube-based projection layout 204 has a face height faceHeight and a face width faceWidth.

[0033] In another embodiment, a cube-based projection with five square projection planes (including one full face and four half faces) representing a 180°x180° omnidirectional video (i.e., a part of a spherical 360-degree content) can be performed. Regarding the conversion circuit 114 of the source electronic device 102, the cube-based projection is performed to generate one full face and four half faces of a cube in 3D space. Figure 3 As an illustration, another cube-based projection is depicted according to an embodiment of the present invention. Only half of the 360-degree content of the sphere 200 is projected onto the faces of a cube 201, including: an upper half face (labeled "Top_H"), a lower half face (labeled "Bottom_H"), a left half face (labeled "Left_H"), a front full face (labeled "Front"), and a right half face (labeled "Right_H"). In this example, a hemispherical cube (e.g., half of the cube 201) is used for the hemispherical cube mapping projection, where a hemisphere (e.g., half of the sphere 200) is inscribed in the hemispherical cube (e.g., half of the cube 201). As Figure 3 shown, an image content of half of the north pole region of the sphere 200 is projected onto the upper half face "Top_H", an image content of half of the south pole region of the sphere 200 is projected onto the lower half face "Bottom_H", and an image content of half of the equatorial region of the sphere 200 is projected onto the left half face "Left_H", the front full face "Front", and the right half face "Right_H".

[0034] Forward conversion can be used to convert from 3D space to a 2D plane. Thus, the upper half face "Top_H", the lower half face "Bottom_H", the left half face "Left_H", the front full face "Front", and the right half face "Right_H" of the cube 201 in 3D space are converted to an upper half face (labeled "2"), a lower half face (labeled "3"), a left half face (labeled "5"), a front full face (labeled "0"), and a right half face (labeled "4") of the 2D plane. Additionally, the size of the front full face (labeled "0") is twice the size of each of the upper half face (labeled "2"), the lower half face (labeled "3"), the left half face (labeled "5"), and the right half face (labeled "4").

[0035] Inverse transformation can be used to transform from a 2D plane to a 3D space. Thus, the upper half plane (labeled "2"), lower half plane (labeled "3"), left half plane (labeled "5"), front full plane (labeled "0"), and right half plane (labeled "4") of the 2D plane are transformed into the upper half "Top_H", lower half "Bottom_H", left half "Left_H", front "Front", and right half "Right_H" of the cube 201 in the 3D space.

[0036] Inverse transformation can be used by the transformation circuit 114 of the source electronic device 102 to generate the upper half plane "2", lower half plane "3", left half plane "5", front full plane "0", and right half plane "4". The upper half plane "2", lower half plane "3", left half plane "5", front full plane "0", and right half plane "4" on the 2D plane are packed to form the projection-based frame IMG encoded by the video encoder 116.

[0037] The video decoder 122 receives the bitstream BS from the transmission medium 103 and decodes a part of the received bitstream BS to generate the decoded frame IMG', which has the same projection layout L_VR as that used on the encoder side. Regarding the graphics generation circuit 124 of the destination electronic device, forward transformation can be used to transform from a 3D space to a 2D plane to determine the pixel values of the pixels in any of the upper half "Top_H", lower half "Bottom_H", left half "Left_H", front "Front", and right half "Right_H". Alternatively, inverse transformation can be used to transform from a 2D plane to a 3D space to remap the sample positions of a projection-based frame onto a sphere.

[0038] As described above, the upper half plane "2", the lower half plane "3", the left half plane "5", the front full plane "0", and the right half plane "4" are packed to form a projection-based frame IMG. For example, the conversion circuit 114 may select a packing type such that the projection-based frame IMG may have the projected video data arranged as in the cube-based projection layout 302, where the projection-based frame IMG having the cube-based projection layout 302 has a picture height cmpPicHeight and a picture width cmpPicWidth, and the square face (i.e., the full face) packed in the cube-based projection layout 302 has a face height faceHeight and a face width faceWidth. In another example, the conversion circuit 114 may select another packing type such that the projection-based frame IMG may have the projected video data arranged as in the cube-based projection layout 304, which is different from the cube-based projection layout 302, where the projection-based frame IMG having the cube-based projection layout 304 has a picture height cmpPicHeight and a picture width cmpPicWidth, and the square face (i.e., the full face) packed in the cube-based projection layout 304 has a face height faceHeight and a face width faceWidth. Hereinafter, the cube-based projection layout 302 is also referred to as a horizontally packed hemispherical cube map projection layout, in which all projection faces are horizontally packed; and the cube-based projection layout 304 is also referred to as a vertically packed hemispherical cube map projection layout, in which all projection faces are vertically packed. In this embodiment, the front face is selected to be packed in the full face of the cube-based projection layout 302 / 304. In this approach, the full face packed in the cube-based projection layout 302 / 304 can be any one of the upper, lower, front, rear, left, and right faces, and the four half faces packed in the cube-based projection layout 302 / 304 depend on the selection of the full face.

[0039] Regarding the embodiment as Figure 2 shown, the projection faces are packed in a general CMP layout 202 / 204 without guard bands (or padding). Regarding the embodiment as Figure 3In the illustrated embodiment, the projection plane is packed in a half-spherical CMP layout 302 / 304 without a guard band (or fill). However, due to the discontinuous layout boundary (boundary) of the CMP layout (which may be a general CMP layout or a half-spherical CMP layout), and / or the discontinuous edge (edge) of the CMP layout (which may be a general CMP layout or a half-spherical CMP layout), the projection-based frame IMG may have artifacts after encoding and decoding. For example, the CMP layout without a guard band (or fill) has an upper discontinuous layout boundary, a lower discontinuous layout boundary, a left discontinuous layout boundary, and a right discontinuous layout boundary. In addition, there is at least one video content discontinuous edge between two adjacent projection planes of the CMP layout packed without a guard band (or fill). Taking the cube-based projection layout 202 / 204 as an example, a discontinuous edge exists between the boundary of the face "3" below and the boundary of the face "5" on the left, a discontinuous edge exists between the boundary of the face "1" at the back and the boundary of the face "0" in the front, and a discontinuous edge exists between the boundary of the face "2" above and the boundary of the face "4" on the right. Taking the cube-based projection layout 302 / 304 as an example, a discontinuous edge exists between the boundary of the face "3" below and the boundary of the face "5" on the left, and a discontinuous edge exists between the boundary of the face "4" on the right and the boundary of the face "2" above.

[0040] To address this issue, the 360VR projection layout L_VR can be set by a projection layout (such as a cube-based projection layout) having at least one guard band (or fill). For example, near the layout boundary and / or the discontinuous edge, additional guard bands generated by pixel padding, for example, can be inserted to reduce seam artifacts. Alternatively, near the layout boundary and / or the continuous edge, additional guard bands generated by pixel padding can be inserted. In short, the position of each guard band added to a projection layout can depend on actual design considerations. Regarding a projection-based frame having a constrained guard band size, a constrained projection plane size, and / or a constrained picture size, when a projection layout with a guard band (or fill) is implemented, the present invention places no restrictions on the number of guard bands and the position of each guard band.

[0041] In this embodiment, the conversion circuit 114 has a padding circuit 115 arranged to generate at least one padding area (i.e., at least one guard band). The conversion circuit 114 builds a projection-based frame IMG by packing the projection plane of the 360VR projection layout L_VR and at least one padding area (i.e., at least one guard band). For example, the conversion circuit 114 determines a guard band configuration for the projection-based frame IMG, and the above frame IMG is composed of a projection plane derived from a cube projection (such as Figure 2 the general cube map projection shown or Figure 3 the hemispherical cube map projection shown), and the decoded frame IMG’ is a projection-based frame generated from the video decoder 122 and has a guard band configuration identical to that of the projection-based frame IMG received and encoded by the video encoder 116.

[0042] Figures 4-6 For illustration, a specification of the guard band packed in a general cube map projection or a hemispherical cube map projection according to an embodiment of the present invention is shown. As Figure 4 shown in the upper part, a first guard band is added to the lower boundary of a first projection plane packed at a face position with position index 2, and a second guard band is added to the upper boundary of a second projection plane packed at a face position with position index 3, where the first guard band and the second guard band have the same guard band size D (i.e., the same number of guard band samples). If the lower boundary of the first projection plane is directly connected to the upper boundary of the second projection plane, an edge (such as a discontinuous edge or a continuous edge) exists between the first projection plane and the second projection plane. The guard band can be added to the edge between the first projection plane and the second projection plane. For example, regarding a cube in a 3D space, the lower boundary of the first projection plane (which is a square face of the cube) may or may not be connected to the upper boundary of the second projection plane (which is another square face of the cube); and regarding a cube-based projection layout on a 2D plane, the lower boundary of the first projection plane is parallel to the upper boundary of the second projection plane, and both the first guard band and the second guard band are located between the first projection plane and the second projection plane to isolate the lower boundary of the first projection plane from the upper boundary of the second projection plane, where the first guard band is connected to the lower boundary of the first projection plane and the second guard band, and the second guard band is connected to the first guard band and the upper boundary of the second projection plane. Therefore, the width of a guard band area (which is composed of the first guard band and the second guard band) inserted between the first projection plane (packed at the face position with position index 2) and the second projection plane (packed at the face position with position index 3) is equal to 2*D.

[0043] As Figure 4As shown in the lower part, a first protective band is added to the lower boundary of a projected surface packed at a surface position with position index 2, a second protective band is added to the upper boundary of a projected surface packed at a surface position with position index 3, a third protective band is added to the upper boundary of a projected surface packed at a surface position with position index 0, a fourth protective band is added to the lower boundary of a projected surface packed at a surface position with position index 5, a fifth protective band is added to the left boundary of the projected surface packed at the surface positions with position indices 0 - 5, and a sixth protective band is added to the right boundary of the projected surface packed at the surface positions with position indices 0 - 5, wherein the first protective band, the second protective band, the third protective band, the fourth protective band, the fifth protective band, and the sixth protective band have the same protective band size D (i.e., the same number of protective band samples). Specifically, the third protective band, the fourth protective band, the fifth protective band, and the sixth protective band serve as the boundaries of the cube-based projection layout. In addition, the width of a protective band region (which consists of two protective bands) inserted between two projected surfaces (which are packed at the surface positions with position indices 2 and 3) is equal to 2*D.

[0044] Since those skilled in the art can easily understand after reading the above paragraphs Figure 5 and Figure 6 the details of other protective band configurations shown, for the sake of brevity, they will not be further elaborated here. A projection-based frame of a cube mapping projection layout with protective bands (such as one of the cube mapping projection layouts shown Figures 4-6 ) can have a picture width excluding the protective band samples, and this picture width is equal to the picture width of another projection-based frame, and the other projection-based frame has a cube mapping projection layout without protective bands in the width direction, and / or can have a picture height excluding the protective band samples, and this picture height is equal to the picture height of another projection-based frame, and the other projection-based frame has a cube mapping projection layout without protective bands in the height direction.

[0045] Figures 7-9 For illustration, another specification of the protective bands packed in a general cube mapping projection or a hemispherical cube mapping projection is shown according to an embodiment of the present invention. Figures 4-6 The main difference between the shown protective band configuration and Figures 7-9 the shown protective band configuration is that a single protective band is added to an edge (such as a discontinuous edge or a continuous edge) between two adjacent projected surfaces packed in a cube-based projection layout. As shown in Figure 7As shown in the upper part, a protective band is inserted between the lower boundary of a first projection plane packed at a surface position with position index 2 and the upper boundary of a second projection plane packed at a surface position with position index 3, where the protective band has a protective band size D. If the lower boundary of the first projection plane is directly connected to the upper boundary of the second projection plane, an edge (such as a discontinuous edge or a continuous edge) exists between the first projection plane and the second projection plane. A protective band can be added to the edge between the first projection plane and the second projection plane. For example, regarding a cube in a 3D space, the lower boundary of the first projection plane (which is a square face of the cube) may or may not be connected to the upper boundary of the second projection plane (which is another square face of the cube); and regarding a cube-based projection layout on a 2D plane, the lower boundary of the first projection plane is parallel to the upper boundary of the second projection plane, and the protective band is located between the first projection plane and the second projection plane to isolate the lower boundary of the first projection plane from the upper boundary of the second projection plane, where the protective band is connected to the lower boundary of the first projection plane and also connected to the upper boundary of the second projection plane. Therefore, the width of a protective band region (which consists of a single protective band) inserted between the first projection plane (packed at the surface position with position index 2) and the second projection plane (packed at the surface position with position index 3) is equal to D.

[0046] As Figure 7 As shown in the lower part, a first protective band is added between the lower boundary of a projection plane packed at a surface position with position index 2 and the upper boundary of a projection plane packed at a surface position with position index 3, a second protective band is added to the upper boundary of a projection plane packed at a surface position with position index 0, a third protective band is added to the lower boundary of a projection plane packed at a surface position with position index 5, a fourth protective band is added to the left boundary of the projection planes packed at surface positions with position indices 0 - 5, and a fifth protective band is added to the right boundary of the projection planes packed at surface positions with position indices 0 - 5, where the first protective band, the second protective band, the third protective band, the fourth protective band, and the fifth protective band have the same protective band size D (i.e., the same number of protective band samples). Specifically, the second protective band, the third protective band, the fourth protective band, and the fifth protective band serve as the boundaries of the cube-based projection layout. In addition, the width of a protective band region (which consists of a single protective band) inserted between two projection planes (packed at the surface positions with position indices 2 and 3) is equal to D.

[0047] Since those skilled in the art can easily understand after reading the above paragraphs Figure 8 and Figure 9Details of other guard band configurations shown are not further elaborated here for the sake of brevity. A projection-based frame of a cube map projection layout with guard bands (e.g., Figures 7-9 one of the cube map projection layouts shown) can have a picture width that excludes guard band samples and is equal to the picture width of another projection-based frame, and the other projection-based frame has a modified cube map projection layout without guard bands, and / or can have a picture height that excludes guard band samples and is equal to the picture height of another projection-based frame, and the other projection-based frame has a modified cube map projection layout without guard bands.

[0048] As described above, chroma subsampling is a type of compression that reduces the color information in a signal to reduce bandwidth without significantly affecting picture quality. For example, the conversion circuit 114 can perform RGB-to-YCrCb conversion and chroma subsampling, and the video encoder 116 can encode the projection-based frame IMG in a 4:2:0 chroma format or a 4:2:2 chroma format into the bitstream BS for transmission. Regarding the guard band configuration, the conversion circuit 114 can impose a guard band size D on each guard band packed in the projection-based frame IMG, and the video decoder 122 can ignore / discard the bitstream that does not comply with these constraints. Assuming that the projection-based frame IMG is in a 4:2:0 chroma format or a 4:2:2 chroma format, the guard band size D of each guard band packed in the projection-based frame IMG can be constrained to an even number of luminance samples. The projection-based frame IMG in a 4:2:0 chroma format (or 4:2:2 chroma format) is encoded into the bitstream BS. Since the decoded frame IMG’ is a projection-based frame derived from the decoded bitstream BS, the decoded frame IMG’ is also in a 4:2:0 chroma format (or 4:2:2 chroma format), and the guard band size D of each guard band packed in the decoded frame IMG’ is also equal to an even number of luminance samples.

[0049] Video encoder 116 may signal syntax elements related to the guard band configuration of the projection-based frame IMG via the bitstream BS, where the guard band configuration includes the guard band size D. Thus, video decoder 122 may parse the syntax elements related to the guard band configuration from the bitstream BS. For example, the syntax element gcmp_guard_band_samples_minus1 is arranged to provide the size information of each guard band packed in the projection-based frame (IMG or IMG’) with a cube-based projection layout. For example, gcmp_guard_band_samples_minus1 plus 1 indicates the number of guard band samples (in terms of luma samples) used in the frame after cube mapping projection. When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), gcmp_guard_band_samples_minus1 plus 1 should correspond to an even number of luma samples. That is, if the 4:2:0 chroma format or 4:2:2 chroma format is used, gcmp_guard_band_samples_minus1 plus 1 should correspond to an even number of luma samples.

[0050] It should be noted that the proposed guard band size constraint can be applied to any projection layout with guard bands to ensure that when the projection-based frame is a 4:2:0 chroma format or a 4:2:2 chroma format, the guard band size of each guard band packed in the projection-based frame is constrained to an even number of luma samples.

[0051] In some embodiments of the present invention, when the projection-based frame IMG is chroma subsampled, conversion circuit 114 may impose constraints on the projection plane size, and video decoder 122 may ignore / discard the bitstream that does not follow these constraints. For example, when a projection-based frame (IMG or IMG’) is a 4:2:0 chroma format or a 4:2:2 chroma format, the width of one side of each projection plane packed in the projection-based frame (IMG or IMG’) is equal to an even number of luma samples. In another example, when a projection-based frame (IMG or IMG’) is a 4:2:0 chroma format, the height of one side of each projection plane packed in the projection-based frame (IMG or IMG’) is equal to an even number of luma samples. It should be noted that the proposed projection plane size constraint can be applied to any projection layout with / without guard bands.

[0052] Consider a first case where the constraint is applied to a face that is packed in a projection-based frame (IMG or IMG') and has a cube-based projection layout (e.g., a general CMP layout or a hemispherical CMP layout). When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), the width of each face packed in the projection-based frame (IMG or IMG') should correspond to an even number of luma samples. When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format), the height of each face packed in the projection-based frame (IMG or IMG') should correspond to an even number of luma samples.

[0053] Consider a second case where the constraint is applied to a face that is packed in a projection-based frame (IMG or IMG') and has a cube-based projection layout (e.g., a general CMP layout or a hemispherical CMP layout), and faceWidth and faceHeight represent the width and height of a square projection face included in the projection face packed in the cube-based projection layout. When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), faceWidth should be an integer multiple of 4 (in luma samples). When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format), faceHeight should be an integer multiple of 4 (in luma samples).

[0054] Consider a third case where the constraint is applied to a face that is packed in a projection-based frame (IMG or IMG') and has a cube-based projection layout (e.g., a general CMP layout or a hemispherical CMP layout), and faceWidth and faceHeight represent the width and height of a square projection face included in the projection face packed in the cube-based projection layout. If a horizontally packed hemispherical cube map layout is implemented, the following constraints apply:

[0055] When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), faceWidth should be an integer multiple of 4 (in luma samples); and

[0056] When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format), faceHeight should be an integer multiple of 2 (in luma samples).

[0057] If a vertically packed hemispherical cube map layout is implemented, the following constraints apply:

[0058] When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), faceWidth shall be an integer multiple of 2 (in units of luma samples); and

[0059] When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format), faceHeight shall be an integer multiple of 4 (in units of luma samples).

[0060] If the general cube map layout is applied, the following constraints apply:

[0061] When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), faceWidth shall be an integer multiple of 2 (in units of luma samples); and

[0062] When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format), faceHeight shall be an integer multiple of 2 (in units of luma samples).

[0063] Consider a fourth case where the constraints are applied to the faces that are packed in a projection-based frame (IMG or IMG’) and have a projection layout of a hemispherical cube map including one full face and four half faces (e.g., a horizontally packed hemispherical cube map projection layout or a vertically packed hemispherical cube map projection layout). When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), the width of each face packed in the hemispherical cube map projection layout shall be an integer multiple of 2 (in units of luma samples). When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format), the height of each face packed in the hemispherical cube map projection layout shall be an integer multiple of 2 (in units of luma samples).

[0064] In some embodiments of the present invention, when a projection-based frame IMG includes a plurality of projection planes packed in a projection layout having M projection plane columns and N projection plane rows (M and N are positive integers), the conversion circuit 114 can impose constraints on the projection size, and the video decoder 122 can ignore / discard a bitstream that does not comply with these constraints. For example, when a projection-based frame (IMG or IMG') includes MxN projection planes, a picture width cmpPicWidth excluding guard band samples is an integer multiple of M, and a picture height cmpPicHeight excluding guard band samples is an integer multiple of N. It should be noted that the proposed projection size constraints can be applied to any projection layout with / without guard bands.

[0065] Consider a first case where constraints are imposed on the faces packed in a projection-based frame (IMG or IMG') having a general cube map projection layout, where cmpPicWidth is the picture width excluding guard band samples and cmpPicHeight is the picture height excluding guard band samples. If the general cube map projection layout is set to a 1x6 layout as Figure 10 shown, cmpPicHeight should be an integer multiple of 6 (N = 6), and cmpPicWidth should be equal to cmpPicHeight / 6. If the general cube map projection layout is set to a 2x3 layout as Figure 11 shown, cmpPicWidth should be an integer multiple of 2 (M = 2) and cmpPicHeight should be an integer multiple of 3 (N = 3), and cmpPicWidth / 2 should be equal to cmpPicHeight / 3. If the general cube map projection layout is set to a 3x2 layout as Figure 12 shown, cmpPicWidth should be an integer multiple of 3 (M = 3) and cmpPicHeight should be an integer multiple of 2 (N = 2), and cmpPicWidth / 3 should be equal to cmpPicHeight / 2. If the general projection layout is set to a 6x1 layout as Figure 13 shown, cmpPicWidth should be an integer multiple of 6 (M = 6), and cmpPicWidth / 6 should be equal to cmpPicHeight.

[0066] Consider a second case where a projection-based frame IMG includes one full face and four half faces packed in a hemispherical cube map projection layout. The conversion circuit 114 can impose constraints on the projection size, and the video decoder 122 can ignore / discard a bitstream that does not comply with these constraints. For example, a projection-based frame (IMG or IMG’) uses a hemispherical cube map projection layout, and one of the picture width excluding guard band samples and the picture height excluding guard band samples is equal to an integer multiple of 6. When the hemispherical cube map projection layout is a horizontally packed hemispherical cube map projection layout, cmpPicWidth is the picture width excluding guard band samples, and cmpPicHeight is the picture height excluding guard band samples. cmpPicWidth should be an integer multiple of 6, and cmpPicWidth / 3 should be equal to cmpPicHeight. When the hemispherical cube map projection layout is a vertically packed hemispherical cube map projection layout, cmpPicWidth is the picture width excluding guard band samples, and cmpPicHeight is the picture height excluding guard band samples. cmpPicHeight should be an integer multiple of 6, and cmpPicWidth should be equal to cmpPicHeight / 3.

[0067] Those skilled in the art will have no difficulty in observing that various modifications and changes can be made to the apparatus and method, yet still retain the teachings of the present invention. Therefore, the interpretation of the foregoing disclosure should be limited only to the boundaries of the appended patent claims.

Claims

1. A video decoding method, comprising: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including at least one projection plane and at least one guard band packed in a projection layout, and at least a portion of a spherical 360-degree content is projected and mapped onto the at least one projection plane; wherein the decoded frame is in a 4:2:0 chroma format or a 4:2:2 chroma format, and a guard band size of each of the at least one guard band is equal to an even number of luminance samples, and the video decoder discards the bitstream that does not conform to the guard band size constraint.

2. A video decoding method, comprising: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including a plurality of projection planes packed in a projection layout, and at least a portion of a spherical 360-degree content is projected and mapped onto the plurality of projection planes; wherein the decoded frame is in a 4:2:0 chroma format or a 4:2:2 chroma format, and a width of each of the plurality of projection planes is equal to an even number of luminance samples, and the video decoder discards the bitstream that does not conform to the plane width constraint.

3. The video decoding method according to claim 2, wherein: the projection is a cube-based projection, and the projection layout is a cube-based projection layout.

4. The video decoding method according to claim 3, wherein: the cube-based projection is a hemispherical cube map projection, and the cube-based projection layout is a hemispherical cube map projection layout.

5. The video decoding method according to claim 3, wherein: it includes that a width of each square projection plane of the plurality of projection planes is equal to an integer multiple of 4 luminance samples.

6. The video decoding method according to claim 5, wherein: the cube-based projection is a hemispherical cube map projection, and the cube-based projection layout is a horizontally packed hemispherical cube map projection layout.

7. The video decoding method according to claim 3, wherein: the cube-based projection is a hemispherical cube map projection, and the cube-based projection layout is a vertically packed hemispherical cube map projection layout.

8. The video decoding method according to claim 3, wherein: the cube-based projection is a cube map projection, and the cube-based projection layout is a cube map projection layout.

9. A video decoding method, comprising: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame including a plurality of projection planes packed in a projection layout, and at least a portion of a spherical 360-degree content is projected and mapped onto the plurality of projection planes; wherein the decoded frame is in a 4:2:0 chroma format, and a height of each of the plurality of projection planes is equal to an even number of luminance samples, and the video decoder discards the bitstream that does not conform to the plane height constraint.

10. The video decoding method according to claim 9, wherein, the projection is a projection based on a cube, and the projection layout is a projection layout based on a cube.

11. The video decoding method according to claim 10, wherein, the cube-based projection is a hemispherical cube mapping projection, and the cube-based projection layout is a hemispherical cube mapping projection layout.

12. The video decoding method according to claim 11, wherein, the hemispherical cube mapping projection layout is a horizontally-packaged hemispherical cube mapping projection layout, and a height of one side of a square projection plane among the plurality of projection planes is equal to an even number of luminance samples.

13. The video decoding method according to claim 10, wherein, a height of one side of each square projection plane among the plurality of projection planes is equal to an integer multiple of 4 luminance samples.

14. The video decoding method according to claim 13, wherein, the cube-based projection is a hemispherical cube mapping projection, and the cube-based projection layout is a vertically-packaged hemispherical cube mapping projection layout.

15. The video decoding method according to claim 10, wherein, the cube-based projection is a cube mapping projection, and the cube-based projection layout is a cube mapping projection layout.

Citation Information

Patent Citations

  • Method for processing projection-based frame that includes at least one projection face packed in 360-degree virtual reality projection layout

    CN109906468A

  • Video encoding method with syntax element signaling of guard band configuration of projection-based frame and associated video decoding method and apparatus

    CN113906755A

  • Systems and methods for signaling information associated with a constituent picture

    WO2019065587A1