Video encoding and decoding method and device
By adopting a cube-based projection layout and inserting guard bands in 360-degree virtual reality video coding, the low coding efficiency and artifact problems caused by a fixed projection layout are solved, achieving more efficient video coding and better image quality.
Patent Information
- Application Number
- CN202080040638.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-01
- Filing Date
- 2020-07-02
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2040-07-02
AI Technical Summary
When processing 360-degree virtual reality content, existing video coding technologies have limited flexibility in fixed projection layouts, resulting in low coding efficiency and possible artifacts.
A cube-based projection layout is adopted, and guard bands are inserted between the projection planes to reduce artifacts. The guard band configuration is indicated by syntax elements sent through the bitstream, allowing flexible adjustment of the projection layout.
It improves the flexibility of video encoding, reduces artifacts, and improves encoding efficiency and image quality.
Smart Images

Figure CN113906755B_ABST
Abstract
Description
[0001] Related references
[0002] This application claims priority to U.S. Provisional Application No. 62 / 869,627, filed on July 2, 2019, U.S. Provisional Application No. 62 / 870,139, filed on July 3, 2019, U.S. Provisional Application No. 62 / 903,056, filed on September 20, 2019, and U.S. Provisional Application No. 62 / 954,814, filed on December 30, 2019. The entire contents of the related applications, including U.S. Provisional Application No. 62 / 869,627, U.S. Provisional Application No. 62 / 870,139, U.S. Provisional Application No. 62 / 903,056, and U.S. Provisional Application No. 62 / 954,814, are incorporated herein by reference in their entirety. Technical Field
[0003] The present invention relates to video encoding and video decoding, and more particularly to a video encoding method for transmitting syntax elements with a guard band configuration based on a projection frame, and a related video decoding method and apparatus. Background Art
[0004] Virtual reality (VR) with a head-mounted display (HMD) is associated with a variety of applications. The ability to display content with a wide field of view to the user can be used to provide an immersive visual experience. The real-world environment must be captured from all directions to produce omnidirectional video corresponding to the field of view. With the advancement of camera equipment and HMDs, the delivery of VR content may soon become a bottleneck due to the high bit rate required to present this 360-degree content. When the resolution of omnidirectional video is 4K or higher, data compression / encoding is crucial to reduce the bit rate.
[0005] Generally speaking, omnidirectional video corresponding to a sphere is converted into frames having 360-degree image content represented by projection surfaces arranged in a 360-degree virtual reality (360VR) projection layout, and the resulting frames are then encoded into a bitstream for transmission. If the configuration of the adopted 360VR projection layout is fixed and no adjustment is allowed, the video encoder has limited flexibility in encoding the 360-degree image content. Therefore, a flexible design that allows for determining / selecting the projection-based guard band configuration of frames and the signal syntax elements associated with the projection-based guard band configuration is required. Summary of the Invention
[0006] One of the objectives of the present invention is to provide a video encoding method for transmitting syntax elements with a guard band configuration based on a projection frame, and a related video decoding method and apparatus.
[0007] According to a first aspect of the present invention, an exemplary video decoding method is disclosed. The exemplary video decoding method includes: decoding a portion of a bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame, which includes a projection surface packed in a cube-based projection layout, and 360-degree content of at least a portion of a sphere is mapped to the projection surface by the cube-based projection; and parsing at least one syntax element from the bitstream, wherein the at least one syntax element indicates a guard band configuration of the projection-based frame.
[0008] According to a second aspect of the present invention, an exemplary electronic device is disclosed. The exemplary electronic device includes decoding circuitry. The decoding circuitry is configured to decode a portion of a bitstream to generate a decoded frame and parse at least one syntax element from the bitstream. The decoded frame is a projection-based frame comprising a projection surface packed in a cube-based projection layout. At least a portion of 360-degree content of a sphere is mapped to the projection surface using the cube-based projection. The at least one syntax element indicates a guard band configuration for the projection-based frame.
[0009] According to a third aspect of the present invention, an exemplary video encoding method is disclosed. The exemplary video encoding method includes: encoding a projection-based frame to generate a portion of a bitstream, wherein at least a portion of 360-degree content of a sphere is mapped to a projection surface via a cube-based projection, and the projection-based frame has the projection surface packed in a cube-based projection layout; and transmitting at least one syntax element via the bitstream, wherein the at least one syntax element indicates a guard band configuration of the projection-based frame.
[0010] These and other objects of the present invention will no doubt become apparent to those of ordinary skill in the art after reading the following detailed description of the preferred embodiments illustrated in the various figures and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 A diagram illustrating a 360-degree virtual reality (360VR) system according to an embodiment of the present invention.
[0012] Figure 2 Diagram showing cube-based projection according to an embodiment of the present invention.
[0013] Figure 3 Diagram showing another cube-based projection according to an embodiment of the present invention.
[0014] Figure 4-Figure 6 A diagram showing a specification of a combination of two syntax elements according to an embodiment of the present invention.
[0015] Figure 7-Figure 9 A diagram showing another specification of a combination of two syntax elements according to an embodiment of the present invention. DETAILED DESCRIPTION
[0016] Throughout the following description and claims, specific terms are used to refer to specific components. As one skilled in the art will appreciate, electronic device manufacturers may refer to a component by different names. This document does not intend to distinguish between components that differ in name but function. In the following description and claims, the terms "including" and "comprising" are used in an open-ended manner and, therefore, should be interpreted to mean "including, but not limited to..." . Additionally, the term "coupled" is intended to mean either an indirect or direct electrical connection. Thus, if one device is coupled to another device, that connection may be through a direct electrical connection or through an indirect electrical connection via other devices and connections.
[0017] Figure 1 A diagram illustrates a 360-degree virtual reality (360VR) system according to an embodiment of the present invention. The 360VR system 100 includes a source electronic device 102 and a destination electronic device 104. The source electronic device 102 includes a video capture device 112, a conversion circuit 114, and a video encoder 116. For example, the video capture device 112 may be an omnidirectional camera. The conversion circuit 114 generates a projection-based frame IMG having a 360-degree virtual reality (360VR) projection layout L_VR based on an omnidirectional video frame S_IN corresponding to a sphere, wherein the omnidirectional video frame S_IN contains 360-degree content of the sphere. The video encoder 116 is an encoding circuit that encodes the projection-based frame IMG (having projection surfaces packaged in the 360VR projection layout L_VR) to generate a portion of a bitstream BS, and outputs the bitstream BS to the destination electronic device 104 via a transmission device 103, such as a wired / wireless communication link or a storage medium.
[0018] The destination electronic device 104 may be a head mounted display (HMD) device. Figure 1As shown, the destination electronic device 104 includes a video decoder 122, graphics rendering circuitry 124, and a display device 126. The video decoder 122 is a decoding circuit that receives a bitstream BS from a transmission device 103 (e.g., a wired / wireless communication link or a storage medium) and decodes a portion of the received bitstream BS to generate a decoded frame IMG'. In this embodiment, the projection-based frame IMG encoded by the video encoder 116 has a 360VR projection layout L_VR. Therefore, after the portion of the bitstream BS is decoded by the video decoder 122, the decoded frame (i.e., the reconstructed frame) IMG' has the same 360VR projection layout L_VR. In other words, the decoded frame IMG' is also a projection-based frame, with its projection planes packed within the 360VR projection layout L_VR. The graphics rendering circuitry 124 is configured to drive the display device 126 to display the image content of the viewport area selected by the user. The graphics rendering circuitry 124 may include a conversion circuit 125, which is configured to process a portion of the image content carried by the decoded frame IMG' to obtain pixel data associated with the image content of the selected viewport area.
[0019] In this embodiment, the 360VR projection layout L_VR is a cube-based projection layout. Therefore, at least a portion (i.e., part or all) of the spherical 360-degree content is mapped to the projection surface through a cube-based projection, and the projection surfaces from different faces of a three-dimensional object (e.g., a cube or a hemispherical cube) are packed into the two-dimensional cube-based projection layout adopted by the projection-based frame IMG / decoded frame IMG'.
[0020] In one embodiment, a cube-based projection with six square projection surfaces representing 360° x 180° omnidirectional video (i.e., all 360-degree content of a sphere) may be employed. With respect to the conversion circuitry 114 of the source electronic device 102, the cube-based projection is used to generate the square projection surfaces of a cube in three-dimensional (3D) space. Figure 2 A diagram illustrating a cube-based projection according to an embodiment of the present invention is shown. The entire 360-degree content on the sphere 200 is projected onto the six square faces of the cube 201, including the top face (labeled as "top"), the bottom face (labeled as "bottom"), the left face (labeled as "left"), the front face (labeled as "front"), the right face (labeled as "right"), and the back face (labeled as "back"). Figure 2 As shown, the image content of the north pole region of the sphere 200 is projected onto the top surface “top”, the image content of the south pole region of the sphere 200 is projected onto the bottom surface “bottom”, and the image content of the equatorial region of the image content sphere 200 is projected onto the left surface “left”, the front surface “front”, the right surface “right” and the back surface “back”.
[0021] In the 3D space defined by the x-axis, y-axis, and z-axis, each point on the six projection surfaces is located at (x,y,z), where x,y,z∈[-1,1]. Figure 2 In the example shown, the front face "front" is on the x-plane of x=1, the back face "back" is on the x-plane of x=-1, the top face "top" is on the z-plane of z=1, the bottom face "bottom" is on the z-plane of z=-1, the left face "left" is on the y-plane of y=1, and the right face "right" is on the y-plane of y=-1. In an alternative design, the front face "front" may be on the x-plane of x=1, the back face "back" may be on the x-plane of x=-1, the top face "top" may be on the y-plane of y=1, the bottom face "bottom" may be on the y-plane of y=-1, the right face "right" may be on the z-plane of z=1, and the left face "left" may be on the z-plane of z=-1.
[0022] The forward transform is used to transform from 3D space (x, y, z) to a 2D plane (u, v). Therefore, the top face "top", bottom face "bottom", left face "left", front face "front", right face "right", and back face "back" of the cube 201 in 3D space are transformed into the top face (labeled "2"), bottom face (labeled "3"), left face (labeled "5"), front face (labeled "0"), right face (labeled "4"), and back face (labeled "1") on a two-dimensional plane. Each face is located on a 2D plane defined by the u-axis and the v-axis, and each point is located at (u, v).
[0023] The inverse transformation is used to transform from the 2D plane (u, v) to the 3D space (x, y, z). Therefore, the top face (marked as "2"), bottom face (marked as "3"), left face (marked as "5"), front face (marked as "0"), right face (marked as "4"), and back face (marked as "1") on the two-dimensional plane are converted to the top face "top", bottom face "bottom", left face "left", front face "front", right face "right", and back face "back" of the cube 201 in the 3D space.
[0024] The inverse transform may be used by the conversion circuit 114 of the source electronic device 102 to generate the top face “2”, the bottom face “3”, the left face “5”, the front face “0”, the right face “4”, and the back face “1”. The top face “2”, the bottom face “3”, the left face “5”, the front face “0”, the right face “4”, and the back face “1” on the two-dimensional plane are packed to form a projection-based frame IMG to be encoded by the video encoder 116.
[0025] The video decoder 122 receives a bitstream BS from the transmission device 103 and decodes a portion of the received bitstream BS to generate a decoded frame IMG' identical to the projection layout L_VR used by the encoder. For the conversion circuit 125 of the destination electronic device 104, a forward transform is used to transform from 3D space (x, y, z) to a 2D plane (u, v) to determine the pixel values of pixels in any of the top plane "top", bottom plane "bottom", left plane "left", front plane "front", right plane "right", and back plane "back". Alternatively, an inverse transform is used to transform from the 2D plane (u, v) to 3D space (x, y, z) to remap the sample positions of the projection-based frame to the sphere.
[0026] As described above, the top surface "2", the bottom surface "3", the left surface "5", the front surface "0", the right surface "4", and the back surface "1" are compressed to form the projection-based frame IMG. For example, the conversion circuit 114 can select a packaging type so that the projection-based frame IMG can have projection image data arranged in the cube-based projection layout 202. For another example, the conversion circuit 114 can select another packaging type so that the projection-based frame IMG can have projection image data arranged in the cube-based projection layout 204, which is different from the cube-based projection layout 202.
[0027] In another embodiment, a cube-based projection with five projection planes (including one full plane and four half planes) representing 180°x180° omnidirectional video (i.e., a portion of a 360-degree content of a sphere) may be employed. With respect to the conversion circuitry 114 of the source electronic device 102, the cube-based projection is used to generate one full plane and four half planes of a cube in 3D space. Figure 3 A diagram illustrating another cube-based projection according to an embodiment of the present invention is shown. Only half of the 360-degree content on sphere 200 is projected onto faces of cube 201, including the top half face (labeled "top_H"), the bottom half face (labeled "bottom_H"), the left half face (labeled "left_H"), the right half face (labeled "right_H"), and the right half face (labeled "right_H"). In this example, a hemispherical cube (e.g., half of cube 201) is used for hemispherical cubemap projection, where the hemisphere (e.g., half of sphere 200) is inscribed in the hemispherical cube (e.g., half of cube 201). As Figure 3 As shown, half of the image content of the North Pole region of the sphere 200 is projected onto the top half surface "Top_H", half of the image content of the South Pole region of the sphere 200 is projected onto the bottom half surface "Bottom_H", and half of the image content of the equatorial region of the sphere 200 is projected onto the left half surface "Left_H", the full front surface "Front", and the right half surface "Right_H".
[0028] In the 3D space defined by the x-axis, y-axis, and z-axis, each point on the five projection surfaces is located at (x,y,z), where x,y,z∈[-1,1]. Figure 3 In the example shown, the full face “Front” is on the x-plane with x=1, the top half face “Top_H” is on the z-plane with z=1, the bottom half face “Bottom_H” is on the z-plane with z=-1, the left half face “Left_H” is on the y-plane with y=1, and the right half face “Right_H” is on the y-plane with y=-1. In an alternative design, the full face “Front” can be on the x-plane with x=1, the top half face “Top_H” can be on the y-plane with y=1, the bottom half face “Bottom_H” can be on the y-plane with y=-1, the right half face “Right_H” can be on the z-plane with z=1, and the left half face “Left_H” can be on the z-plane with z=-1.
[0029] The forward transform is used to transform from 3D space (x, y, z) to a 2D plane (u, v). Therefore, the top half face "top_H", the bottom half face "bottom_H", the left half face "left_H", the full face "full face" and the right half face "right_H" of the cube 201 in 3D space are transformed into the top half face (marked as "2"), the bottom half face (marked as "3"), the left half face (marked as "5"), the full face (marked as "0") and the right half face (marked as "4") on the 2D plane. Each face is located on a 2D plane defined by the u-axis and the v-axis, and each point therein is located at (u, v). In addition, the size of the full face (marked as "0") is twice the size of the top half face (marked as "2"), the bottom half face (marked as "3"), the left half face (marked as "5") and the right half face (marked as "4").
[0030] The inverse transform is used to transform from the 2D plane (u, v) to the 3D space (x, y, z). Thus, the top half face (labeled as "2"), the bottom half face (labeled as "3"), the left half face (labeled as "5"), the right half face (labeled as "0"), and the right half face (labeled as "4") on the 2D plane are transformed into the top half face "top_H", the bottom half face "bottom_H", the left half face "left_H", the right half face "right_H" of the cube 201 in the 3D space.
[0031] The inverse transform may be used by the conversion circuitry 114 of the source electronic device 102 to generate a top half plane “2,” a bottom half plane “3,” a left half plane “5,” a full square plane “0,” and a right half plane “4.” The top half plane “2,” the bottom half plane “3,” the left half plane “5,” the full square plane “0,” and the right half plane “4” on the 2D plane are packed to form a projection-based frame IMG to be encoded by the video encoder 116.
[0032] The video decoder 122 receives a bitstream BS from the transmission device 103 and decodes a portion of the received bitstream BS to generate a decoded frame IMG' identical to the projection layout L_VR used by the encoder. For the conversion circuit 125 of the destination electronic device 104, a forward transform is used to transform from 3D space (x, y, z) to a 2D plane (u, v) to determine the pixel values of pixels on any of the top half plane "top_H", the bottom half plane "bottom_H", the left half plane "left_H", the full front plane "front", and the right half plane "right_H". Alternatively, an inverse transform is used to transform from the 2D plane (u, v) to 3D space (x, y, z) to remap the sample positions of the projection-based frame onto the sphere.
[0033] As described above, the top half face "2", the bottom half face "3", the left half face "5", the full face "0" and the right half face "4" are packed to form the projection-based frame IMG. For example, the conversion circuit 114 can select a packing type so that the projection-based frame IMG can have projection image data arranged in the cube-based projection layout 302. For another example, the conversion circuit 114 can select another packing type so that the projection-based frame IMG can have projection image data arranged in the cube-based projection layout 304, which is different from the cube-based projection layout 302. In this embodiment, the front face is selected as the full face packed in the cube-based projection layout 302 / 304. In fact, the full face packed in the cube-based projection layout 302 / 304 can be any of the top face, bottom face, front face, back face, left face and right face, and the four half faces filled in the cube-based projection layout 302 / 304 depend on the selection of the full face.
[0034] about Figure 2 In the embodiment shown, the projection surface adopts a conventional CMP layout without a guard band (or filler) 202 / 204. Figure 3In the illustrated embodiment, the projection surface is padded in a hemispherical CMP layout without a guard band (or padding) 302 / 304. However, the encoded and decoded projection-based frame IMG may produce artifacts due to discontinuous layout boundaries of the CMP layout (which may be a conventional CMP layout or a hemispherical CMP layout) and / or discontinuous edges of the CMP layout (which may be a conventional CMP layout or a hemispherical CMP layout). For example, the CMP layout without a guard band (or padding) has a top discontinuous layout boundary, a bottom discontinuous layout boundary, a left discontinuous layout boundary, and a right discontinuous layout boundary. In addition, there is at least one image content discontinuous edge between two adjacent projection surfaces in the CMP layout without a guard band (or padding). Taking the cube-based projection layout 202 / 204 as an example, a discontinuous edge exists between a surface boundary of the bottom surface "3" and a surface boundary of the left surface "5", a discontinuous edge exists between a surface boundary of the back surface "1" and a surface boundary of the front surface "0", and a discontinuous edge exists between a surface boundary of the top surface "2" and a surface boundary of the right surface "4". Taking the cube-based projection layout 302 / 304 as an example, a discontinuous edge exists between a face boundary of the bottom face "3" and a face boundary of the left face "5", and a discontinuous edge exists between a face boundary of the right face "4" and a face boundary of the top face "2".
[0035] To address this issue, the 360VR projection layout L_VR can be configured using a cube-based projection layout with guard bands (or padding). For example, around layout boundaries and / or discontinuous edges, additional guard bands, such as those generated by pixel padding, can be inserted to reduce seam artifacts. Alternatively, around layout boundaries and / or continuous edges, additional guard bands, such as those generated by pixel padding, can be inserted.
[0036] In this embodiment, the conversion circuit 114 determines a guard band configuration for a projection-based frame IMG that is formed from a cube-based projection (e.g., Figure 2 The regular cubemap projection shown or Figure 3 The video encoder 116 transmits the syntax element SE associated with the guard band configuration of the projection-based frame IMG via the bitstream BS. Therefore, the video decoder 122 can parse the syntax element SE associated with the guard band configuration from the bitstream BS.
[0037] In order to better understand the technical features of the present invention, an exemplary syntax transmission method is described below. The video encoder 116 can adopt the proposed syntax transmission method to transmit a syntax element SE, which indicates the configuration information of the guard band added by the conversion circuit 114, and the video decoder 122 can parse the syntax element SE' transmitted by a proposed syntax transmission method adopted by the video encoder 116, and can provide the parsed syntax element SE' to the graphics rendering circuit 124 (particularly, the conversion circuit 125), so that the graphics rendering circuit 124 (particularly, the conversion circuit 125) is informed of the guard band configuration information added by the conversion circuit 114. In this way, the conversion circuit 125 can refer to the guard band configuration information when determining the image content of the viewport area selected by the user and perform the conversion correctly. Ideally, the syntax element SE encoded into the bitstream BS by the video encoder 116 is the same as the syntax element SE' parsed from the bitstream BS by the video decoder 122.
[0038] It should be noted that the descriptors in the following exemplary syntax table specify the parsing process for each syntax element. For example, a syntax element can be encoded using a fixed-length codec (e.g., u(n)). Taking the descriptor u(n) as an example, it describes an unsigned integer using n bits. However, this is for illustrative purposes only and is not intended to limit the present invention. In practice, syntax elements can be encoded using fixed-length encoding (e.g., f(n), i(n), or u(n)) and / or variable-length encoding (e.g., ce(v), se(v), or ue(v)). The descriptor f(n) represents a fixed-pattern bit string, written with n bits written left-bit first (from left to right). The descriptor i(n) represents an n-bit signed integer. The descriptor u(n) represents an n-bit unsigned integer. The descriptor ce(v) represents a context-adaptive variable-length entropy coded syntax element with the left bit first. The descriptor se(v) represents an Exp-Golomb coded syntax element with a signed integer with the left bit first. The syntax element ue(v) represents an unsigned integer Exp-Golomb coded syntax element with the left bit first.
[0039] According to the proposed syntax sending method, the following syntax table can be adopted.
[0040]
[0041] The syntax element gcmp_guard_band_flag is arranged to indicate whether a projection-based frame (e.g., IMG or IMG') contains at least one guard band. If the syntax element gcmp_guard_band_flag is equal to 0, it indicates that the coded picture does not contain a guard band area. If the syntax element gcmp_guard_band_flag is equal to 1, it indicates that the coded picture contains a guard band area of a size specified by the syntax element gcmp_guard_band_samples_minus1.
[0042] The syntax element gcmp_guard_band_boundary_exterior_flag is arranged to indicate whether at least one guard band packed in a projection based frame (eg, IMG or IMG') includes a guard band serving as a boundary of a cube based projection layout.
[0043] The syntax element gcmp_guard_band_samples_minus1 is arranged to provide size information for each guard band packed in a projection-based frame (e.g., IMG or IMG'). For example, gcmp_guard_band_samples_minus1 plus 1 specifies the number of guard band samples used in the cubemap projection picture, in units of luma samples. When the parameter ChromaFormatIdc is equal to 1 (4:2:0 chroma format) or 2 (4:2:2 chroma format), gcmp_guard_band_samples_minus1 plus 1 corresponds to an even number of luma samples. That is, when using the 4:2:0 chroma format or the 4:2:2 chroma format, gcmp_guard_band_samples_minus1 plus 1 corresponds to an even number of luma samples.
[0044] Figure 4-6A diagram illustrating a specification of a combination of two syntax elements, gcmp_packing_type and gcmp_guard_band_boundary_exterior_flag, according to an embodiment of the present invention. In this example, the syntax element gcmp_packing_type specifies the packing type for packing projection faces in a cube-based projection layout, and further specifies a predetermined arrangement of position indices assigned to face positions under the selected packing type. When the value of gcmp_packing_type is in the range of 0 to 3 (inclusive), conventional cubemap packing with six faces is used, where each packing type is associated with six face positions, each of which is assigned position indices {0, 1, 2, 3, 4, 5}. When gcmp_packing_type is 4 or 5, hemispherical cubemap packing with one full face and four half faces is used, where each packing type is associated with five face positions, each of which is assigned position indices {0, 1, 2, 3, 4}. The value of gcmp_packing_type should be in the range of 0 to 5 (inclusive). Other values of gcmp_packing_type are reserved for future use. In addition, the syntax element gcmp_face_index[i] (not shown) may specify the face index of the position index i under the packing type specified by the syntax element gcmp_packing_type.
[0045] Taking conventional cubemap projection as an example, the front face may be assigned face index gcmp_face_index[i]==0, the back face may be assigned face index gcmp_face_index[i]==1, the top face may be assigned face index gcmp_face_index[i]==2, the bottom face may be assigned face index gcmp_face_index[i]==3, the right face may be assigned face index gcmp_face_index[i]==4, and the left face may be assigned face index gcmp_face_index[i]==5. When the syntax element gcmp_packing_type is set to 0, 1, 2, or 3, the syntax element gcmp_face_index[i] specifies the projection face (e.g., Figure 2 The face index of the projection surface is front "0", back "1", top "2", bottom "3", right "4" or left "5" as shown, where the projection surface with the face index specified by the syntax element gcmp_face_index[i] is packed at the face position with position index i under the selected packing type.
[0046] Taking the hemispherical cubemap projection as an example, the full front face can be assigned face index gcmp_face_index[i]==0, the upper half face can be assigned face index gcmp_face_index[i]==2, the lower half face can be assigned face index gcmp_face_index[i]==3, the right half face can be assigned face index gcmp_face_index[i]==4, and the left half face can be assigned face index gcmp_face_index[i]==5. When the syntax element gcmp_packing_type is set to 4 or 5, the syntax element gcmp_face_index[i] (not shown) specifies the face index of the projection face (e.g. Figure 3 front face "0", top face "2", bottom face "3", right face "4" or left face "5" in the selected grid), where the projection face with the face index specified by the syntax element gcmp_face_index[i] is packed at the face position with position index i under the selected packing type.
[0047] Since the present invention focuses on guard band syntax transmission, further description of the packing of projection surfaces based on cube projection is omitted here for the sake of brevity.
[0048] like Figure 4As shown, in the case where the syntax element gcmp_guard_band_flag is set to 1, when the syntax element gcmp_packing_type is set to 0 and the syntax element gcmp_guard_band_boundary_exterior_flag is set to 0, a first guard band is added to the bottom boundary of the first projected surface packed at the surface position of position index i=2, and a second guard band is added to the top boundary of the second projected surface packed at the surface position of position index i=3, wherein the first guard band and the second guard band have the same guard band size D specified by gcmp_guard_band_samples_minus1 plus 1. If the bottom boundary of the first projected surface is directly connected to the top boundary of the second projected surface, an edge (e.g., a discontinuous edge or a continuous edge) exists between the first projected surface and the second projected surface. A guard band may be added to the edge between the first projected surface and the second projected surface. For example, for a cube in 3D space, the bottom boundary of the first projection plane (a square face of the cube) may be connected to the top boundary of the second projection plane or not (this is another aspect of the cube); for a projection layout based on a cube on a 2D plane, the bottom boundary of the first projection plane is parallel to the top boundary of the second projection plane, and the first guard band and the second guard band are both between the first projection plane and the second projection plane to isolate the bottom boundary of the first projection plane from the top boundary of the second projection plane, wherein the first guard band connects the bottom boundary of the first projection plane and the second guard band, and the second guard band connects the first guard band and the top boundary of the second projection plane. In addition, the width of a guard band area (composed of the first guard band and the second guard band, inserted between the first projection plane (packed at the face position with position index i=2) and the second projection plane (packed at the face position with position index i=3)) is equal to 2*D.
[0049] like Figure 4As shown, in the case where the syntax element gcmp_guard_band_flag is set to 1, when the syntax element gcmp_packing_type is set to 0 and the syntax element gcmp_guard_band_boundary_exterior_flag is set to 1, the first guard band is added to the bottom boundary of the projected surface packed at the surface position of position index i=2, the second guard band is added to the top boundary of the projected surface packed at the surface position of position index i=3, and the third guard band is added to the top boundary of the projected surface packed at the surface position of position index i=0. The top surface boundary of the projected surface packed at the surface position of position index i=5, the fourth guard band is added to the bottom surface boundary of the projected surface packed at the surface position of position index i=5, the fifth guard band is added to the multiple left surface boundaries of the multiple projected surfaces packed at the multiple surface positions of position index i=0-5, and the sixth guard band is added to the multiple right surface boundaries of the multiple projected surfaces packed at the multiple surface positions of position index i=0-5, wherein the first guard band, the second guard band, the third guard band, the fourth guard band, the fifth guard band, and the sixth guard band have the same guard band size D as specified by gcmp_guard_band_samples_minus1 plus 1. Specifically, the third guard band, the fourth guard band, the fifth guard band, and the sixth guard band serve as boundaries of the cube-based projection layout. In addition, the width of a guard band area (consisting of two guard bands) inserted between two projected surfaces (packed at the surface positions of position index i=2 and i=3) is equal to 2*D.
[0050] Since those skilled in the art can easily understand after reading the above paragraphs Figure 5 and Figure 6 For the sake of brevity, details of the other guard band configurations shown are not repeated here.
[0051] As described above, the syntax element gcmp_guard_band_samples_minus1 is arranged to provide size information of each guard band packed in a projection-based frame (e.g., IMG or IMG'). In some embodiments of the present invention, the decoder-side conversion circuitry 125 may refer to the syntax element gcmp_guard_band_samples_minus1 parsed from the bitstream BS to Figures 4 to 6 Calculations are performed for a specific application under one of the guard band configurations shown in .
[0052] For example, the size information specified by the syntax element gcmp_guard_band_samples_minus1 may relate to the calculation of the size of each projection face packed in a cube-based projection layout with guard bands. The input to this process may include the width pictureWidth and height pictureHeight of the projection-based frame (e.g., decoded frame IMG'). The output of this process may include the width faceWidth and height faceHeight of the projection face packed in the projection-based frame (e.g., decoded frame IMG').
[0053] Outputs can be exported as follows:
[0054]
[0055]
[0056]
[0057] As another example, the size information specified by the syntax element gcmp_guard_band_samples_minusl may involve a conversion from sample positions within the decoded frame IMG' to local sample positions within one of the projection planes packed in a cube-based projection layout with a guard band. The input to this process may include the width faceWidth and height faceHeight of the projection plane, and may also include the sample positions (hPos, vPos) of the projection-based frame (e.g., the decoded frame IMG'). The output of this process may include the local sample positions (hPosFace, vPosFace) of the projection plane packed in the projection-based frame (e.g., the decoded frame IMG').
[0058] Outputs can be exported as follows:
[0059]
[0060]
[0061] Figure 7-9 A diagram showing another specification of a combination of two syntax elements gcmp_packing_type and gcmp_guard_band_boundary_exterior_flag according to an embodiment of the present invention. Figures 4 to 6 Same as the example shown in . Figure 4-6, the syntax element gcmp_packing_type specifies the packing type for packing of projection faces in a cube-based projection layout, and further specifies a predetermined configuration of position indices assigned to face positions under the selected packing type. Thus, when the value of gcmp_packing_type is in the range of 0 to 3 (inclusive), conventional cubemap packing with six faces is used, where each packing type is associated with six face positions assigned with position indices {0, 1, 2, 3, 4, 5}, respectively. When gcmp_packing_type is 4 or 5, hemispherical cubemap packing with one full face and four half faces is used, where each packing type is associated with five face positions assigned with position indices {0, 1, 2, 3, 4}, respectively. In addition, the syntax element gcmp_face_index[i] (not shown) can specify the face index of position index i under the packing type specified by the syntax element gcmp_packing_type.
[0062] Since the present invention focuses on guard band syntax transmission, further description of the packing of projection surfaces based on cube projection is omitted here for the sake of brevity.
[0063] Figure 4-6 The guard band configuration shown in Figure 7-9 The main difference between the guard band configurations shown in is that a single guard band is added at the edge (e.g., discontinuous edge or continuous edge) between two adjacent projection faces packed in a cube-based projection layout with a gcmp_packing_type value in the range of 0 to 3. Figure 7As shown, when the syntax element gcmp_guard_band_flag is set to 1, when the syntax element gcmp_packing_type is set to 0 and the syntax element gcmp_guard_band_boundary_exterior_flag is set to 0, a guard band is inserted between the bottom boundary of the first projected surface (packed at the surface position with surface index i=2) and the top boundary of the second projected surface (packed at the surface position with surface index i=3), wherein the guard band has a guard band size D specified by gcmp_guard_band_samples_minus1 plus 1. If the bottom boundary of the first projected surface is directly connected to the top boundary of the second projected surface, there is an edge (e.g., a discontinuous edge or a continuous edge) between the first projected surface and the second projected surface. A guard band is added to the edge between the first projected surface and the second projected surface. For example, for a cube in 3D space, the bottom boundary of a first projection surface (a square surface of the cube) may be connected to or not connected to the top boundary of a second projection surface (another surface of the cube); for a cube-based projection layout on a 2D plane, the bottom boundary of the first projection surface is parallel to the top boundary of the second projection surface, and a guard band is located between the first projection surface and the second projection surface to isolate the bottom boundary of the first projection surface from the top boundary of the second projection surface, wherein the guard band is connected to the bottom boundary of the first projection surface and also to the top boundary of the second projection surface. Therefore, the width of the guard band region (consisting of a single guard band) inserted between the first projection surface (packed at the surface position of position index i=2) and the second projection surface (packed at the surface position of position index i=3) is equal to D.
[0064] like Figure 7As shown, in the case where the syntax element gcmp_guard_band_flag is set to 1, when the syntax element gcmp_packing_type is set to 0 and the syntax element gcmp_guard_band_boundary_exterior_flag is set to 1, the first guard band is added between the bottom surface boundary of the projection surface (packed at the surface position of position index i=2) and the top surface boundary of the projection surface (packed at the surface position of position index i=3), and the second guard band is added between the bottom surface boundary of the projection surface (packed at the surface position of position index i=0). The top boundary of the projection surface, the third guard band is added at the bottom boundary of the projection surface packed at the surface position of position index i=5, the fourth guard band is added at the multiple left boundaries of the multiple projection surfaces packed at the multiple surface positions of position index i=0-5, and the fifth guard band is added at the multiple right boundaries of the multiple projection surfaces packed at the multiple surface positions of position index i=0-5, wherein the first guard band, the second guard band, the third guard band, the fourth guard band and the fifth guard band have the same guard band size D as specified by gcmp_guard_band_samples_minus1 plus 1. Specifically, the second guard band, the third guard band, the fourth guard band and the fifth guard band serve as the boundaries of the cube-based projection layout. In addition, the width of a guard band area (consisting of a single guard band) inserted between two projection surfaces (packed at the surface positions of position index i=2 and i=3) is equal to D.
[0065] Since those skilled in the art can easily understand after reading the above paragraphs Figure 8 and Figure 9 Details of the other guard band configurations shown are omitted for brevity.
[0066] As described above, the syntax element gcmp_guard_band_samples_minus1 is arranged to provide size information of each guard band packed in a projection-based frame (e.g., IMG or IMG'). In some embodiments of the present invention, the decoder-side conversion circuitry 125 may refer to the syntax element gcmp_guard_band_samples_minus1 parsed from the bitstream BS to Figure 7-9 Calculations are performed for a specific application under one of the guard band configurations shown in .
[0067] For example, the size information specified by the syntax element gcmp_guard_band_samples_minus1 may relate to the calculation of the size of each projection face packed in a cube-based projection layout with guard bands. The input to this process may include the width pictureWidth and height pictureHeight of the projection-based frame (e.g., decoded frame IMG'). The output of this process may include the width faceWidth and height faceHeight of the projection face packed in the projection-based frame (e.g., decoded frame IMG').
[0068] Outputs can be exported as follows:
[0069]
[0070]
[0071]
[0072] For another example, the size information specified by the syntax element gcmp_guard_band_samples_minusl may involve a conversion from sample positions within the decoded frame IMG' to local sample positions within one of the projection planes packed in a cube-based projection layout with a guard band. The input to the process may include the width faceWidth and height faceHeight of the projection plane, and may also include the sample positions (hPos, vPos) of the projection-based frame (e.g., the decoded frame IMG'). The output of the process may include the local sample positions (hPosFace, vPosFace) of the projection plane packed in the projection-based frame (e.g., the decoded frame IMG').
[0073] Outputs can be exported as follows:
[0074]
[0075]
[0076] Those skilled in the art will readily observe that various modifications and variations can be made to the apparatus and method while retaining the teachings of the present invention.Accordingly, the above disclosure should be construed as being limited only by the metes and bounds of the appended claims.
Claims
1. A video decoding method, comprising: decoding a portion of the bitstream to generate a decoded frame, wherein the decoded frame is a projection-based frame, the projection-based frame including a plurality of projection planes packed in a cube-based projection layout, and at least a portion of the 360-degree content of the sphere is mapped to the plurality of projection planes by the cube-based projection; as well as parsing at least one syntax element from a bitstream, wherein the at least one syntax element indicates a guard band configuration for a projection-based frame, Wherein parsing the at least one syntax element from the bitstream comprises: parsing one of the at least one syntax element from the bitstream; and In response to the one syntax element being set to indicate that the projection-based frame includes at least one guard band, parsing another syntax element of the at least one syntax element from the bitstream, wherein the one syntax element and the another syntax element are separate syntax elements sent via the bitstream, Therein, the further syntax element is arranged to indicate whether at least one guard band packed in the projection-based frame includes a guard band serving as a boundary of the cube-based projection layout.
2. The video decoding method according to claim 1, wherein: The packing of the plurality of projection surfaces is selected from the group consisting of packing of regular cubemap projection surfaces and packing of hemispherical cubemap projection surfaces.
3. The video decoding method according to claim 1, wherein: The multiple projection surfaces respectively correspond to multiple surfaces of an object in three-dimensional space, including a first projection surface and a second projection surface, wherein the object is a cube or a hemispherical cube; The at least one guard band packed in the cube-based projection layout includes a first guard band, wherein, with respect to the cube-based projection layout on a two-dimensional plane, a face boundary of the first projection surface is parallel to a face boundary of the second projection surface, and the first guard band is located between the first projection surface and the second projection surface.
4. The video decoding method according to claim 3, wherein: The at least one protection band also includes a second protection band, wherein with respect to the cube-based projection layout on the two-dimensional plane, the first protection band and the second protection band are both located between the first projection surface and the second projection surface, for isolating the one surface boundary of the first projection surface from the one surface boundary of the second projection surface, the first protection band is connected to the one surface boundary of the first projection surface and the second protection band, and the second protection band is connected to the first protection band and the one surface boundary of the second projection surface.
5. The video decoding method according to claim 3, wherein: Regarding the cube-based projection layout on the two-dimensional plane, the first protection band connects the one surface boundary of the first projection surface and the one surface boundary of the second projection surface, and is used to isolate the one surface boundary of the first projection and the one surface boundary of the second projection surface.
6. The video decoding method according to claim 1, wherein: The at least one syntax element further comprises: A further syntax element is configured to provide size information for each of the at least one guard band.
7. The video decoding method according to claim 6, wherein: The size information relates to calculating the size of each projection surface of the plurality of projection surfaces, or relates to a conversion from a sample position within the projection-based frame to a local sample position within one of the plurality of projection surfaces.
8. An electronic device comprising: a decoding circuit configured to decode a portion of the bitstream to generate a decoded frame and to parse at least one syntax element from the bitstream; wherein the decoded frame is a projection-based frame, the projection-based frame comprising a plurality of projection surfaces packed in a cube-based projection layout, and at least a portion of the 360-degree content of the sphere is mapped to the plurality of projection surfaces by the cube-based projection, wherein the at least one syntax element indicates a guard band configuration of the projection-based frame, Wherein, the decoding circuit is arranged as follows: parsing one of the at least one syntax element from the bitstream; and In response to the one syntax element being set to indicate that the projection-based frame includes at least one guard band, parsing another syntax element of the at least one syntax element from the bitstream, wherein the one syntax element and the another syntax element are separate syntax elements sent via the bitstream, Therein, the further syntax element is arranged to indicate whether at least one guard band packed in the projection-based frame includes a guard band serving as a boundary of the cube-based projection layout.
9. The electronic device according to claim 8, wherein The packing of the plurality of projection surfaces is selected from the group consisting of packing of regular cubemap projection surfaces and packing of hemispherical cubemap projection surfaces.
10. The electronic device according to claim 8, wherein The multiple projection surfaces respectively correspond to multiple surfaces of an object in three-dimensional space, including a first projection surface and a second projection surface, wherein the object is a cube or a hemispherical cube; the at least one protection band packaged in the cube-based projection layout includes a first protection band, wherein with respect to the cube-based projection layout on a two-dimensional plane, a surface boundary of the first projection surface is parallel to a surface boundary of the second projection surface, and the first protection band is located between the one surface boundary of the first projection surface and the one surface boundary of the second projection surface.
11. The electronic device according to claim 10, wherein: The at least one protection band also includes a second protection band, wherein with respect to the cube-based projection layout on the two-dimensional plane, the first protection band and the second protection band are both located between the first projection surface and the second projection surface, for isolating the one surface boundary of the first projection surface from the one surface boundary of the second projection surface, the first protection band is connected to the one surface boundary of the first projection surface and the second protection band, and the second protection band is connected to the first protection band and the one surface boundary of the second projection surface.
12. The electronic device according to claim 10, wherein: Regarding the cube-based projection layout on the two-dimensional plane, the first protection band connects the one surface boundary of the first projection surface and the one surface boundary of the second projection surface, and is used to isolate the one surface boundary of the first projection and the one surface boundary of the second projection surface.
13. The electronic device according to claim 8, wherein The at least one syntax element further comprises: A further syntax element is configured to provide size information for each of the at least one guard band.
14. The electronic device according to claim 13, wherein: The size information relates to calculating the size of each projection surface of the plurality of projection surfaces, or relates to a conversion from a sample position within the projection-based frame to a local sample position within one of the plurality of projection surfaces.
15. A video encoding method, comprising: encoding a projection-based frame to generate a portion of a bitstream, wherein at least a portion of 360-degree content of a sphere is mapped to a plurality of projection planes via a cube-based projection, and the projection-based frame has the plurality of projection planes packed in a cube-based projection layout; as well as sending at least one syntax element via a bitstream, wherein the at least one syntax element indicates a guard band configuration of the projection-based frame, The sending of the at least one syntax element through the bitstream includes: transmitting one of the at least one syntax element via the bitstream; and in response to the one syntax element being set to indicate that the projection-based frame includes at least one guard band, transmitting another syntax element of the at least one syntax element via the bitstream, wherein the one syntax element and the another syntax element are separate syntax elements transmitted via the bitstream, Therein, the further syntax element is arranged to indicate whether at least one guard band packed in the projection-based frame includes a guard band serving as a boundary of the cube-based projection layout.
Citation Information
Patent Citations
Hybrid cubemap projection for 360-degree video coding
WO2018218028A1