Displacement data codec for dynamic grid codec
By using a dynamic mesh coding and decoding method, the basic mesh, displacement vector and texture data are compressed using 2D video coding and decoding standards, and displacement data is transmitted using only the luma channel. This solves the problems of increased bandwidth requirements and complexity in dynamic mesh coding and decoding technology, and achieves high coding and decoding efficiency and throughput.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DOUYIN CO LTD
- Filing Date
- 2024-09-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing dynamic mesh coding and decoding technologies are difficult to compress and transmit 3D video data efficiently, leading to increased bandwidth requirements and processing complexity.
The dynamic mesh coding and decoding method is adopted, which uses 2D video coding and decoding standards to compress the basic mesh, displacement vector and texture data. The coding and decoding process of displacement data is optimized by wavelet transform and arithmetic coding and decoding, and displacement data is transmitted only using the luminance channel.
It improves the encoding and decoding efficiency and throughput of dynamic grid data, reduces processing complexity, and lowers bandwidth requirements.
Smart Images

Figure CN121970334A_ABST
Abstract
Description
Displacement data encoding and decoding for dynamic mesh encoding and decoding
[0001] Cross-reference to related applications
[0002] This patent application claims the benefit of U.S. Patent Application No. 63 / 588,615, filed October 6, 2023, which is incorporated herein by reference. Technical Field
[0003] This disclosure relates to the generation, storage, and use of digital audio and video media information in file formats. Background Technology
[0004] Digital video accounts for the largest share of bandwidth used on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is likely to continue to grow. Summary of the Invention
[0005] The first aspect relates to a method for processing video data, comprising: determining, when displacement data is encoded or decoded in a 4:4:4 video format, to use only the luminance channel to transmit the displacement data; and performing a conversion between visual media data and a bitstream based on the displacement data.
[0006] Alternatively, in any of the above aspects, another embodiment of the aspect provides displacement data including three-dimensional (3D) displacement data encoded and decoded into video, wherein all 3D displacement data is in the luminance channel.
[0007] Alternatively, in any of the above aspects, another implementation of that aspect provides that displacement data is packaged into the video, regardless of the color format.
[0008] Alternatively, in any of the above aspects, another implementation of that aspect provides that, when displacement data is packaged into a video, only the luminance channel is used to transmit the displacement data, regardless of the color format.
[0009] Alternatively, in any of the above aspects, another implementation of that aspect provides that, regardless of the color format, when the displacement dimension (DisplacementDim) is equal to 1, the first displacement component is derived from the first color component of the video, and the second and third color components are assumed to be 0.
[0010] Alternatively, in any of the above aspects, another implementation of that aspect provides that, regardless of the color format, when the displacement dimension (DisplacementDim) is equal to 3, the first displacement component, the second displacement component, and the third displacement component are derived from the first color component of the video.
[0011] Alternatively, in any of the above aspects, another embodiment of the aspect provides that, when the base grid corresponding to the second time (t2) is in the reference list of the base grid corresponding to the first time (t1), the displacement data at the first time is only allowed to use the displacement data at the second time as a reference.
[0012] Alternatively, in any of the above aspects, another embodiment of that aspect provides that, when the base grid at the second time point is a reference to the base grid at the first time point, the displacement data at the first time point is only allowed to use the displacement data at the second time point as a reference.
[0013] Alternatively, in any of the above aspects, another embodiment of the aspect provides that, when the displacement at the second time (t2) is in a reference list of the displacement at the first time (t1), the base grid at the first time is only allowed to use the base grid at the second time.
[0014] Alternatively, in any of the above aspects, another embodiment of that aspect provides that, when the base grid at the second time point is a reference to the base grid at the first time point, the displacement data at the first time point is only allowed to use the displacement data at the second time point as a reference.
[0015] Optionally, in any of the above aspects, another embodiment of the aspect provides that at least one of the displacement reference list and the base mesh reference list is set to correspond to the atlas reference list.
[0016] Optionally, in any of the above aspects, another embodiment of the aspect provides that, when the atlas corresponding to the second time (t2) is in the reference list of the atlas at the first time (t1), the base grid and displacement data at the first time are only allowed to use the base grid and displacement data at the second time as references.
[0017] Alternatively, in any of the above aspects, another embodiment of the aspect provides that, when the atlas at the second time (t2) is a reference to the atlas at the first time (t1), the base grid and displacement data at the first time are only allowed to use the base grid and displacement data at the second time.
[0018] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a synchronization method for all sub-bitstreams corresponding to a video codec standard, wherein the video codec standard includes one of video-based point cloud compression (V-PCC) and video-based dynamic mesh codec (V-DMC).
[0019] Alternatively, in any of the above aspects, another implementation of that aspect provides that all sub-bitstreams have the same reference list.
[0020] Alternatively, in any of the above aspects, another implementation of that aspect provides that all sub-bitstreams have the same reference structure.
[0021] Optionally, in any of the foregoing aspects, another implementation of that aspect provides one or more syntax elements used to indicate the maximum allowed number of decoding buffers for the decoding process and each sub-bitstream, wherein the number of decoding buffers is not greater than the maximum allowed number of decoding buffers for the decoding process.
[0022] Optionally, in any of the foregoing aspects, another implementation of that aspect provides one or more syntax elements used to indicate the maximum allowed number of reordered frames for the decoding process and each sub-bitstream, wherein the number of reordered frames is not greater than the maximum allowed number of reordered frames for the decoding process.
[0023] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides one or more syntax elements used to indicate the maximum allowed number of atlas frames with an atlas frame output flag equal to 1 that are allowed to precede any atlas frame with an atlas frame output flag equal to 1 in the output order at a particular time-domain layer.
[0024] Optionally, in any of the foregoing aspects, another implementation of that aspect provides one or more syntax elements used to indicate the maximum allowed number of delayed frames for the decoding process and each sub-bitstream, wherein the number of delayed frames is not greater than the maximum allowed number of delayed frames for the decoding process.
[0025] Optionally, in any of the foregoing aspects, another implementation of the aspect provides one or more syntax elements used to indicate the maximum allowed number of atlas frames with AtlasFrameOutputFlag equal to 1 that are allowed in a particular time-domain layer, in output order, before any atlas frame with AtlasFrameOutputFlag equal to 1 and after any atlas frame with AtlasFrameOutputFlag equal to 1.
[0026] Alternatively, in any of the above aspects, another implementation of that aspect provides that all sub-bit streams at a specific time have the same time-domain identifier.
[0027] Alternatively, in any of the foregoing aspects, another implementation of that aspect provides a conversion that includes encoding media data into a bitstream.
[0028] Alternatively, in any of the above aspects, another implementation of that aspect provides a conversion that includes decoding media data from a bitstream.
[0029] The second aspect relates to an apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform any of the disclosed methods.
[0030] The third aspect relates to a non-transitory computer-readable medium including a computer program product for use by a video codec apparatus, the computer program product including computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform any of the disclosed methods when executed by a processor.
[0031] The fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining, when displacement data is encoded or decoded in a 4:4:4 video format, to use only the luminance channel to transmit the displacement data; and generating a bitstream based on the displacement data.
[0032] The fifth aspect relates to a method for storing a bitstream of video, comprising: determining, when displacement data is encoded or decoded in a 4:4:4 video format, to use only the luminance channel to transmit the displacement data; generating a bitstream based on the displacement data; and storing the bitstream in a non-transitory computer-readable recording medium.
[0033] The sixth aspect relates to a method, apparatus, or system described in this disclosure.
[0034] For clarity, any of the embodiments described above may be combined with any one or more of the other embodiments described above to create new embodiments within the scope of this disclosure.
[0035] These and other features will become clearer from the following detailed description by referring to the accompanying drawings and claims. Attached Figure Description
[0036] To gain a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals denote like parts.
[0037] Figure 1 is a block diagram illustrating the decoder design of the dynamic mesh codec.
[0038] Figure 2 is a block diagram showing the structure of the dynamic mesh encoding / decoding test model.
[0039] Figure 3 is a block diagram illustrating an example video processing system.
[0040] Figure 4 is a block diagram of an example video processing device.
[0041] Figure 5 is a flowchart of an example method for video processing.
[0042] Figure 6 is a block diagram illustrating an example video codec system.
[0043] Figure 7 is a block diagram showing an example encoder.
[0044] Figure 8 is a block diagram showing an example decoder.
[0045] Figure 9 is a schematic diagram of an example encoder. Detailed Implementation
[0046] First, it should be understood that although illustrative implementations of one or more embodiments are provided below, the disclosed systems and / or methods can be implemented using any number of techniques, whether currently known or yet to be developed. This disclosure should not be limited in any way to the illustrative implementations, drawings, and techniques described below, including the exemplary designs and implementations described herein, but can be modified within the scope of the appended claims and their full equivalents.
[0047] 1. Preliminary Discussion
[0048] This disclosure relates to improvements in dynamic mesh coding and decoding based on Motion Picture Experts Group Immersive (MPEG-I) video. It is also applicable to other immersive video coding and decoding standards or codecs.
[0049] 2. Further discussion
[0050] In computer graphics, three-dimensional (3D) / immersive content can typically be represented by 3D meshes and texture maps. These mesh and texture data can be generated by machines or converted from images captured by multiple cameras from different angles. Similar to two-dimensional (2D) video, the mesh and texture data change as these 3D contents change over time and form a dynamic mesh sequence. Dynamic meshes are typically large in data and difficult to store and transmit. To meet the needs of applications using dynamic meshes, the Motion Picture Experts Group (MPEG) issued a call for proposals [1]. One of the key requirements for effectively using existing 2D codecs is to use current 2D video codec standards to compress most of the data and keep the rest simple and low in complexity. Such a requirement ensures that the representation can take advantage of 2D video hardware / software systems without spending a lot of effort redesigning specific systems dedicated to dynamic meshes.
[0051] MPEG received five responses to the solicitation of proposals. Among them, proposal [2] showed better performance compared to the other proposals. Therefore, based on [2], a test model was built for the development of the planned dynamic mesh coding standard.
[0052] Until the draft document was prepared, the latest test model for dynamic mesh coding and decoding could be found at this link: http: / / mpegx.int-evry.fr / software / MPEG / dmc / mpeg-vmesh-tm / - / tags / v4.0; the latest working draft document is WD3.0[3].
[0053] 2.1 Data Representation in Dynamic Mesh Encoding and Decoding
[0054] Figure 1 is a block diagram showing the decoder design of the dynamic mesh codec. Figure 1 shows the decoder design as described in WD1.0[3]. It can be seen that the dynamic mesh decoder receives three bit streams and performs decoding to reconstruct the dynamic mesh plus texture signal. The first bit stream represents the base mesh, which is a decimated version of the original mesh. The second bit stream represents the displacement vector between the reconstructed base mesh and the original mesh. The displacement vector is arranged as 2D video and compressed with a codec that conforms to the 2D video codec standard. The third bit stream represents the texture (or attribute map). The attribute map is also arranged as 2D video and compressed with a codec that conforms to the 2D video codec standard. The design concept is to make the base mesh part small enough so that the module for processing the base mesh can be easily implemented. On the other hand, the displacement vector and attribute map account for most of the volume of the entire dynamic mesh data, which can be processed by the current dedicated high-efficiency 2D video codec system. Such a design can reduce the extra effort required to implement the dynamic mesh codec system and ensure high throughput and encoding / decoding efficiency of the dynamic mesh data.
[0055] 2.2 Test Model for Dynamic Mesh Encoding and Decoding
[0056] Figure 2 is a block diagram illustrating the structure of a dynamic mesh codec test model. Figure 2 shows the structure of an example dynamic mesh codec model. In this model, Draco is used to compress the base mesh, and a High-Efficiency Video Coding (HEVC) test model (e.g., HM) is used to compress the displacement vector and property map. However, it should be noted that other mesh or video codec systems can also be used in dynamic mesh codecs.
[0057] The base mesh m is generated from the original mesh using a downsampling scheme. Its quantized version m' is then encoded and decoded using Draco. The reconstructed base mesh m'' can be obtained by inverse quantizing m'. The displacement vector d' is generated by calculating the difference between the subdivision version of the original mesh and m'' using a subdivision scheme.
[0058] 2.3 Encoding and Decoding of Displacement Vectors
[0059] After obtaining the displacement vector d' (the difference between the original mesh and the subdivided base mesh), a lifting-based wavelet transform is applied to further compact the energy. The wavelet transform coefficients are then traversed from low to high frequencies using Morton order to form 2D coefficient blocks. These various 2D coefficient blocks comprise an image to be processed by a 2D codec.
[0060] 2.4 Sports Field Encoding and Decoding
[0061] In the test model, arithmetic coding and decoding are used to directly encode and decode the motion field between the base grids. Proposal [4] also investigated the use of a standard-compliant 2D coding and decoding system to encode and decode the motion field and showed that the loss of coding and decoding efficiency is negligible. Therefore, it makes sense to further transfer the motion field coding and decoding process to a 2D video codec.
[0062] 2.5 chroma format
[0063] Different chroma formats are supported in H.264 / Advanced Video Coding (AVC), H.265 / HEVC, and H.266 / Versatile Video Coding (VVC). The format can be transmitted via the signal using the syntax element `sps_chroma_format_idc` and is represented by the variable `ChromaFormatIdc`. The following table shows the chroma formats corresponding to the different `sps_chroma_format_idc` values:
[0064]
[0065] 2.6 Lossless Encoding and Decoding
[0066] In H.264 / AVC and H.266 / VVC, lossless encoding and decoding can be achieved by setting the quantization parameters (QP) to 4 and applying a reversible spatial transformation or transform skip to the block. In H.265 / HEVC, in addition to the above methods, lossless encoding and decoding can also be achieved by setting the cu_transquant_bypass_flag of the encoding / decoding unit to 1.
[0067] 2.7 Other Designs
[0068] In our earlier sample designs, we proposed the idea of combining multiple attributes (including texture, displacement data, occupancy data) into a single video for encoding / decoding without requiring multiple encoding / decoding capabilities on the device; color space for lossless texture encoding / decoding and sub-block size signals for displacement encoding / decoding based on arithmetic encoding / decoding.
[0069] 2.8 Depackaging of displacement data
[0070] In the example V-DMC design, when the displacement data is encoded as 4:0:0 video, the first displacement component is derived from the first color component, and the second and third displacement components are assumed to be 0; when the displacement data is encoded as 4:4:4 video, the first, second, and third displacement components are derived from the first, second, and third color components of the video, respectively; when the displacement data is encoded as 4:2:0 video and asps_vdmc_ext_1d_displacement_flag equals 1, the first displacement component is derived from the first color component of the video, and the second and third displacement components are assumed to be 0; when the displacement data is encoded as 4:2:0 video and asps_vdmc_ext_1d_displacement_flag equals 0, the first, second, and third displacement components are all derived from the first color component of the video. The corresponding description in Working Draft (WD) 3.0 is as follows:
[0071] 8.4.6.1.3 Atlas Sequence Parameter Set vdmc Extended RBSP Syntax
[0072] The `asps_vdmc_ext_subdivision_method` directive indicates the identifier of the method used to subdivide the mesh associated with the current atlas sequence parameter set. Table 2 describes a list of supported subdivision methods and their relationship to `asps_vdmc_ext_subdivision_method`.
[0073] Table 2
[0074]
[0075] `asps_vdmc_ext_subdivision_iteration_count` indicates the number of iterations used for subdivision. When it does not exist, the value of `asps_vdmc_ext_subdivision_iteration_count` is presumed to be 0.
[0076] The `asps_vdmc_ext_displacement_coordinate_system` is an identifier indicating the coordinate system of the grid associated with the current atlas sequence parameter set. Table 3 describes a list of supported displacement coordinate systems and their relationship to `asps_vdmc_ext_displacement_coordinate_system`.
[0077] Table 3
[0078]
[0079] The `asps_vdmc_ext_transform_method` directive indicates the identifier of the transformation applied to the displacement. Table 4 describes a list of supported transformations and their relationship to `asps_vdmc_ext_transform_method`.
[0080] Table 4
[0081]
[0082] asps_vdmc_ext_num_attribute_video indicates the number of attributes transmitted via the signal through the video sub-bitstream.
[0083] asps_vdmc_ext_attribute_type_id[i] indicates the attribute type of the Attribute Video Data cell at index i.
[0084] `asps_vdmc_ext_attribute_frame_width[i]` indicates the atlas frame width of the Attribute VideoData unit at index `i`, in integer luminance samples of the atlas with atlas ID `j`. The V3C bitstream consistency requirement is that the value of `asps_vdmc_ext_attribute_frame_width[i]` must be equal to the value of `vps_ext_attribute_frame_width[j][i]`, where `j` is the ID of the current atlas.
[0085] `asps_vdmc_ext_attribute_frame_height[i]` indicates the atlas frame height of the Attribute VideoData unit at index `i`, in integer luminance samples of the atlas with atlas ID `j`. The V3C bitstream consistency requirement is that the value of `asps_vdmc_ext_attribute_frame_height[i]` must be equal to the value of `vps_ext_attribute_frame_height[j][i]`, where `j` is the ID of the current atlas.
[0086] asps_vdmc_ext_attribute_transform_method[i] is the identifier of the transformation of the attribute transmitted via signal in the AttributeVideo Data cell at index i. Table 5 describes the list of supported transformations and their relationship to asps_vdmc_ext_attribute_transform_method.
[0087] Table 5
[0088]
[0089] If `asps_vdmc_ext_direct_attribute_projection_enabled_flag[i]` equals 0, it specifies that the signal-transmitted attribute of the `AttributeVideoData` cell at index `i` in the small data cell or the original small data cell is not transmitted via signal. If `asps_vdmc_ext_direct_attribute_projection_enabled_flag[i]` equals 1, it specifies that the signal-transmitted attribute of the `AttributeVideoData` cell at index `i` in the small data cell or the original small data cell is transmitted via signal.
[0090] When asps_vdmc_ext_packing_method is equal to 0, the displacement component samples are packed in ascending order; when asps_vdmc_ext_packing_method is equal to 1, the displacement component samples are packed in descending order.
[0091] `asps_vdmc_ext_1d_displacement_flag` equal to 1 indicates that only the normal (or x) component of the displacement exists in the compressed geometric video. The remaining two components are presumed to be 0. `asps_vdmc_ext_1d_displacement_flag` equal to 0 indicates that all three components of the displacement exist in the compressed geometric video.
[0092] A value of 0 for `asps_vdmc_ext_projection_textcoord_enable_flag` indicates that texture coordinates can be transmitted within the underlying mesh, while a value of 1 for `asps_vdmc_ext_projection_textcoord_enable_flag` indicates that texture coordinates will be derived using projection parameters from the `meshpatch` data cell.
[0093] The `asps_vdmc_ext_projection_textcoord_mapping_method` directive is the identifier for the variable `FaceToSubPatchMapping`, which indicates the method for mapping a set of faces to sub-patches. Table 6 describes a list of supported face-to-sub-patch mapping methods and their relationship to the variable `FaceToSubPatchMapping`.
[0094] Table 6
[0095]
[0096] asps_vdmc_ext_projection_textcoord_scale_factor indicates the value of the scaling factor variable TextCoordProjectionScaleFactor, which is used to derive texture coordinates from geometric projection.
[0097] 11.5 Inverse Image Packing of Wavelet Coefficients
[0098] The input to this process is:
[0099] `width` is a variable that indicates the width of the video frame being moved.
[0100] `height` is a variable that indicates the height of the displaced video frame.
[0101] bitDepth is a variable that indicates the bit depth of the displaced video frame.
[0102] dispQuantCoeffFrame, a 3D array of size width×height×3, indicates the packed quantized shift wavelet coefficients.
[0103] blockSize is a variable indicating the size of the displacement coefficient block.
[0104] verCoordCount is a variable that indicates the number of vertex coordinates in the subdivided submesh.
[0105] The output of this process is a dispQuantCoeffArray, a 2D array of size verCoordCount×3, which indicates the quantized displacement wavelet coefficients.
[0106] The requirement for bitstream consistency is that when DecGeoChromaFormat equals 4:0:0, asps_vdmc_ext_1d_displacement_flag must be equal to 1. Another requirement for bitstream consistency is that when DecGeoChromaFormat equals 4:4:4, asps_vdmc_ext_1d_displacement_flag must be equal to 0.
[0107] The 2D array `dispQuantCoeffArray` is initialized to 0. The variable `DisplacementDim` is set as follows:
[0108] If asps_vdmc_ext_1d_displacement_flag equals 1, then DisplacementDim is set to 1.
[0109] Otherwise, asps_vdmc_ext_1d_displacement_flag equals 0, and DisplacementDim is set to 3.
[0110] Let the function extracOddBits(x) be defined as follows:
[0111] x = extracOddBits( x ) {
[0112] x = x & 0x55555555
[0113] x = ( x | ( x >> 1 ) ) & 0x33333333
[0114] x = ( x | ( x >> 2 ) ) & 0x0F0F0F0F
[0115] x = ( x | ( x >> 4 ) ) & 0x00FF00FF
[0116] x = ( x | ( x >> 8 ) ) & 0x0000FFFF
[0117] }
[0118] Let the function computeMorton2D(i) be defined as follows:
[0119] ( x, y) = computeMorton2D( i ) {
[0120] x = extracOddBits( i >> 1 )
[0121] y = extracOddBits( i )
[0122] }
[0123] The inverse unpacking process of wavelet coefficients is carried out as follows:
[0124] pixelsPerBlock = blockSize blockSize
[0125] widthInBlocks = width / blockSize
[0126] shift = (1 << bitDepth) >> 1
[0127] blockCount = (verCoordCount + pixelsPerBlock – 1) / pixelsPerBlock
[0128] heightInBlocks = (blockCount + widthInBlocks – 1) / widthInBlocks
[0129] origHeight = heightInBlocks blockSize
[0130] paddedHeight = height - 3 origHeight
[0131] if ( !asps_vdmc_ext_1d_displacement_flag )
[0132] start = (paddedHeight + origHeight) width – 1
[0133] else
[0134] start = (width height) - 1
[0135] for( v = 0; v < verCoordCount; v++ ) {
[0136] v0 = asps_vdmc_ext_packing_method ? start – v : v
[0137] blockIndex = v0 / pixelsPerBlock
[0138] indexWithinBlock = v0 % pixelsPerBlock
[0139] x0 = ( blockIndex % widthInBlocks ) blockSize
[0140] y0 = ( blockIndex / widthInBlocks ) blockSize
[0141] ( x, y ) = computeMorton2D( indexWithinBlock )
[0142] x1 = x0 + x
[0143] y1 = y0 + y
[0144] for( d = 0; d < DisplacementDim; d++ ) {
[0145] if ( DecGeoChromaFormat == 4:2:0 ) {
[0146] dispQuantCoeffArray[ v0 ][ d ] =
[0147] dispQuantCoeffFrame[ x1 ][ d origHeight + y1 ][ d0 ] – shift
[0148] } else {
[0149] dispQuantCoeffArray[ v0 ][ d ] = dispQuantCoeffFrame[x1 ][ y1 ][ d ] – shift
[0150] }
[0151] }
[0152] }
[0153] 2.9 Basic Mesh Submesh Encoding and Decoding Syntax and Semantics
[0154] H.8.1.3.3 Basic Mesh Sub-mesh Layer RBSP Syntax
[0155]
[0156] H.8.1.3.4 Basic Mesh Sub-Mesh Header Syntax
[0157]
[0158]
[0159] H.8.1.3.5 Reference List Structure Syntax
[0160]
[0161] H.8.3.3 Basic Mesh Submesh Layer RBSP Semantics
[0162] H.8.3.4 Basic Mesh Sub-mesh Head Semantics
[0163] When present, the values of the atlas header syntax elements smh_basemesh_frame_parameter_set_id, smh_mesh_output_flag, smh_no_output_of_prior_mesh_frames_flag, and smh_mesh_frm_order_cnt_lsb must be the same in all sub-mesh headers of the encoded / decoded mesh frame.
[0164] The `smh_no_output_of_prior_mesh_frames_flag` affects the output of previously decoded mesh frames in the DAB after decoding the atlas in a CAS AU that is not the first AU in the bitstream. When `smh_no_output_of_prior_mesh_frames_flag` is not present, its value is presumed to be 0.
[0165] The requirement for bitstream consistency is that the value of smh_no_output_of_prior_mesh_frames_flag must be the same for all mesh frames in the AU.
[0166] The value of smh_no_output_of_prior_mesh_frames_flag in the submesh header is also called the output_of_prior_mesh_frames_flag value of AU.
[0167] smh_basemesh_frame_parameter_set_id specifies the value of bfps_basemesh_frame_parameter_set_id for the active base mesh frame parameter set of the current submesh.
[0168] smh_id specifies the subgrid ID associated with the current subgrid. If it does not exist, the value of smh_id is presumed to be 0.
[0169] The following applies:
[0170] The length of –smh_id is bmsi_signalled_submesh_id_length_minus1+1 bits.
[0171] The value of –smh_id must be within the range of values specified by the array SubMeshIndexToID[i], where i is in the range from 0 to bsmi_num_submeshes_minus1 (inclusive).
[0172] The requirement for bitstream consistency applies to the following constraints:
[0173] The value of –smh_id should not be equal to the value of smh_id of any other encoded atlas slice unit of the same encoded atlas frame.
[0174] – The atlas frames must be sorted in ascending order of their smh_id values.
[0175] The `smh_type` parameter specifies the encoding / decoding type for the current subgrid according to Table 8. In bitstreams conforming to this version of the document, the value of `smh_type` must be equal to 0, 1, or 2. Other values for `smh_type` are reserved for future use by ISO / IEC. Decoders conforming to this version of the document must ignore the reserved values for `smh_type`.
[0176] Table 8 – Association between smh_type and its name
[0177]
[0178] The `smh_mesh_output_flag` affects the mesh output and removal process during decoding. When `smh_mesh_output_flag` is not present, it is presumed to be equal to 1.
[0179] `smh_mesh_frm_order_cnt_lsb` specifies the mesh frame order count of the current submesh, modulo `MaxMeshFrmOrderCntLsb`. The length of the `smh_mesh_frm_order_cnt_lsb` syntax element is equal to `Log2MaxMeshFrmOrderCntLsb` bits. The value of `smh_mesh_frm_order_cnt_lsb` must be in the range of 0 to `MaxMeshFrmOrderCntLsb-1` (inclusive).
[0180] `smh_ref_mesh_frame_list_bmsps_flag` equal to 1 indicates that the reference bmesh frame list for the current submesh is deduced based on one of the `bmesh_ref_list_struct(rlsIdx)` syntax structures in the active BMSPS. `smh_ref_mesh_frame_list_bmsps_flag` equal to 0 indicates that the reference bmesh frame list for the current submesh is deduced based on the `bmesh_ref_list_struct(rlsIdx)` syntax structure directly included in the submesh header of the current submesh. When `bmsps_num_ref_mesh_frame_lists_in_bmsps` equals 0, the value of `smh_ref_mesh_frame_list_bmsps_flag` is presumed to be 0.
[0181] `smh_ref_mesh_frame_list_idx` specifies the index of the `bmesh_ref_list_struct(rlsIdx)` syntax structure used to deduce the list of reference mesh frames for the current submesh, within the list of `bmesh_ref_list_struct(rlsIdx)` syntax structures included in the active ASPS. The syntax element `smh_ref_mesh_frame_list_idx` is represented by Ceil(Log2(bmsps_num_ref_mesh_frame_lists_in_bmsps)) bits. When it does not exist, the value of `smh_ref_mesh_frame_list_idx` is presumed to be 0. The value of `smh_ref_mesh_frame_list_idx` must be in the range of 0 to `bmsps_num_ref_mesh_frame_lists_in_bmsps-1` (inclusive). When smh_ref_mesh_frame_list_bmsps_flag equals 1 and bmsps_num_ref_mesh_frame_lists_in_bmsps equals 1, the value of smh_ref_mesh_frame_list_idx is presumed to be 0.
[0182] The variable RlsIdx for the current atlas piece is derived as follows:
[0183] RlsIdx=smh_ref_mesh_frame_list_bmsps_flag?
[0184] smh_ref_mesh_frame_list_idx:bmsps_num_ref_mesh_frame_lists_in_bmsps
[0185] A value of 1 for smh_additional_mfoc_lsb_present_flag[j] indicates that smh_additional_mfoc_lsb_val[j] exists for the current subgrid. A value of 0 for smh_additional_mfoc_lsb_present_flag[j] indicates that smh_additional_mfoc_lsb_val[j] does not exist.
[0186] smh_additional_mfoc_lsb_val[j] specifies the value of FullMeshFrmOrderCntLsbLt[RlsIdx][j] for the current atlas piece as follows:
[0187] FullMeshFrmOrderCntLsbLt[RlsIdx][j]= smh_additional_mfoc_lsb_val[j] MaxMeshFrmOrderCntLsb+mfoc_lsb_lt[RlsIdx][j]
[0188] The syntax element smh_additional_mfoc_lsb_val[j] is represented by smh_additional_lt_mfoc_lsb_len bits. When it does not exist, the value of smh_additional_mfoc_lsb_val[j] is presumed to be 0.
[0189] A value of 1 for `smh_num_ref_idx_active_override_flag` indicates that the syntax element `smh_num_ref_idx_active_minus1` exists for the current submesh. A value of 0 for `smh_num_ref_idx_active_override_flag` indicates that the syntax element `smh_num_ref_idx_active_minus1` does not exist. If `smh_num_ref_idx_active_override_flag` does not exist, its value is presumed to be 0.
[0190] `smh_num_ref_idx_active_minus1` is used to derive the variable `NumRefIdxActive`, as specified in Equation 5 for the current submesh. The value of `smh_num_ref_idx_active_minus1` must be in the range of 0 to 14 (inclusive).
[0191] When the current submesh is a P_SUBMESH submesh, smh_num_ref_idx_active_override_flag is equal to 1, and smh_num_ref_idx_active_minus1 does not exist, smh_num_ref_idx_active_minus1 is presumed to be equal to 0.
[0192] The variable NumRefIdxActive is derived as follows:
[0193] if( smh_type == P_SUBMESH || smh_type == SKIP_SUBMESH ) {
[0194] if( smh_num_ref_idx_active_override_flag == 1 )
[0195] NumRefIdxActive = smh_num_ref_idx_active_minus1 + 1 (5)
[0196] else {
[0197] if( num_ref_entries[ RlsIdx ] >= bfps_num_ref_idx_default_active_minus1 + 1 )
[0198] NumRefIdxActive = bfps_num_ref_idx_default_active_minus1 + 1
[0199] else
[0200] NumRefIdxActive = num_ref_entries[RlsIdx]
[0201] }
[0202] }
[0203] else
[0204] NumRefIdxActive = 0
[0205] NumRefIdxActive minus 1 specifies the maximum value of the atlas reference frame index that can be used to decode the current atlas piece.
[0206] H.8.3.5 Reference List Structure Semantics
[0207] `num_ref_entries[rlsIdx]` specifies the number of entries in the `bmesh_ref_list_struct(rlsIdx)` syntax structure, where `rlsIdx` is the index of the mesh frame reference list. For `P_SUBMESH` and `SKIP_SUBMESH`, the value of `num_ref_entries[rlsIdx]` must be in the range of 1 to `bmsps_max_dec_mesh_frame_buffering_minus1+1`. Otherwise, the value of `num_ref_entries[rlsIdx]` must be in the range of 0 to `bmsps_max_dec_mesh_frame_buffering_minus1+1`.
[0208] A value of 1 for `st_ref_mesh_frame_flag[rlsIdx][i]` indicates that the i-th entry in the `bmesh_ref_list_struct(rlsIdx)` syntax structure is a short-term reference mesh frame entry. A value of 0 for `st_ref_mesh_frame_flag[rlsIdx][i]` indicates that the i-th entry in the `ref_list_struct(rlsIdx)` syntax structure is a long-term reference mesh frame entry. When `st_ref_mesh_frame_flag[rlsIdx][i]` does not exist, its value is presumed to be 1.
[0209] The variable NumLtrMeshFrmEntries[rlsIdx] is derived as follows:
[0210] NumLtrMeshFrmEntries[rlsIdx] = 0
[0211] for( i = 0; i < num_ref_entries[ rlsIdx ]; i++ )
[0212] if( !st_ref_mesh_frame_flag[ rlsIdx ][ i ] ) (6)
[0213] NumLtrMeshFrmEntries[ rlsIdx ]++
[0214] abs_delta_mfoc_st[rlsIdx][i] specifies the absolute difference between the grid frame order count of the current mesh piece and the grid frame order count of the mesh frame referenced by the i-th entry when the i-th entry is the first short-lived reference mesh frame entry in the bmesh_ref_list_struct(rlsIdx) syntax structure; or, when the i-th entry is a short-lived reference mesh frame entry but not the first short-lived reference mesh frame entry in the bmesh_ref_list_struct(rlsIdx) syntax structure, specifies the absolute difference between the grid frame order count of the mesh frame referenced by the i-th entry and the grid frame order count of the mesh frame referenced by the previous short-lived reference mesh frame entry in the bmesh_ref_list_struct(rlsIdx) syntax structure.
[0215] The value of abs_delta_mfoc_st[rlsIdx][i] must be in the range of 0 to 2^15-1 (inclusive).
[0216] A value of 1 for `straf_entry_sign_flag[rlsIdx][i]` indicates that the i-th entry in the syntax structure `bmesh_ref_list_struct(rlsIdx)` has a value greater than or equal to 0. A value of 0 for `straf_entry_sign_flag[rlsIdx][i]` indicates that the i-th entry in the syntax structure `bmesh_ref_list_struct(rlsIdx)` has a value less than 0. When `straf_entry_sign_flag[rlsIdx][i]` does not exist, the value of `straf_entry_sign_flag[rlsIdx][i]` is presumed to be 1.
[0217] The list DeltaMfocSt[rlsIdx][i] is derived as follows:
[0218] for( i = 0; i < num_ref_entries[ rlsIdx ]; i++ )
[0219] if( st_ref_mesh_frame_flag[ rlsIdx ][ i ] )
[0220] DeltaMfocSt[ rlsIdx ][ i ] =
[0221] (2) straf_entry_sign_flag[ rlsIdx ][ i ] – 1 ) abs_delta_ mfoc_st[ rlsIdx ][ i ] (7)
[0222] else
[0223] DeltaMfocSt[ rlsIdx ][ i ] = 0
[0224] `mfoc_lsb_lt[rlsIdx][i]` specifies the grid frame sequence number modulo `MaxMeshFrmOrderCntLsb` of the grid frame referenced by the `i`th entry in the `bmesh_ref_list_struct(rlsIdx)` syntax structure. The length of the `mfoc_lsb_lt[rlsIdx][i]` syntax element is `Log2MaxMeshFrmOrderCntLsb` bits.
[0225] 2.10 Arithmetic Encoding and Decoding of Displacement Data
[0226] In the example working draft of dynamic mesh coding [3], displacement data can be encoded and decoded using arithmetic coding and decoding. We refer to this method as Alternating Current (AC) based displacement coding and decoding. The following syntax table and semantics illustrate the design:
[0227] J.7 Syntax and Semantics
[0228] J.7.1 Grammar in Tabular Form
[0229] J.7.1.1 General NAL Unit Syntax
[0230]
[0231] J.7.1.3 Raw Byte Sequence Payload, Trailing Bit, and Byte Alignment Syntax
[0232] J.7.1.3.1 Displacement Sequence Parameter Set RBSP Syntax
[0233] J.7.1.3.1.1 General Displacement Sequence Parameter Set (RBSP) Syntax
[0234]
[0235] J.7.1.3.1.2 Displacement levels, layers, and grades syntax
[0236]
[0237] J.7.1.3.1.3 Level Toolset Constraint Information Syntax
[0238]
[0239] J.7.1.3.2 Displacement Frame Parameter Set (RBSP) Syntax
[0240] J.7.1.3.2.1 General Displacement Frame Parameter Set (RBPS) Syntax
[0241]
[0242] J.7.1.3.3 Displacement Reference List Structure Syntax
[0243]
[0244] J.7.1.3.4 Displacement Layer RBSP Syntax
[0245]
[0246] J.7.1.3.5 Displacement Head Syntax
[0247]
[0248]
[0249] J.7.1.3.6 Displacement Data Unit Syntax
[0250]
[0251] J.7.1.3.7 Syntax for Displacement Intra-Frame Data Units
[0252]
[0253]
[0254] J.7.1.3.8 Syntax for Displacement Inter-Frame Data Units
[0255]
[0256] J.7.2 Semantics
[0257] J.7.2.1 NAL Unit Semantics
[0258] J.7.2.1.1 General NAL Unit Semantics
[0259] NumBytesInNalUnit specifies the size of the NAL unit in bytes. This value is necessary for decoding the NAL unit. Some form of NAL unit boundary delimitation is necessary to make inference using NumBytesInNalUnit possible. One such delimitation method is specified for the sample stream format. Other delimitation methods may be specified outside of this document.
[0260] Note 1 – The Displacement Coding Layer (DCL) is specified to efficiently represent the content of displacement data. The NAL is specified to format the data and provide header information in a manner suitable for transmission over various communication channels or storage media. All data is contained within NAL units, each containing an integer number of bytes. NAL units specify a general format for both packet-oriented and bitstream systems. The format of NAL units for packet-oriented transmission and sample streams is the same, except that in the sample stream format specified in Appendix TBD, each NAL unit may be preceded by an additional element specifying the size of the NAL unit.
[0261] rbsp_byte[i] is the i-th byte of the RBSP. The RBSP is specified as an ordered sequence of bytes as follows:
[0262] RBSP contains the following string of data bits (SODB):
[0263] - If SODB is empty (i.e., its length is zero bits), then RBSP is also empty.
[0264] - Otherwise, the RBSP contains the following SODB:
[0265] 1) The first byte of the RBSP contains the first (most prominent, leftmost) eight bits of the SODB; the next byte of the RBSP contains the next eight bits of the SODB, and so on, until the remaining SODB is less than eight bits.
[0266] 2) The rbsp_trailing_bits() syntax structure exists after SODB, as shown below:
[0267] i) The first (most significant, leftmost) bit of the final RBSP byte contains the remaining bits of SODB (if any).
[0268] ii) The next bit consists of a single bit equal to 1 (i.e., rbsp_stop_one_bit).
[0269] iii) When rbsp_stop_one_bit is not the last bit of a byte that is byte aligned, there are one or more bits equal to 0 (i.e., instances of rbsp_alignment_zero_bit) to achieve byte alignment.
[0270] In the syntax table, the suffix "_rbsp" is used to indicate syntax structures with these RBSP attributes. These structures are carried within NAL units as the contents of rbsp_byte[i] data bytes. The association between RBSP syntax structures and NAL units is specified as follows.
[0271] Note 2 – When the boundaries of the RBSP are known, the decoder can extract the SODB from the RBSP by concatenating the bits of the RBSP bytes and discarding the rbsp_stop_one_bit (i.e., the last (least significant, rightmost) bit that is equal to 1) and any bits equal to 0 that follow it (least significant, further right). The data required for the decoding process is contained in the SODB portion of the RBSP.
[0272] J.7.2.1.2 NAL Unit Header Semantics
[0273] Similar to the atlas example, a similar NAL cell type is defined for displacements, enabling similar functionality for random access, and specific NAL cells corresponding to encoded / decoded displacement data are defined. Additionally, NAL cells that can include metadata such as SEI messages are also defined.
[0274] Specifically, the supported displacement NAL element types are specified as follows:
[0275]
[0276]
[0277] J.7.2.1.3 The order of NAL units and displacement frames, and the relationship between encoded / decoded displacement frames, access units, and encoded / decoded displacement sequences.
[0278] J.7.3 Raw Byte Sequence Payload, Trailing Bits, and Byte Alignment Semantics
[0279] J.7.3.1 Semantics of RBSP (Role-Based Sequence Parameter Set)
[0280] J.7.3.1.1 Semantics of the General Displacement Sequence Parameter Set (RBSP)
[0281] dsps_sequence_parameter_set_id provides an identifier for the displacement sequence parameter set for reference by other syntax elements.
[0282] `dsps_codec_id` indicates the identifier of the codec used for compression shifting. `dsps_codec_id` must be in the range of 0 to 255 (inclusive). This codec can be identified by the level, component codec mapping SEI message as defined herein, or by means outside this document. It can be associated with a specific shift codec by the level specified in the corresponding specification, or it can be explicitly indicated by the SEI message as in the V3C specification for video sub-bitstreams.
[0283] The value of dsps_range_log2_minus2 plus 2 indicates the range of geometric displacement coordinates. dsps_range_log2_minus2 must be within the range of 0 to 3 (inclusive).
[0284] `dsps_single_dimension_flag` indicates the number of dimensions of the displacement associated with it. `dsps_single_dimension_flag` equal to 0 indicates that all three components of the displacement are used. `dsps_single_dimension_flag` equal to 1 indicates that only the normal component of the displacement is used.
[0285] The dsps_msb_align_flag indicates how to convert decoded shift samples into samples at the bit depth of the shift range.
[0286] The value of dsps_log2_max_displ_frame_order_cnt_lsb_minus4 plus 4 specifies the values of the variables Log2MaxDisplFrmOrderCntLsb and MaxDisplFrmOrderCntLsb used in the decoding process for offset frame order counting, as follows:
[0287] Log2MaxDisplFrmOrderCntLsb=
[0288] dsps_log2_max_displ_frame_order_cnt_lsb_minus4+4 (2)
[0289] MaxDisplFrmOrderCntLsb=2 Log2MaxDisplFrmOrderCntLsb (3)
[0290] The value of dsps_log2_max_displ_frame_order_cnt_lsb_minus4 must be in the range of 0 to 12 (inclusive).
[0291] The value of dsps_max_dec_displ_frame_buffering_minus1 plus 1 specifies the maximum required size of the shift frame buffer for CDS decoding, in units of shift frame storage buffer. The value of dsps_max_dec_displ_frame_buffering_minus1 must be in the range of 0 to 15 (inclusive).
[0292] A value of 0 for `dsps_long_term_ref_displ_frames_flag` indicates that no long-term reference displacement is used for inter-frame prediction of any encoded or decoded displacement frames in the CDS. A value of 1 for `dsps_long_term_ref_displ_frames_flag` indicates that a long-term reference displacement frame is available for inter-frame prediction of one or more encoded or decoded displacement frames in the CDS.
[0293] `dsps_num_ref_displ_frame_lists_in_dsps` specifies the number of `displ_ref_list_struct(rlsIdx)` syntax structures included in the displacement sequence parameter set. The value of `dsps_num_ref_displ_frame_lists_in_dsps` must be in the range of 0 to 64 (inclusive).
[0294] Note 1 – The decoder allocates a total of (dsps_num_ref_displ_frame_lists_in_dsps+1) memory for the displ_ref_list_struct(rlsIdx) syntax structure because a displ_ref_list_struct(rlsIdx) syntax structure can be directly transmitted via signal in the displacement header of the current displacement frame.
[0295] The value of dsps_extension_present_flag equal to 1 indicates that dsps_extension_count_minus1 and dsps_extension_length_minus1 exist in the displacement sequence parameter set.
[0296] Incrementing 1 to dsps_extension_count_minus1 specifies the number of extensions present in the current displacement sequence parameter set. When none exist, dsps_extension_count_minus1 is presumed to be -1.
[0297] Incrementing 1 to dsps_extension_length_minus1 specifies the length of the dsps_extension_data_byte element following this syntax element. If it does not exist, dsps_extension_length_minus1 is presumed to be equal to -1.
[0298] dsps_extension_data_byte can have any value.
[0299] J.7.3.1.2 Semantics of Displacement Levels, Layers, and Grades
[0300] The dptl_tier_flag specifies the layer context used to interpret dptl_level_idc.
[0301] The dptl_tier_flag specifies the layer context used to interpret dptl_level_idc.
[0302] `dptl_profile_codec_group_idc` indicates the codec group profile component to which the CDS conforms. The bitstream should not contain any `dptl_profile_codec_group_idc` values other than those specified herein. Other values for `dptl_profile_codec_group_idc` are reserved for future use by ISO / IEC.
[0303] `dptl_profile_toolset_idc` indicates the toolset profile component to which the CDS conforms. The bitstream should not contain any `dptl_profile_toolset_idc` values other than those specified herein. Other values for `dptl_profile_toolset_idc` are reserved for future use by ISO / IEC.
[0304] The `dptl_profile_reconstruction_idc` directive indicates the recommended CDS-compliant reconstruction profile component. The decoder may choose to use a reconstruction profile different from the one indicated in the bitstream. The bitstream should not contain any `dptl_profile_reconstruction_idc` values other than those specified herein. Other values for `dptl_profile_reconstruction_idc` are reserved for future use by ISO / IEC.
[0305] The value of dptl_reserved_zero_16bits (if present) must be equal to 0 in the bitstream conforming to this version of the document. Other values for dptl_reserved_zero_16bits are reserved for future use by ISO / IEC. The decoder must ignore the value of dptl_reserved_zero_16bits.
[0306] The dptl_reserved_0xffff_16bits value (if present) must be equal to 0xFFFF in the bitstream conforming to this version of the document. Other values for dptl_reserved_0xffff_16bits are reserved for future use by ISO / IEC. The decoder must ignore the value of dptl_reserved_0xffff_16bits.
[0307] dptl_level_idc indicates the level that the CDS conforms to. The bitstream should not contain any dptl_level_idc values other than those specified herein. Other values for dptl_level_idc are reserved for future use by ISO / IEC.
[0308] dptl_num_sub_profiles indicates the number of syntax elements in dptl_sub_profile_idc[i].
[0309] A value of 1 for dptl_extended_sub_profile_flag specifies that the dptl_sub_profile_idc[i] syntax element (if present) should be represented using 64 bits. A value of 0 for dptl_extended_sub_profile_flag specifies that the dptl_sub_profile_idc[i] syntax element (if present) should be represented using 32 bits.
[0310] dptl_sub_profile_idc[i] indicates the i-th interoperability metadata registered as specified in Recommendation ITU-T T.35, the contents of which are not specified in this document. The number of bits used to represent dptl_sub_profile_idc[i] is equal to (dptl_extended_sub_profile_flag == 0 ? 32 : 64).
[0311] A value of 1 for `dptl_toolset_constraints_present_flag` indicates that the additional structure `dptl_profile_toolset_constraints_information()` exists in the bitstream. A value of 0 for `dptl_toolset_constraints_present_flag` indicates that the structure `dptl_profile_toolset_constraints_information()` does not exist.
[0312] J.7.3.1.3 Semantic Constraint Information of Displacement Level Toolset
[0313] The `dptc_one_displacement_frame_only_flag` (if present) has the semantics specified in this document, where the profile indicated by `dptl_profile_toolset_idc` is the profile specified in this document. When it does not exist, `dptc_one_displacement_frame_only_flag` is presumed to be equal to 0.
[0314] The dptc_reserved_zero_7bits value must be equal to 0 in the bitstream conforming to this version of the document. Other values for dptc_reserved_zero_7bits are reserved for future use by ISO / IEC and should not exist in the bitstream conforming to this version of the document. Decoders conforming to this version of the document must ignore dptc_reserved_zero_7bits values other than 0.
[0315] `dptc_num_reserved_constraint_bytes` specifies the number of bytes reserved as constraint. In bitstreams conforming to this version of the document, the value of `dptc_num_reserved_constraint_bytes` must be 0. Other values of `dptc_num_reserved_constraint_bytes` are reserved for future use by ISO / IEC and should not exist in bitstreams conforming to this version of the document. Decoders conforming to this version of the document must ignore `dptc_num_reserved_constraint_bytes` values other than 0.
[0316] dptc_reserved_constraint_byte[i] can have any value. Its presence and value do not affect the consistency of the decoder with the grade specified in this version of the document. Decoders conforming to this version of the document must ignore the values of all dptc_reserved_constraint_byte[i] syntax elements.
[0317] J.7.3.2 Semantics of the Frame Parameter Set (RBSP)
[0318] J.7.3.2.1 Semantics of the General Displacement Frame Parameter Set (RBSP)
[0319] dfps_displ_sequence_parameter_set_id specifies the value of dsps_sequence_parameter_set_id for the active displacement sequence parameter set.
[0320] dfps_displ_parameter_set_id identifies the set of displacement frame parameters for reference by other syntax elements.
[0321] A value of 1 for `dfps_output_flag_present_flag` indicates that the `displ_output_flag` syntax element exists in the associated displacement header. A value of 0 for `dfps_output_flag_present_flag` indicates that the `displ_output_flag` syntax element does not exist in the associated displacement header.
[0322] Increasing 1 to dfps_num_ref_idx_default_active_minus1 specifies the estimated value of the variable NumRefIdxActive for slices where displ_num_ref_idx_active_override_flag equals 0. The value of dfps_num_ref_idx_default_active_minus1 must be in the range of 0 to 14 (inclusive).
[0323] The value of the variable MaxLtDisplFrmOrderCntLsb, which specifies the decoding process for the reference atlas frame list, is as follows:
[0324] MaxLtDisplFrmOrderCntLsb =
[0325] 2 (Log2MaxDisplFrmOrderCntLsb + dfps_additional_lt_dfoc_lsb_len) (4)
[0326] The value of dfps_additional_lt_dfoc_lsb_len must be in the range of 0 to 32 –Log2MaxDisplFrmOrderCntLsb (inclusive).
[0327] When dsps_long_term_ref_displ_frames_flag equals 0, the value of dfps_additional_lt_dfoc_lsb_len must be equal to 0.
[0328] A value of 1 for `dfps_extension_present_flag` indicates that the syntax element `dfps_extension_8bits` exists in the shift frame parameter set. A value of 0 for `dfps_extension_present_flag` indicates that the syntax element `dfps_extension_8bits` does not exist. In this version of the document, the value of `dfps_extension_present_flag` must be 0.
[0329] A value of 0 for `dfps_extension_8bits` indicates that no `dfps_extension_data_flag` syntax element exists in the DFPS RBSP syntax structure. When present, `dfps_extension_8bits` must be equal to 0 in the bitstream conforming to this version of the document. Values of `dfps_extension_8bits` that are not equal to 0 are reserved for future use by ISO / IEC. The decoder must allow values of `dfps_extension_8bits` to be non-zero and must ignore all `dfps_extension_data_flag` syntax elements in the DFPS NAL unit. When not present, the value of `dfps_extension_8bits` is presumed to be 0.
[0330] The `dfps_extension_data_flag` flag can have any value. Its presence and value do not affect the consistency of the decoder with the profile specified in this version of the document. Decoders conforming to this version of the document must ignore all `dfps_extension_data_flag` syntax elements.
[0331] The `displ_no_output_of_prior_displ_frames_flag` affects the output of previously decoded shift frames in the DDB after decoding a shift frame in a CDS AU that is not the first AU in the bitstream. When `no_output_of_prior_displ_frames_flag` is not present, its value is presumed to be 0.
[0332] The requirement for bitstream consistency is that the value of no_output_of_prior_displ_frames_flag must be the same for all displacement frames in the AU.
[0333] The value of no_output_of_prior_displ_frames_flag in the displacement header is also known as the output_of_prior_displ_frames_flag value of AU.
[0334] The displ_frame_parameter_set_id specifies the value of dfps_displ_frame_parameter_set_id for the active displacement frame parameter set of the current displacement frame.
[0335] The `dislp_type` parameter specifies the encoding / decoding type for the current shift frame according to Table 10. In bitstreams conforming to this version of the document, the value of `smh_type` must be equal to 0, 1, or 2. Other values of `smh_type` are reserved for future use by ISO / IEC. Decoders conforming to this version of the document must ignore the reserved values of `smh_type`.
[0336] Table 10 – Association between dislp_type and its name
[0337]
[0338] The `displ_output_flag` affects the bit shift output and removal process during decoding. When `displ_output_flag` is not present, it is presumed to be equal to 1.
[0339] `displ_frm_order_cnt_lsb` specifies the displacement frame order count modulo `MaxDisplFrmOrderCntLsb` for the current displacement frame. The length of the `displ_frm_order_cnt_lsb` syntax element is equal to `Log2MaxDisplFrmOrderCntLsb` bits. The value of `displ_frm_order_cnt_lsb` must be in the range of 0 to `MaxDisplFrmOrderCntLsb - 1` (inclusive).
[0340] A value of 1 for `ref_displ_frame_list_dsps_flag` indicates that the list of reference displacement frames for the current displacement frame is derived based on one of the `displ_ref_list_struct(rlsIdx)` syntax structures in the active DSPS. A value of 0 for `ref_displ_frame_list_dsps_flag` indicates that the list of reference displacement frames for the current displacement frame is derived based on the `displ_ref_list_struct(rlsIdx)` syntax structure directly included in the displacement frame header of the current displacement frame. When `dsps_num_ref_displ_frame_lists_in_dsps` equals 0, the value of `ref_displ_frame_list_dsps_flag` is presumed to be 0.
[0341] `ref_displ_frame_list_idx` specifies the index of the `displ_ref_list_struct(rlsIdx)` syntax structure used to deduce the list of reference displacement frames for the current displacement frame in the list of `displ_ref_list_struct(rlsIdx)` syntax structures included in the active DSPS. The syntax element `ref_displ_frame_list_idx` is represented by `Ceil(Log2(dsps_num_ref_displ_frame_lists_in_dsps))` bits. When it does not exist, the value of `ref_displ_frame_list_idx` is presumed to be 0. The value of `ref_displ_frame_list_idx` must be in the range of 0 to `dsps_num_ref_displ_frame_lists_in_dsps - 1` (inclusive). When ref_displ_frame_list_dsps_flag equals 1 and dsps_num_ref_displ_frame_lists_in_dsps equals 1, the value of ref_displ_frame_list_idx is presumed to be equal to 0.
[0342] The variable RlsIdx for the current atlas piece is derived as follows:
[0343] RlsIdx = ref_displ_frame_list_dsps_flag?
[0344] ref_displ_frame_list_idx : dsps_num_ref_displ_frame_lists_in_dsps
[0345] The value of additional_dfoc_lsb_present_flag[j] being equal to 1 indicates that additional_dfoc_lsb_val[j] exists for the current displacement frame. The value of additional_dfoc_lsb_present_flag[j] being equal to 0 indicates that additional_dfoc_lsb_val[j] does not exist.
[0346] additional_dfoc_lsb_val[j] specifies the value of FullFrmOrderCntLsbLt[RlsIdx][j] for the current atlas piece as follows:
[0347] FullDisplFrmOrderCntLsbLt[ RlsIdx ][ j ] =
[0348] additional_dfoc_lsb_val[j] MaxDisplFrmOrderCntLsb +dfoc_lsb_lt[ RlsIdx ][ j ]
[0349] The syntax element additional_dfoc_lsb_val[j] is represented by the bits dfps_additional_lt_dfoc_lsb_len. When it does not exist, the value of additional_dfoc_lsb_val[j] is presumed to be 0.
[0350] A value of 1 for `num_ref_idx_active_override_flag` indicates that the syntax element `num_ref_idx_active_minus1` exists for the current displacement frame. A value of 0 for `num_ref_idx_active_override_flag` indicates that the syntax element `num_ref_idx_active_minus1` does not exist. If `num_ref_idx_active_override_flag` does not exist, its value is presumed to be 0.
[0351] `num_ref_idx_active_minus1` is used to derive the variable `NumRefIdxActive` for the current displacement frame, as specified in Equation 5. The value of `num_ref_idx_active_minus1` must be in the range of 0 to 14 (inclusive).
[0352] When the current displacement frame is a P_DISPLACEMENT displacement frame, num_ref_idx_active_override_flag is equal to 1, and num_ref_idx_active_minus1 does not exist, num_ref_idx_active_minus1 is presumed to be equal to 0.
[0353] The variable NumRefIdxActive is derived as follows:
[0354] if( displ_type == P_DISPLACEMENT ) {
[0355] if(num_ref_idx_active_override_flag == 1)
[0356] NumRefIdxActive = num_ref_idx_active_minus1 + 1 (5)
[0357] else {
[0358] if( num_ref_entries[ RlsIdx ] >= dfps_num_ref_idx_default_active_minus1 + 1 )
[0359] NumRefIdxActive = dfps_num_ref_idx_default_active_minus1 + 1
[0360] else
[0361] NumRefIdxActive = num_ref_entries[RlsIdx]
[0362] }
[0363] }
[0364] else
[0365] NumRefIdxActive = 0
[0366] NumRefIdxActive minus 1 specifies the maximum value of the displacement reference frame index that can be used to decode the current displacement frame.
[0367] J.7.3.3 Reference List Structure Semantics
[0368] `drl_num_ref_entries[ rlsIdx ]` specifies the number of entries in the `displ_ref_list_struct( rlsIdx )` syntax structure, where `rlsIdx` is the index of the displacement frame reference list. For `P_DISPLACEMENT`, the value of `num_ref_entries[ rlsIdx ]` must be in the range of 1 to `dsps_max_dec_displ_frame_buffering_minus1 + 1`. Otherwise, the value of `num_ref_entries[ rlsIdx ]` must be in the range of 0 to `dsps_max_dec_displ_frame_buffering_minus1 + 1`.
[0369] The value of `drl_st_ref_displ_frame_flag[rlsIdx][i]` equal to 1 indicates that the i-th entry in the `displ_ref_list_struct(rlsIdx)` syntax structure is a short-term reference displacement frame entry. The value of `st_ref_displ_frame_flag[rlsIdx][i]` equal to 0 indicates that the i-th entry in the `displ_ref_list_struct(rlsIdx)` syntax structure is a long-term reference displacement frame entry. When it does not exist, the value of `drl_st_ref_displ_frame_flag[rlsIdx][i]` is presumed to be 1.
[0370] The variable NumLtrDisplFrmEntries[rlsIdx] is derived as follows:
[0371] NumLtrDisplFrmEntries[rlsIdx] = 0
[0372] for( i = 0; i < drl_num_ref_entries[ rlsIdx ]; i++ )
[0373] if( !drl_st_ref_displ_frame_flag[ rlsIdx ][ i ] ) (6)
[0374] NumLtrDisplFrmEntries[ rlsIdx ]++
[0375] drl_abs_delta_dfoc_st[ rlsIdx ][ i ], when the i-th entry is the first short-term reference displacement frame entry in the displ_ref_list_struct( rlsIdx ) syntax structure, specifies the absolute difference between the displacement frame order count value of the current displacement frame referenced by the i-th entry, or, when the i-th entry is a short-term reference displacement frame entry but not the first short-term reference displacement frame entry in the displ_ref_list_struct( rlsIdx ) syntax structure, specifies the absolute difference between the displacement frame order count value of the displacement frame referenced by the i-th entry and the displacement frame order count value of the displacement frame referenced by the previous short-term reference displacement frame entry in the displ_ref_list_struct( rlsIdx ) syntax structure.
[0376] The value of drl_abs_delta_dfoc_st[ rlsIdx ][ i ] must be between 0 and 2. 15 -1 (including boundary values).
[0377] The value of `drl_straf_entry_sign_flag[ rlsIdx ][ i ]` being equal to 1 indicates that the i-th entry in the syntax structure `displ_ref_list_struct( rlsIdx )` has a value greater than or equal to 0. The value of `drl_straf_entry_sign_flag[ rlsIdx ][ i ]` being equal to 0 indicates that the i-th entry in the syntax structure `displ_ref_list_struct( rlsIdx )` has a value less than 0. When it does not exist, the value of `drl_straf_entry_sign_flag[ rlsIdx ][ i ]` is presumed to be equal to 1.
[0378] The list DeltaDfocSt[rlsIdx][i] is derived as follows:
[0379] for( i = 0; i < drl_num_ref_entries[ rlsIdx ]; i++ )
[0380] if( drl_st_ref_displ_frame_flag[ rlsIdx ][ i ] )
[0381] DeltaDfocSt[rlsIdx][i] =
[0382] (2) drl_straf_entry_sign_flag[ rlsIdx ][ i ] – 1 ) drl_abs_ delta_dfoc_st[ rlsIdx ][ i ] (7)
[0383] else
[0384] DeltaDfocSt[ rlsIdx ][ i ] = 0
[0385] `drl_dfoc_lsb_lt[ rlsIdx ][ i ]` specifies the displacement frame sequence count modulo `MaxDisplFrmOrderCntLsb` for the displacement frame referenced by the `i`th entry in the `displ_ref_list_struct( rlsIdx )` syntax structure. The length of the `drl_dfoc_lsb_lt[ rlsIdx ][ i ]` syntax element is `Log2MaxDisplFrmOrderCntLsb` bits.
[0386] J.7.3.4 Semantics of Displacement Data Units
[0387] The `displ_intra_unit(unitSize)` method contains a stream of displacement units of size `unitSize` (in bytes), as an ordered stream of bytes or bits, where the positions of the unit boundaries can be identified from patterns in the data. The format of this displacement unit stream is determined by 4CC codes as defined by `dptl_profile_codec_group_idc` or by SEI message identifiers mapped by component codecs.
[0388] The `displ_inter_unit(unitSize)` method contains a stream of displacement units of size `unitSize` (in bytes), as an ordered stream of bytes or bits, where the positions of the unit boundaries can be identified from patterns in the data. The format of this displacement unit stream is determined by 4CC codes as defined by `dptl_profile_codec_group_idc` or by SEI message identifiers mapped by component codecs.
[0389] J.7.3.5 Semantics of Displacement Intra-Frame Data Units
[0390] The arithmetic decoding engine is a context-separated binary arithmetic decoder that performs binary renormalization and produces binary output.
[0391] The displacement value is derived from arithmetic decoding.
[0392] diu_last_sig_coeff[k] indicates the index of the last position of the non-zero displacement coefficient level in the k-th component.
[0393] diu_coded_block_flag[k][b] indicates whether the block at index b has any non-zero displacement coefficient level in the k-th component (when 1), or not (when 0).
[0394] diu_coded_subblock_flag[k][b][s] indicates whether the subblock at index s of the block at index b has any non-zero displacement coefficient level in the k-th component (when 1), or no (when 0).
[0395] diu_coeff_abs_level_gt0[k][b][s][v] indicates whether the k-th component of the displacement coefficient level associated with the vertex of the sub-block at index v of index s of the block at index b has an absolute value greater than 0 (when 1), or does not have one (when 0).
[0396] `diu_coeff_abs_level_gt1[k][b][s][v]` indicates whether the k-th component of the displacement coefficient level associated with the vertex of the sub-block at index `s` of the block at index `b` has an absolute value greater than 1 (when 1), or does not (when 0). If `diu_coeff_abs_level_gt1[k][b][s][v]` does not exist, it must be presumed to be equal to 0.
[0397] `diu_coeff_sign[k][b][s][v]` indicates whether the k-th component of the displacement coefficient level associated with the vertex at index v of the sub-block at index s of the block at index b has a positive sign (when 1) or no sign (when 0). If `diu_coeff_sign[k][b][s][v]` does not exist, it must be presumed to be equal to 1.
[0398] `diu_coeff_abs_level_rem[k][b][s][v]` indicates the absolute value of the k-th component of the displacement coefficient level associated with the vertex at index v of the block with index b minus 2. If `diu_coeff_abs_level_rem[k][b][s][v]` does not exist, it must be presumed to be equal to 0.
[0399] J.7.3.6 Semantics of Inter-Frame Displacement Data Units
[0400] The arithmetic decoding engine is a context-separated binary arithmetic decoder that performs binary renormalization and produces binary output.
[0401] The displacement residuals are derived from arithmetic decoding. Same as A.7.3.5.
[0402] 3. The technical problem solved by the disclosed technical solution
[0403] The example design for dynamic mesh encoding and decoding has the following problems:
[0404] First, it's unclear how to process the displacement data encoded and decoded into 4:2:2 video.
[0405] Second, the sub-block size of AC-based encoding and decoding should be constrained.
[0406] Third, the encoding / decoding type and / or reference index of AC-based displacement encoding / decoding may not match the encoding / decoding type and / or reference index of subgrid encoding / decoding, resulting in unnecessary decoding delays.
[0407] Fourth, when the displacement is encoded as 4:0:0 video, only one displacement component can be sent in the example design.
[0408] Fifth, when encoding and decoding displacement to 4:4:4 video, in the example design, the chroma channel can be used to transmit non-zero displacement information. However, in some system designs, the chroma information may be distorted due to format conversion or other post-processing, resulting in poor encoding and decoding performance.
[0409] Sixth, in the example design, the packing method depends on the color format used in the codec. However, in some hardware decoders, output in a specific color format cannot be guaranteed.
[0410] Seventh, there are multiple sub-bitstreams, including atlas sub-bitstreams, base grid sub-bitstreams, displacement sub-bitstreams, and attribute sub-bitstreams. Currently, these sub-bitstreams lack synchronization, which may lead to significant latency.
[0411] Eighth, the lack of synchronization between different sub-bit streams also makes it difficult to support the time-domain layer.
[0412] 4. List of solutions and implementation examples
[0413] The detailed designs below should be considered as examples for explaining general concepts. These examples should not be interpreted in a narrow sense. Furthermore, these examples can be combined in any way. Combinations between this disclosure and other disclosures also apply.
[0414] 1. To solve problem 1, when the displacement data is encoded and decoded as 4:2:2 video, it can be processed in the same way as 4:2:0 video.
[0415] a. In one example, when DisplacementDim equals 1, the first displacement component is derived from the first color component of the video, and the second and third displacement components are assumed to be 0.
[0416] b. In one example, when DisplacementDim equals 3, the first displacement component, the second displacement component, and the third displacement component are derived from the first color component of the video.
[0417] 2. To solve problem 2, the sub-block size of AC-based displacement encoding and decoding must be constrained.
[0418] a. In one example, the sub-block size of AC-based displacement encoding and decoding must be greater than 0.
[0419] b. In one example, the sub-block size of AC-based displacement encoding and decoding must be greater than 1.
[0420] 3. To address problem 3, the encoding / decoding type of the AC-based displacement encoding / decoding can be aligned with the encoding / decoding type of the submesh.
[0421] a. In one example, when smh_type indicates intra-frame encoding (e.g., I_SUBMESH), it is not necessary to transmit dislp_type via signals, and dislp_type can be presumed to be intra-frame encoding (e.g., I_DISPLACEMENT).
[0422] i. Alternatively, when dislp_type indicates intra-frame encoding / decoding (e.g., I_DISPLACEMENT), smh_type does not need to be transmitted via signaling, and smh_type can be presumed to be intra-frame encoding / decoding (e.g., I_SUBMESH).
[0423] b. In one example, when smh_type indicates inter-frame encoding / decoding (e.g., P_SUBMESH or SKIP_SUBMESH), it is not necessary to transmit dislp_type via signaling, and dislp_type can be presumed to be inter-frame encoding / decoding (e.g., P_DISPLACEMENT).
[0424] i. Alternatively, when dislp_type indicates inter-frame encoding / decoding (e.g., P_DISPLACEMENT), smh_type can be presumed to be intra-frame encoding / decoding (e.g., P_SUBMESH or SKIP_SUBMESH).
[0425] 4. To solve problem 3, when the subgrid at time t1 uses the subgrid at time t2 as a reference, the displacement data at time t1 can use only the displacement data at time t2 as a reference.
[0426] a. In one example, the reference index of the displacement is set to be equal to the reference index of the submesh.
[0427] b. In one example, the displacement reference list structure is set to be the same as the base mesh reference list structure, for example, bmesh_ref_list_struct.
[0428] 5. To solve problem 4, encoding and decoding the displacement data into 4:0:0 video can be processed in the same way as 4:2:0 video.
[0429] a. In one example, when the displacement data is encoded as a 4:0:0 video and DisplacementDim equals 1, the first displacement component is derived from the first color component of the video, and the second and third displacement components are assumed to be 0.
[0430] b. In one example, when the displacement data is encoded as a 4:0:0 video and DisplacementDim equals 3, the first displacement component, the second displacement component, and the third displacement component are derived from the first color component of the video.
[0431] 6. To solve problem 5, when the displacement information is encoded and decoded into a 4:4:4 video, only the luminance channel is used to transmit the displacement information.
[0432] a. In one example, when the displacement has three-dimensional information and is encoded and decoded into video, all of this three-dimensional information is in the luminance channel.
[0433] b. In one
[0434] 7. To address problem 6, displacement information is packaged into the video, regardless of the color format.
[0435] a. In one example, regardless of the color format, when displacement information is packed into the video, only the luminance channel is used to transmit information.
[0436] b. In one example, regardless of the color format, when DisplacementDim equals 1, the first displacement component is derived from the first color component of the video, while the second and third displacement components are assumed to be 0.
[0437] c. In one example, regardless of the color format, when DisplacementDim equals 3, the first displacement component, the second displacement component, and the third displacement component are derived from the first color component of the video.
[0438] 8. To solve problem 3, displacement data at time t1 can be used as a reference only if the base grid corresponding to time t2 is in the reference list of the base grid corresponding to time t1.
[0439] a. Alternatively, displacement data at time t1 may be used as a reference only if the base grid corresponding to time t2 is a reference to the base grid corresponding to time t1.
[0440] 9. To solve problem 3, the base grid at time t1 can be used as a reference only if the displacement corresponding to time t2 is in the reference list of displacements corresponding to time t1.
[0441] a. Alternatively, displacement data at time t1 may be used as a reference only if the base grid corresponding to time t2 is a reference to the base grid corresponding to time t1.
[0442] 10. To address issue 3, the displacement and / or base mesh reference list is set to correspond to the atlas reference list.
[0443] a. In one example, the base grid and / or displacement data at time t1 can be used as a reference only if the atlas corresponding to time t2 is in the reference list of the atlas corresponding to time t1.
[0444] b. Alternatively, the base grid and / or displacement data at time t1 may be used as a reference only if the atlas corresponding to time t2 is a reference to the atlas corresponding to time t1.
[0445] 11. To address problem 7, a method for synchronizing all sub-bitstreams within a system, such as video-based point cloud compression (V-PCC) or video-based dynamic mesh coding and decoding (V-DMC), is proposed.
[0446] 12. In one example, all sub-bitstreams are required to have the same reference list.
[0447] a. In one example, it is required that all sub-bitstreams have the same reference structure.
[0448] b. In one example, one or more syntax elements are used to indicate the maximum allowed number of decode buffers for the entire decoding system and for each sub-bitstream, and require that the number of decode buffers should not exceed the maximum number of decode buffers for the entire decoding system.
[0449] c. In one example, one or more syntax elements are used to indicate the maximum allowed number of reordered frames for the entire decoding system and for each sub-bitstream, and require that the number of reordered frames should not exceed that maximum allowed number.
[0450] i. In one example, one or more syntax elements are used to indicate the maximum allowed number of atlas frames with AtlasFrameOutputFlag equal to 1, preceding any atlas frame with AtlasFrameOutputFlag equal to 1 in the output order at a particular time-domain layer.
[0451] d. In one example, one or more syntax elements are used to indicate the maximum allowed number of delayed frames for the entire decoding system and for each sub-bitstream, and require that the number of delayed frames should not exceed that maximum allowed number.
[0452] i. In one example, one or more syntax elements are used to indicate the maximum allowed number of atlas frames with AtlasFrameOutputFlag equal to 1, in output order before any atlas frame with AtlasFrameOutputFlag equal to 1 and in decoding order after that frame with AtlasFrameOutputFlag equal to 1, at a particular time-domain layer.
[0453] 13. To solve problem 8, it is required that all sub-bit streams at a specific time have the same time domain ID.
[0454] 5. Examples
[0455] The following are some example embodiments of the aspects outlined in Section 4.
[0456] The most relevant parts that have been added or modified are shown in bold, and some deleted parts are shown in bold and italics. There may be other editable changes that are not indicated.
[0457] The following text changes are based on WD 3.0 based on V-DMC [3].
[0458] 5.1 Example 1
[0459] This embodiment refers to item 1, which is outlined in Section 4.
[0460] 11.5 Inverse Image Packaging of Wavelet Coefficients
[0461] ...
[0462] The wavelet coefficient depacking process is performed as follows:
[0463] pixelsPerBlock = blockSize blockSize
[0464] widthInBlocks = width / blockSize
[0465] shift = (1 << bitDepth) >> 1
[0466] blockCount = (verCoordCount + pixelsPerBlock - 1) / pixelsPerBlock
[0467] heightInBlocks = (blockCount + widthInBlocks - 1) / widthInBlocks
[0468] origHeight = heightInBlocks blockSize
[0469] paddedHeight = height - 3 origHeight
[0470] if ( !asps_vdmc_ext_1d_displacement_flag )
[0471] start = (paddedHeight + origHeight) width - 1
[0472] else
[0473] start = (width height) - 1
[0474] for( v = 0; v < verCoordCount; v++ ) {
[0475] v0 = asps_vdmc_ext_packing_method ? start - v : v
[0476] blockIndex = v0 / pixelsPerBlock
[0477] indexWithinBlock = v0 % pixelsPerBlock
[0478] x0 = ( blockIndex % widthInBlocks ) blockSize
[0479] y0 = ( blockIndex / widthInBlocks ) blockSize
[0480] ( x, y ) = computeMorton2D( indexWithinBlock )
[0481] x1 = x0 + x
[0482] y1 = y0 + y
[0483] for( d = 0; d < DisplacementDim; d++ ) {
[0484] if ( DecGeoChromaFormat == 4:2:0 || DecGeoChromaFormat == 4:2:2 ) {
[0485] dispQuantCoeffArray[ v0 ][ d ] =
[0486] dispQuantCoeffFrame[ x1 ][ d origHeight + y1 ][ d0 ] - shift
[0487] } else {
[0488] dispQuantCoeffArray[ v0 ][ d ] = dispQuantCoeffFrame[ x1][ y1 ][ d ] - shift
[0489] }
[0490] }
[0491] }
[0492] 5.2 Example 2
[0493] This embodiment refers to item 5, which is outlined in Section 4.
[0494] 11.5 Inverse Image Packaging of Wavelet Coefficients
[0495] ...
[0496] The wavelet coefficient depacking process is performed as follows:
[0497] pixelsPerBlock = blockSize blockSize
[0498] widthInBlocks = width / blockSize
[0499] shift = (1 << bitDepth) >> 1
[0500] blockCount = (verCoordCount + pixelsPerBlock - 1) / pixelsPerBlock
[0501] heightInBlocks = (blockCount + widthInBlocks - 1) / widthInBlocks
[0502] origHeight = heightInBlocks blockSize
[0503] paddedHeight = height - 3 origHeight
[0504] if ( !asps_vdmc_ext_1d_displacement_flag )
[0505] start = (paddedHeight + origHeight) width - 1
[0506] else
[0507] start = (width height) - 1
[0508] for( v = 0; v < verCoordCount; v++ ) {
[0509] v0 = asps_vdmc_ext_packing_method ? start - v : v
[0510] blockIndex = v0 / pixelsPerBlock
[0511] indexWithinBlock = v0 % pixelsPerBlock
[0512] x0 = ( blockIndex % widthInBlocks ) blockSize
[0513] y0 = ( blockIndex / widthInBlocks ) blockSize
[0514] ( x, y ) = computeMorton2D( indexWithinBlock )
[0515] x1 = x0 + x
[0516] y1 = y0 + y
[0517] for( d = 0; d < DisplacementDim; d++ ) {
[0518] if ( DecGeoChromaFormat == 4:2:0 || DecGeoChromaFormat == 4:0:0 ) {
[0519] dispQuantCoeffArray[ v0 ][ d ] =
[0520] dispQuantCoeffFrame[ x1 ][ d origHeight + y1 ][ d0 ] - shift
[0521] } else {
[0522] dispQuantCoeffArray[ v0 ][ d ] = dispQuantCoeffFrame[ x1][ y1 ][ d ] - shift
[0523] }
[0524] }
[0525] }
[0526] 5.3 Example 3
[0527] This embodiment refers to items 1 and 5 as outlined in Section 4.
[0528] 11.5 Inverse Image Packaging of Wavelet Coefficients
[0529] ...
[0530] The wavelet coefficient depacking process is performed as follows:
[0531] pixelsPerBlock = blockSize blockSize
[0532] widthInBlocks = width / blockSize
[0533] shift = (1 << bitDepth) >> 1
[0534] blockCount = (verCoordCount + pixelsPerBlock - 1) / pixelsPerBlock
[0535] heightInBlocks = (blockCount + widthInBlocks - 1) / widthInBlocks
[0536] origHeight = heightInBlocks blockSize
[0537] paddedHeight = height - 3 origHeight
[0538] if ( !asps_vdmc_ext_1d_displacement_flag )
[0539] start = (paddedHeight + origHeight) width - 1
[0540] else
[0541] start = (width height) - 1
[0542] for( v = 0; v < verCoordCount; v++ ) {
[0543] v0 = asps_vdmc_ext_packing_method ? start - v : v
[0544] blockIndex = v0 / pixelsPerBlock
[0545] indexWithinBlock = v0 % pixelsPerBlock
[0546] x0 = ( blockIndex % widthInBlocks ) blockSize
[0547] y0 = ( blockIndex / widthInBlocks ) blockSize
[0548] ( x, y ) = computeMorton2D( indexWithinBlock )
[0549] x1 = x0 + x
[0550] y1 = y0 + y
[0551] for( d = 0; d < DisplacementDim; d++ ) {
[0552] if ( DecGeoChromaFormat == 4:2:0 DecGeoChromaFormat != 4:4:4 ) {
[0553] dispQuantCoeffArray[ v0 ][ d ] =
[0554] dispQuantCoeffFrame[ x1 ][ d origHeight + y1 ][ d0 ] - shift
[0555] } else {
[0556] dispQuantCoeffArray[ v0 ][ d ] = dispQuantCoeffFrame[ x1][ y1 ][ d ] - shift
[0557] }
[0558] }
[0559] }
[0560] 5.4 Example 4
[0561] This embodiment refers to item 6, which is outlined in Section 4.
[0562] The following text changes are based on WD 4.0 based on V-DMC [5].
[0563] 11.3 Inverse Image Packaging of Wavelet Coefficients
[0564] ...
[0565] The requirement for bitstream consistency is that when DecGeoChromaFormat equals 4:0:0, asps_vdmc_ext_1d_displacement_flag must be equal to 1. Another requirement for bitstream consistency is that when DecGeoChromaFormat equals 4:4:4, asps_vdmc_ext_1d_displacement_flag must be equal to 0.
[0566] ...
[0567] The wavelet coefficient depacking process is performed as follows:
[0568] pixelsPerBlock = blockSize blockSize
[0569] widthInBlocks = width / blockSize
[0570] shift = (1 << bitDepth) >> 1
[0571] lodExtraPixels = 0
[0572] numLod = min( 1, subdivisionIterationCount + 1 )
[0573] if( numLod == 1 ) {
[0574] numBlocksPerLod[ 0 ] = ( verCoordCount + pixelsPerBlock - 1) / pixelsPerBlock
[0575] lodExtraPixels = lodExtraPixels +
[0576] ( numBlocksPerLod[ 0 ] pixelsPerBlock - verCoordCount )
[0577] vStart[ 0 ] = 0
[0578] vEnd [ 0 ] = verCoordCount
[0579] startBlock[ 0 ] = 0
[0580] } else {
[0581] numBlocksPerLod[ 0 ] =
[0582] ( levelOfDetailVertexCounts[ 0 ] + pixelsPerBlock -1) / pixelsPerBlock
[0583] lodExtraPixels = lodExtraPixels +
[0584] ( numBlocksPerLod[ 0 ] pixelsPerBlock - levelOfDetailVertexCounts[ 0 ] )
[0585] vStart[ 0 ] = 0
[0586] vEnd [ 0 ] = levelOfDetailVertexCounts[ 0 ]
[0587] startBlock[ 0 ] = 0
[0588] for( i= 1; i < numLods; i++ ){
[0589] numPointsInLod =
[0590] levelOfDetailVertexCounts[ i ] -levelOfDetailVertexCounts[ i - 1 ]
[0591] numBlocksPerLod[ i ] = ( numPointsInLod + pixelsPerBlock - 1) / pixelsPerBlock
[0592] lodExtraPixels = lodExtraPixels +
[0593] ( numBlocksPerLod[ i ] pixelsPerBlock - numPointsInLod )
[0594] vStart[ i ] = levelOfDetailVertexCounts[ i - 1 ]
[0595] vEnd[ i ] = levelOfDetailVertexCounts[ i ]
[0596] startBlock[ i ] = numBlocksPerLod[ i - 1 ] + startBlock[ i -1 ]
[0597] }
[0598] }
[0599] blockCount = (verCoordCount + lodExtraPixels + pixelsPerBlock - 1) / pixelsPerBlock
[0600] heightInBlocks = (blockCount + widthInBlocks - 1) / widthInBlocks
[0601] origHeight = heightInBlocks blockSize
[0602] totalBlocksInVideoFrame = ( width origHeight ) / pixelsPerBlock
[0603] for( lodIdx = 0; lodIdx < numLod; lodIdx++ ) {
[0604] for( v = vStart[ lodIdx ]; v < vEnd[ lodIdx ]; v++ ) { blockIndex= ( v – vStart[ lodIdx ] ) / pixelsPerBlock+ startBlock[ lodIdx ]
[0605] indexWithinBlock = ( v – vStart[ lodIdx ] ) % pixelsPerBlock
[0606] if( asps_vdmc_ext_packing_method ){
[0607] blockIndex = totalBlocksInVideoFrame – 1 - blockIndex
[0608] indexWithinBlock = pixelsPerBlock – 1 - indexWithinBlock
[0609] }
[0610] x0 = ( blockIndex % widthInBlocks ) blockSize
[0611] y0 = ( blockIndex / widthInBlocks ) blockSize
[0612] ( x, y ) = computeMorton2D( indexWithinBlock )
[0613] x1 = x0 + x
[0614] y1 = y0 + y
[0615] for( d = 0; d < DisplacementDim; d++ ) {
[0616] if ( DecGeoChromaFormat == 4:2:0 || DecGeoChromaFormat ==4:2:2 || DecGeoChromaFormat == 4:4:4 ) {
[0617] dispQuantCoeffArray[ v ][ d ] =
[0618] dispQuantCoeffFrame[ x1 ][ d origHeight + y1 ][ d0 ] – shift
[0619] } else {
[0620] dispQuantCoeffArray[ v ][ d ] =
[0621] dispQuantCoeffFrame[ x1 ][ y1 ][ d ] – shift
[0622] }
[0623] }
[0624] }
[0625] }
[0626] 5.5 Example 5
[0627] This embodiment refers to item 7, which is outlined in Section 4.
[0628] The following text changes are based on WD 4.0 based on V-DMC [5].
[0629] 11.3 Inverse Image Packaging of Wavelet Coefficients
[0630] ...
[0631] The requirement for bitstream consistency is that when DecGeoChromaFormat equals 4:0:0, asps_vdmc_ext_1d_displacement_flag must be equal to 1. Another requirement for bitstream consistency is that when DecGeoChromaFormat equals 4:4:4, asps_vdmc_ext_1d_displacement_flag must be equal to 0.
[0632] ...
[0633] The wavelet coefficient depacking process is performed as follows:
[0634] pixelsPerBlock = blockSize blockSize
[0635] widthInBlocks = width / blockSize
[0636] shift = (1 << bitDepth) >> 1
[0637] lodExtraPixels = 0
[0638] numLod = min(1, subdivisionIterationCount + 1)
[0639] if (numLod == 1) {
[0640] numBlocksPerLod[ 0 ] = ( verCoordCount + pixelsPerBlock - 1) / pixelsPerBlock
[0641] lodExtraPixels = lodExtraPixels +
[0642] (numBlocksPerLod[0]) pixelsPerBlock - verCoordCount )
[0643] vStart[0] = 0
[0644] vEnd[0] = verCoordCount
[0645] startBlock[0] = 0
[0646] } else {
[0647] numBlocksPerLod[ 0 ] =
[0648] ( levelOfDetailVertexCounts[ 0 ] + pixelsPerBlock -1) / pixelsPerBlock
[0649] lodExtraPixels = lodExtraPixels +
[0650] ( numBlocksPerLod[ 0 ] pixelsPerBlock - levelOfDetailVertexCounts[ 0 ] )
[0651] vStart[ 0 ] = 0
[0652] vEnd [ 0 ] = levelOfDetailVertexCounts[ 0 ]
[0653] startBlock[ 0 ] = 0
[0654] for( i= 1; i < numLods; i++ ){
[0655] numPointsInLod =
[0656] levelOfDetailVertexCounts[ i ] -levelOfDetailVertexCounts[ i - 1 ]
[0657] numBlocksPerLod[ i ] = ( numPointsInLod + pixelsPerBlock - 1) / pixelsPerBlock
[0658] lodExtraPixels = lodExtraPixels +
[0659] ( numBlocksPerLod[ i ] pixelsPerBlock - numPointsInLod )
[0660] vStart[ i ] = levelOfDetailVertexCounts[ i - 1 ]
[0661] vEnd[ i ] = levelOfDetailVertexCounts[ i ]
[0662] startBlock[ i ] = numBlocksPerLod[ i - 1 ] + startBlock[ i -1 ]
[0663] }
[0664] }
[0665] blockCount = (verCoordCount + lodExtraPixels + pixelsPerBlock - 1) / pixelsPerBlock
[0666] heightInBlocks = (blockCount + widthInBlocks - 1) / widthInBlocks
[0667] origHeight = heightInBlocks blockSize
[0668] totalBlocksInVideoFrame = ( width origHeight ) / pixelsPerBlock
[0669] for( lodIdx = 0; lodIdx < numLod; lodIdx++ ) {
[0670] for( v = vStart[ lodIdx ]; v < vEnd[ lodIdx ]; v++ ) { blockIndex= ( v – vStart[ lodIdx ] ) / pixelsPerBlock+ startBlock[ lodIdx ]
[0671] indexWithinBlock = ( v – vStart[ lodIdx ] ) % pixelsPerBlock
[0672] if( asps_vdmc_ext_packing_method ){
[0673] blockIndex = totalBlocksInVideoFrame – 1 - blockIndex
[0674] indexWithinBlock = pixelsPerBlock – 1 - indexWithinBlock
[0675] }
[0676] x0 = ( blockIndex % widthInBlocks ) blockSize
[0677] y0 = ( blockIndex / widthInBlocks ) blockSize
[0678] ( x, y ) = computeMorton2D( indexWithinBlock )
[0679] x1 = x0 + x
[0680] y1 = y0 + y
[0681] for( d = 0; d < DisplacementDim; d++ ) {
[0682] if ( DecGeoChromaFormat == 4:2:0 || DecGeoChromaFormat ==4:2:2 ) {
[0683] dispQuantCoeffArray[ v ][ d ] =
[0684] dispQuantCoeffFrame[ x1 ][ d origHeight + y1 ][ d0 ] – shift
[0685] } else {
[0686] dispQuantCoeffArray[ v ][ d ] =
[0687] dispQuantCoeffFrame[ x1 ][ y1 ][ d ] – shift
[0688] }
[0689] }
[0690] }
[0691] }
[0692] 5.6 Example 6
[0693] This embodiment refers to item 8, which is outlined in Section 4.
[0694] The following text changes are based on WD 4.0 based on V-DMC [5].
[0695] J.7.1.2.1.1 General Displacement Sequence Parameter Set (RBSP) Syntax
[0696]
[0697] 5.7 Example 7
[0698] This embodiment refers to item 9, which is outlined in Section 4.
[0699] The following text changes are based on WD 4.0 based on V-DMC [5].
[0700] H.8.1.3.1.1 General Basic Mesh Sequence Parameter Set (RBSP) Syntax
[0701]
[0702]
[0703] 5.8 Example 8
[0704] This embodiment refers to item 9, which is outlined in Section 4.
[0705] The following text changes are based on WD 4.0 based on V-DMC [5].
[0706] H.8.1.3.1.1 General Basic Mesh Sequence Parameter Set (RBSP) Syntax
[0707]
[0708]
[0709] J.7.1.2.1.1 General Displacement Sequence Parameter Set (RBSP) Syntax
[0710]
[0711] 6. References
[0712] [1]MPEG technical requirements, “CfP for. Dynamic Mesh Coding,” ISO / IEC JTC 1 / SC 29 / WG 2 doc. no. N145, in Oct. 2021.
[0713] [2]K. Mammou, J. Kim, A. Tourapis and D. Podborski, "[V-CG] Apple'sDynamic Mesh Coding CfP Response," ISO / IEC JTC 1 / SC 29 / WG 7 doc. no. m59281, in Apr. 2022.
[0714] [3]MPEG output document, “WD 3.0 of V-DMC,” ISO / IEC JTC 1 / SC 29 / WG 7doc. no. N00611, in Apr. 2023.
[0715] [4] C. Huang, X. Xu, X. Zhang, J. Tian and S. Liu, “Investigation of video coding of motion fields,” ISO / IEC JTC 1 / SC 29 / WG 7 doc. no. m61005, inJul. 2022.
[0716] [5]MPEG output document, “WD 4.0 of V-DMC,” ISO / IEC JTC 1 / SC 29 / WG 7doc. no. N00680, in Sep. 2023.
[0717] Figure 3 is a block diagram illustrating an example video processing system 4000 in which various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, Passive Optical Networking (PON), etc.) and wireless interfaces (such as Wi-Fi or cellular interfaces).
[0718] System 4000 may include an encoding component 4004 capable of implementing the various encoding / decoding or encoding methods described in this disclosure. Encoding component 4004 can reduce the average bit rate from the video input 4002 to the output of encoding component 4004 to produce an encoded representation of the video. Encoding techniques are therefore sometimes referred to as video compression or video transcoding techniques. The output of encoding component 4004 may be stored or transmitted via a communication connection such as that represented by component 4006. The bitstream (or encoded) representation of the video received at input 4002, whether stored or transmitted via communication, may be used by component 4008 to generate pixel values or displayable video that is sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “encoding” operations or tools, it is understood that encoding tools or operations are used by encoders, and corresponding decoding tools or operations that reverse the encoded result will be performed by decoders.
[0719] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronic Devices (IDE), etc. The technologies described in this disclosure can be embodied in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.
[0720] Figure 4 is a block diagram of an example video processing apparatus 4100. Apparatus 4100 can be used to implement one or more methods described herein. Apparatus 4100 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Apparatus 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. The processors(multiple) 4102 can be configured to implement one or more methods described herein. The memories(multiple) 4104 can be used to store data and code for implementing the methods and techniques described herein. The video processing circuitry 4106 can be used to implement some of the techniques described herein in hardware circuitry. In some embodiments, the video processing circuitry 4106 may be at least partially included in the processor 4102, such as a graphics coprocessor.
[0721] Figure 5 is a flowchart of an example method 4200 for video processing. Method 4200 includes: in block 4202, determining that when the displacement data is encoded and decoded in a 4:4:4 video format, only the luminance channel is used to transmit the displacement data. In block 4204, performing a conversion between visual media data and a bitstream based on the displacement data. The conversion may include encoding at the encoder, decoding at the decoder, or a combination thereof.
[0722] It should be noted that method 4200 can be implemented in an apparatus for processing video data, including a processor and a non-transitory memory having instructions thereon, such as a video encoder 4400, a video decoder 4500, and / or an encoder 4600. In this case, the instructions cause the processor to execute method 4200 when executed by the processor. Furthermore, method 4200 can be executed by a non-transitory computer-readable medium including a computer program product for use by a video codec device. The computer program product includes computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec device to execute method 4200 when executed by a processor.
[0723] Figure 6 is a block diagram illustrating an example video encoding / decoding system 4300 that can utilize the techniques of this disclosure. The video encoding / decoding system 4300 may include a source device 4310 and a target device 4320. The source device 4310 generates encoded video data, and this source device 4310 may be referred to as a video encoding device. The target device 4320 can decode the encoded video data generated by the source device 4310, and this target device 4320 may be referred to as a video decoding device.
[0724] Source device 4310 may include video source 4312, video encoder 4314, and input / output (I / O) interface 4316. Video source 4312 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits forming a coded representation of the video data. The bitstream may include encoded pictures and associated data. Encoded pictures are coded representations of pictures. Associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 4320 via network 4330 through I / O interface 4316. Encoded video data may also be stored on storage medium / server 4340 for access by target device 4320.
[0725] Target device 4320 may include I / O interface 4326, video decoder 4324, and display device 4322. I / O interface 4326 may include a receiver and / or a modem. I / O interface 4326 may acquire encoded video data from source device 4310 or storage medium / server 4340. Video decoder 4324 may decode the encoded video data. Display device 4322 may display the decoded video data to a user. Display device 4322 may be integrated with target device 4320 or may be external to target device 4320, wherein target device 4320 may be configured to interface with an external display device.
[0726] The video encoder 4314 and the video decoder 4324 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or further standards.
[0727] Figure 7 is a block diagram illustrating an example of a video encoder 4400, which may be the video encoder 4314 in the system 4300 shown in Figure 6. The video encoder 4400 may be configured to perform any or all of the techniques of this disclosure. The video encoder 4400 includes multiple functional components. The techniques described in this disclosure may be shared among the various components of the video encoder 4400. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0728] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402 (which may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405, and an intra-frame prediction unit 4406), a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a buffer 4413, and an entropy coding unit 4414.
[0729] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, where at least one reference picture is the picture in which the current video block is located.
[0730] Furthermore, some components such as the motion estimation unit 4404 and the motion compensation unit 4405 can be highly integrated, but for illustrative purposes, these components are represented separately in the example of the video encoder 4400.
[0731] The segmentation unit 4401 can segment an image into one or more video blocks. The video encoder 4400 and the video decoder 4500 can support various video block sizes.
[0732] The mode selection unit 4403 can select one of several encoding / decoding modes (intra-frame encoding / decoding or inter-frame encoding / decoding), for example, based on error results, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 4403 can select an intra-frame / inter-frame joint prediction (CIIP) mode, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the block based on motion vectors (e.g., sub-pixel precision or integer pixel precision).
[0733] To perform inter-frame prediction on the current video block, motion estimation unit 4404 can generate motion information for the current video block by comparing one or more reference frames from buffer 4413 with the current video block. Motion compensation unit 4405 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 4413 other than the image associated with the current video block.
[0734] The motion estimation unit 4404 and the motion compensation unit 4405 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.
[0735] In some examples, motion estimation unit 4404 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Motion estimation unit 4404 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.
[0736] In other examples, motion estimation unit 4404 can perform bidirectional prediction on the current video block. Motion estimation unit 4404 can search for a reference video block for the current video block in the reference images of list 0, and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 4404 can then generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.
[0737] In some examples, the motion estimation unit 4404 can output a complete set of motion information for use in the decoder's decoding process. In some examples, the motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, the motion estimation unit 4404 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 4404 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0738] In one example, the motion estimation unit 4404 may indicate a value to the video decoder 4500 in the syntax structure associated with the current video block, which indicates that the current video block has the same motion information as another video block.
[0739] In another example, motion estimation unit 4404 may identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0740] As discussed above, the video encoder 4400 can transmit motion vectors via signaling in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.
[0741] Intra-prediction unit 4406 can perform intra-prediction on the current video block. When intra-prediction unit 4406 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples of other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.
[0742] The residual generation unit 4407 can generate residual data for the current video block by subtracting (or more) predicted video blocks from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0743] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 4407 may not perform subtraction operations.
[0744] The transform processing unit 4408 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0745] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0746] The inverse quantization unit 4410 and the inverse transform unit 4411 can apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 4412 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block and store it in the buffer 4413.
[0747] After the video block is reconstructed by reconstruction unit 4412, a loop filtering operation can be performed to reduce video block artifacts in the video block.
[0748] The entropy coding unit 4414 can receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, it can perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0749] Figure 8 is a block diagram illustrating an example of a video decoder 4500, which may be the video decoder 4324 in the system 4300 shown in Figure 6. The video decoder 4500 may be configured to perform any or all of the techniques disclosed herein. In the example shown, the video decoder 4500 includes multiple functional components. The techniques described in this disclosure may be shared among the various components of the video decoder 4500. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0750] In the example shown, the video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-frame prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, the video decoder 4500 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 4400.
[0751] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-decoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by executing AMVP and Merge modes.
[0752] The motion compensation unit 4502 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter to be used, with sub-pixel accuracy, can be included in the syntax element.
[0753] The motion compensation unit 4502 can use the interpolation filter used by the video encoder 4400 during the encoding of the video block to calculate the interpolation for sub-integer pixels of the reference block. The motion compensation unit 4502 can determine the interpolation filter used by the video encoder 4400 based on the received syntax information, and the motion compensation unit 4502 can use the interpolation filter to generate the prediction block.
[0754] The motion compensation unit 4502 may use some syntax information to determine the size of the blocks used to encode (multiple) frames and / or (multiple) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame codec block, and other information for decoding the encoded video sequence.
[0755] Intra-prediction unit 4503 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 4504 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 4501. Inverse transform unit 4505 applies the inverse transform.
[0756] The reconstruction unit 4506 can sum the residual block with the corresponding predicted block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 4507, which provides a reference block for subsequent motion compensation / intra-frame prediction, and also generates decoded video for presentation on a display device.
[0757] Figure 9 is a schematic diagram of an example encoder 4600. Encoder 4600 is suitable for implementing VVC techniques. Encoder 4600 includes three loop filters: a deblocking filter (DF) 4602, a sample adaptive compensation (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike DF 4602, which uses predefined filters, SAO 4604 and ALF 4606 utilize the original samples of the current image, respectively, by adding an offset and by applying a finite impulse response (FIR) filter, and by utilizing the encoded / decoded side information through signal transmission offset and filter coefficients to reduce the mean square error between the original and reconstructed samples. ALF 4606 is located in the last processing stage of each image and can be considered as a tool to attempt to capture and repair artifacts caused by previous stages.
[0758] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference images obtained from a reference image buffer 4612. Residual blocks from inter-frame or intra-frame prediction are fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are then fed into an entropy coding component 4618. The entropy coding component 4618 entropy-codes the prediction results and the quantized transform coefficients and transmits them to a video decoder (not shown). The quantized components output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. REC component 4624 is able to output images to DF 4602, SAO 4604 and ALF 4606 for filtering before these images are stored in reference image buffer 4612.
[0759] The following is a list of some preferred solutions.
[0760] The following solutions illustrate examples of the techniques discussed in this article.
[0761] 1. A method for processing media data, comprising: determining that the video is processed in the same manner as a 4:2:0 video when the displacement data is encoded and decoded into 4:2:2 video; and performing a conversion between visual media data and a bitstream based on the displacement data.
[0762] 2. The method according to Solution 1, wherein when DisplacementDim equals 1, the first displacement component is derived from the first color component of the video, and the second and third displacement components are assumed to be 0.
[0763] 3. The method according to any one of solutions 1-2, wherein when DisplacementDim equals 3, the first displacement component, the second displacement component, and the third displacement component are derived from the first color component of the video.
[0764] 4. The method according to any one of solutions 1-3, wherein the sub-block size of the AC-based displacement encoding / decoding must be constrained.
[0765] 5. The method according to any one of solutions 1-4, wherein the sub-block size of the AC-based displacement encoding / decoding must be greater than 0 or greater than 1.
[0766] 6. The method according to any one of solutions 1-5, wherein the encoding / decoding type of the AC-based displacement encoding / decoding is aligned with the encoding / decoding type of the submesh.
[0767] 7. The method according to any one of solutions 1-6, wherein when smh_type indicates intra-frame encoding / decoding according to I_SUBMESH, it is not necessary to transmit dislp_type via signaling, and dislp_type can be presumed to be intra-frame encoding / decoding according to I_DISPLACEMENT.
[0768] 8. The method according to any one of solutions 1-7, wherein when dislp_type indicates intra-frame encoding / decoding according to I_DISPLACEMENT, it is not necessary to transmit smh_type via signaling, and smh_type can be presumed to be intra-frame encoding / decoding according to I_SUBMESH.
[0769] 9. The method according to any one of solutions 1-8, wherein when smh_type indicates inter-frame encoding / decoding according to P_SUBMESH or SKIP_SUBMESH, it is not necessary to transmit dislp_type via signaling, and dislp_type can be presumed to be inter-frame encoding / decoding according to P_DISPLACEMENT.
[0770] 10. The method according to any one of solutions 1-9, wherein when dislp_type indicates inter-frame coding / decoding according to P_DISPLACEMENT, smh_type can be presumed to be intra-frame coding / decoding according to P_SUBMESH or SKIP_SUBMESH.
[0771] 11. The method according to any one of solutions 1-10, wherein when the subgrid at time t1 uses the subgrid at time t2 as a reference, the displacement data at time t1 may use only the displacement data at time t2 as a reference.
[0772] 12. The method according to any one of solutions 1-11, wherein the reference index of the displacement is set to be equal to the reference index of the subgrid.
[0773] 13. The method according to any one of solutions 1-12, wherein the displacement reference list structure is set to be the same as the base mesh reference list structure according to bmesh_ref_list_struct.
[0774] 14. The method according to any one of solutions 1-13, wherein the displacement data is encoded and decoded into 4:0:0 video and processed in the same manner as 4:2:0 video.
[0775] 15. According to any one of solutions 1-14, when the displacement data is encoded and decoded into a 4:0:0 video and DisplacementDim equals 1, the first displacement component is derived from the first color component of the video, and the second and third displacement components are presumed to be 0.
[0776] 16. According to any one of solutions 1-15, when the displacement data is encoded and decoded into a 4:0:0 video and DisplacementDim equals 3, the first displacement component, the second displacement component, and the third displacement component are derived from the first color component of the video.
[0777] 17. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of solutions 1-16.
[0778] 18. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform the method according to any one of solutions 1-16 when executed by a processor.
[0779] 19. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining that the video is processed in the same manner as 4:2:0 video when displacement data is encoded and decoded into 4:2:2 video; and generating a bitstream based on the determination.
[0780] 20. A method for storing a bitstream of video, comprising: determining, when shift data is encoded or decoded into 4:2:2 video, that the video is processed in the same manner as 4:2:0 video; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.
[0781] 21. A method, apparatus or system described in this disclosure.
[0782] In the described solution, the encoder conforms to the format rules by generating an encoded representation based on those rules. In the described solution, the decoder parses the syntax elements in the encoded representation using known information about their presence or absence, based on the format rules, to generate the decoded video.
[0783] In this disclosure, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits at co-positions or propagated at different positions in the bitstream defined by the syntax. For example, a macroblock can be encoded based on the error residual value after transformation and encoding / decoding, and also using bits from the header and other fields in the bitstream. Furthermore, during the conversion, the decoder can, based on this determination, parse the bitstream knowing whether certain fields may or may not be present, as described in the solutions above. Similarly, the encoder can determine whether to include or exclude specific syntax fields, and generate the encoding / decoding representation accordingly by including or excluding syntax fields from the encoding / decoding representation.
[0784] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more computer program instruction modules encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition affecting machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" includes all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for an associated computer program, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are artificially generated signals, such as machine-generated electrical signals, optical signals, or electromagnetic signals, which are generated to encode information to be transmitted to a suitable receiver device.
[0785] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the related program, or in multiple co-located files (e.g., a file storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at a single site or distributed across multiple sites and interconnected by a communication network.
[0786] The processing and logic flows described in this disclosure can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the apparatus can also be implemented as special-purpose logic circuitry, such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs).
[0787] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor storage devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable hard disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. Processors and memory may be supplemented by or incorporated into special-purpose logic circuitry.
[0788] While this disclosure contains numerous details, these details should not be construed as limiting any subject matter or the scope of the claims, but rather as descriptions of features specific to particular embodiments of this disclosure. Certain features described in the context of individual embodiments in this disclosure may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Furthermore, although features may function in certain combinations as described above, and even were originally claimed in this manner, in some cases one or more features in a claimed combination may be removed from that combination, and the claimed combination may be for sub-combinations or variations thereof.
[0789] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed sequentially in the particular order or sequence shown, or requiring all shown operations to be performed in order to achieve the desired result. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be construed as requiring such partitioning in all embodiments.
[0790] Only a few implementations and examples are described, and other implementations, improvements and variations may be made based on what is described and shown in this disclosure.
[0791] When there is no intermediate component (other than a line, trace, or other medium between the first and second components), the first component is directly coupled to the second component. When there is an intermediate component between the first and second components (other than a line, trace, or other medium), the first component is indirectly coupled to the second component. The term "coupled" and its variations include direct coupling and indirect coupling. The use of the term "about" means including a range of ±10% of the following figures, unless otherwise specified.
[0792] While several embodiments have been provided in this disclosure, it should be understood that the disclosed systems and methods may be embodied in many other specific forms without departing from the spirit or scope of this disclosure. The present examples are to be considered illustrative rather than restrictive and are not intended to be limited to the details set forth herein. For example, various elements or components may be combined or integrated into another system, or certain features may be omitted or not implemented.
[0793] Furthermore, the technologies, systems, subsystems, and methods described and illustrated in the various embodiments as discrete or separate may be combined or integrated with other systems, modules, technologies, or methods without departing from the scope of this disclosure. Other items shown or discussed as couplings may be directly connected or indirectly coupled or communicated through some interface, device, or intermediate component (whether electrical, mechanical, or other). Examples of other changes, substitutions, and modifications can be identified by those skilled in the art and may be made without departing from the spirit and scope of this disclosure.
Claims
1. A method for processing media data, comprising: When the displacement data is encoded and decoded in a 4:4:4 video format, it is determined that only the luminance channel is used to transmit the displacement data. And perform the conversion between visual media data and bitstream based on the displacement data.
2. The method of claim 1, wherein the displacement data includes three-dimensional (3D) displacement data encoded and decoded into video, and wherein all of the 3D displacement data is in the luminance channel.
3. The method of claim 1, wherein the displacement data is packaged into the video, regardless of the color format.
4. The method of claim 3, wherein when the displacement data is packaged into the video, regardless of the color format, only the luminance channel is used to transmit the displacement data.
5. The method of claim 3, wherein regardless of the color format, when the displacement dimension (DisplacementDim) is equal to 1, the first displacement component is derived from the first color component of the video, and the second and third color components are presumed to be 0.
6. The method of claim 3, wherein regardless of the color format, when the displacement dimension (DisplacementDim) is equal to 3, the first displacement component, the second displacement component, and the third displacement component are derived from the first color component of the video.
7. The method of claim 1, wherein when the base grid corresponding to the second time (t2) is in the reference list of the base grid corresponding to the first time (t1), the displacement data at the first time is only allowed to use the displacement data at the second time as a reference.
8. The method of claim 7, wherein when the base grid at the second time is a reference to the base grid at the first time, the displacement data at the first time is only allowed to use the displacement data at the second time as a reference.
9. The method of claim 1, wherein when the displacement at the second time (t2) is in a reference list of the displacement at the first time (t1), the base grid at the first time is only allowed to use the base grid at the second time.
10. The method of claim 9, wherein when the base grid at the second time is a reference to the base grid at the first time, the displacement data at the first time is only allowed to use the displacement data at the second time as a reference.
11. The method of claim 1, wherein at least one of the displacement reference list and the base mesh reference list is configured to correspond to the atlas reference list.
12. The method of claim 11, wherein when the atlas corresponding to the second time (t2) is in the reference list of the atlas at the first time (t1), at least one of the base grid at the first time and the displacement data at the first time is only allowed to use at least one of the base grid at the second time and the displacement data at the second time as a reference.
13. The method of claim 11, wherein when the atlas at the second time (t2) is a reference to the atlas at the first time (t1), the base grid at the first time and the displacement data at the first time are only permitted to use the base grid at the second time and the displacement data at the second time.
14. The method of claim 1, further comprising using a synchronization method for all sub-bitstreams corresponding to a video codec standard, wherein the video codec standard includes one of video-based point cloud compression (V-PCC) and video-based dynamic mesh codec (V-DMC).
15. The method of claim 14, wherein all said sub-bit streams have the same reference list.
16. The method according to any one of claims 14-15, wherein all said sub-bit streams have the same reference structure.
17. The method according to any one of claims 14-16, wherein one or more syntax elements are used to indicate the maximum allowed number of decoding buffers for the decoding process and each sub-bitstream, and wherein the number of decoding buffers is not greater than the maximum allowed number of decoding buffers for the decoding process.
18. The method of any one of claims 14-17, wherein one or more syntax elements are used to indicate the maximum allowed number of reordered frames for the decoding process and each sub-bitstream, and wherein the number of reordered frames is not greater than the maximum allowed number of reordered frames for the decoding process.
19. The method of claim 18, wherein one or more syntax elements are used to indicate the maximum allowed number of atlas frames having the atlas frame output flag equal to 1 that precede any atlas frame having the atlas frame output flag equal to 1 in the output order at a particular time-domain layer.
20. The method of any one of claims 14-19, wherein one or more syntax elements are used to indicate the maximum allowed number of delayed frames for the decoding process and each sub-bitstream, and wherein the number of delayed frames is not greater than the maximum allowed number of delayed frames for the decoding process.
21. The method of claim 20, wherein one or more syntax elements are used to indicate, at a particular time-domain layer, the maximum allowed number of atlas frames having the atlas frame output flag (AtlasFrameOutputFlag) equal to 1, preceding any atlas frame having the atlas frame output flag equal to 1 and following any atlas frame having the atlas frame output flag equal to 1, in output order.
22. The method according to any one of claims 1-21, wherein all sub-bit streams at a specific time have the same time-domain identifier.
23. The method according to any one of claims 1-22, wherein the conversion comprises encoding the media data into the bitstream.
24. The method according to any one of claims 1-22, wherein the conversion comprises decoding the media data from the bitstream.
25. An apparatus for processing video data, comprising: processor; And a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-24.
26. A non-transitory computer-readable medium comprising a computer program product for use by a video codec apparatus, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, causing the video codec apparatus to perform the method according to any one of claims 1-24 when executed by a processor.
27. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: When the displacement data is encoded and decoded in a 4:4:4 video format, it is determined that only the luminance channel is used to transmit the displacement data. And the bit stream is generated based on the displacement data.
28. A method for storing a bitstream of video, comprising: When the displacement data is encoded and decoded in a 4:4:4 video format, it is determined that only the luminance channel is used to transmit the displacement data. The bit stream is generated based on the displacement data; And storing the bit stream in a non-transitory computer-readable recording medium.
29. A method, apparatus or system described in this disclosure.