Data arrangement for dynamic grid coding and decoding

By arranging the displacement vector and attribute diagrams of three-dimensional visual media data into a single picture and using a two-dimensional codec to process it, the problem of multi-bit stream processing in dynamic grid codec is solved, and the system throughput and efficiency is improved.

CN120513629APending Publication Date: 2025-08-19DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480006680.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-03
Filing Date
2024-01-03
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing dynamic grid codec technology requires processing multiple video codec bitstreams, resulting in the inability to deploy in systems that can only process one video bitstream, and may reduce system throughput and efficiency.

Method used

The displacement vector and attribute diagram of the three-dimensional visual media data are arranged into a single picture, processed by a two-dimensional codec, and the data arrangement is optimized through splicing, position information indication, independent decoding, bit depth and color format alignment, etc., so as to facilitate processing by a 2D codec.

Benefits of technology

It improves the throughput and efficiency of the dynamic grid codec system, and can efficiently process dynamic grid data in traditional 2D codec systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120513629A_ABST
    Figure CN120513629A_ABST
Patent Text Reader

Abstract

A mechanism for processing video data is disclosed. Displacement vectors and attribute maps of three-dimensional (3D) visual media data are determined to be arranged into a single picture for processing by a two-dimensional (2D) codec. A conversion between the visual media data and the bitstream is performed based on the displacement vector and the attribute graph.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This patent application claims the benefit of U.S. Provisional Patent Application No. 63 / 478,314, filed on January 3, 2023, which is incorporated herein by reference. Technical Field

[0003] The present disclosure relates to the processing of digital images and videos. Background Art

[0004] Digital video accounts for the largest share of bandwidth used on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, the demand for bandwidth used by digital video is likely to continue to grow. Summary of the Invention

[0005] A first aspect relates to a method for processing video or image data, comprising: determining a displacement vector and an attribute map for arranging three-dimensional (3D) visual media data into a single picture for processing by a two-dimensional (2D) codec; and performing conversion between the visual media data and a bitstream based on the displacement vector and the attribute map.

[0006] Optionally, in any one of the above aspects, another embodiment of this aspect provides concatenating the displacement vector and the property map to form the single picture.

[0007] Optionally, in any one of the above aspects, another embodiment of this aspect provides determining to arrange the displacement vector and the property map in different parts of the single picture, and indicating position information or size information of each of the different parts in the bitstream.

[0008] Optionally, in any one of the above aspects, another embodiment of this aspect provides indicating the position information or the size information in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a picture header or a slice header of the bitstream.

[0009] Optionally, in any one of the above aspects, another embodiment of this aspect provides for deriving the position information or the size information from information in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a picture header or a slice header of the bitstream.

[0010] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange the displacement vector and the property map in different parts of the single picture, wherein each of the different independent parts can be decoded independently.

[0011] Optionally, in any one of the above aspects, another embodiment of this aspect provides determining to arrange the displacement vector and the property map in different strips, different slices or different sub-pictures of the single picture.

[0012] Optionally, in any of the above aspects, another embodiment of this aspect provides using constrained intra prediction and / or motion constrained slice sets so that the displacement vector and the property map can be decoded independently.

[0013] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning the displacement vector and the property map to a single bit depth before determining to arrange the displacement vector and the property map into the single picture.

[0014] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning the bit depth of the displacement vector to the property map, so that the displacement vector and the property map are aligned to the single bit depth.

[0015] Optionally, in any of the above aspects, another implementation of this aspect provides aligning the bit depth of the property map to the displacement vector, so that the displacement vector and the property map are aligned to the single bit depth.

[0016] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning a portion of data with a lower bit depth to a portion of data with a higher bit depth, so that the displacement vector and the property map are aligned to the single bit depth.

[0017] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning the displacement vectors and the property map to a single color format before determining to arrange the displacement vectors and the property map into the single picture.

[0018] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning the color format of the displacement vector to the property map, so that the displacement vector and the property map are aligned to the single color format.

[0019] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning the color format of the property map to the displacement vector, so that the displacement vector and the property map are aligned to the single color format.

[0020] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning a portion of data with less color format information to a portion of data with more color format information, so that the displacement vector and the attribute map are aligned to the single color format.

[0021] Optionally, in any one of the above aspects, another embodiment of the aspect provides determining the motion field of the grid with the displacement vector and the attribute Figure 1 The 2D images are arranged together into the single picture for processing by the 2D codec.

[0022] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange any two parts or all parts of the motion field of the grid, the displacement vector and the property map into the single picture for processing by the 2D codec.

[0023] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning any two parts or all parts of the motion field, the displacement vector and the attribute map of the grid to a single bit depth before determining to arrange the motion field, the displacement vector and the attribute map of the grid into the single picture.

[0024] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning any two parts or all parts of the motion field, the displacement vector and the attribute map of the grid to a single color format before determining to arrange the motion field, the displacement vector and the attribute map of the grid into the single picture.

[0025] Optionally, in any of the above aspects, another embodiment of this aspect provides that the single color format includes the motion field of the grid, the displacement vectors and two or more parts of the property map with the most color information.

[0026] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange any data that can be processed by the 2D codec into the single picture for processing by the 2D codec.

[0027] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange two or more parts of the occupancy map, geometry map and property map of the video-based point cloud codec into the single picture for processing by the 2D codec.

[0028] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange the occupancy map and geometry map conforming to the Moving Picture Experts Group (MPEG) immersive video standard into the single picture for processing by the 2D codec.

[0029] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning any data arranged into the single picture to the same bit depth.

[0030] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange only data with the same bit depth into the single picture.

[0031] Optionally, in any of the above aspects, another embodiment of this aspect provides that any data with different bitmaps are not arranged into the single picture.

[0032] Optionally, in any of the above aspects, another embodiment of this aspect provides aligning any data arranged in the single picture to have the same color format.

[0033] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange only data having the same color format into the single picture.

[0034] Optionally, in any of the above aspects, another embodiment of this aspect provides not arranging any data with different bit depths into the single picture.

[0035] Optionally, in any of the above aspects, another embodiment of this aspect provides determining to arrange only data with the same bit depth and the same color format into the single picture.

[0036] Optionally, in any of the above aspects, another embodiment of this aspect provides that any data with different bit depths or different color formats are not arranged into the single picture.

[0037] Optionally, in any of the above aspects, another embodiment of the aspect provides determining any combination of a motion field of a grid arranging 3D visual media data, the displacement vectors and the property map.

[0038] Optionally, in any of the above aspects, another embodiment of this aspect provides that the conversion includes encoding the media data into a bitstream.

[0039] Optionally, in any of the above aspects, another embodiment of this aspect provides that the converting includes decoding the media data from a bitstream.

[0040] A second aspect includes an apparatus for processing media data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any disclosed embodiment.

[0041] A third aspect relates to a non-transitory computer-readable medium, comprising a computer program product for use with a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any disclosed embodiment.

[0042] A fourth aspect relates to a non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method includes the method of any disclosed embodiment.

[0043] A fifth aspect relates to a method for storing a bitstream of a video, including the method of any disclosed embodiment.

[0044] The sixth aspect relates to the method, device or system described in the present disclosure.

[0045] For purposes of clarity, any of the foregoing embodiments may be combined with any one or more of the other foregoing embodiments to create new embodiments within the scope of the present disclosure.

[0046] These and other features will be more clearly understood from the following detailed description, taken in conjunction with the accompanying drawings and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] For a more complete understanding of this disclosure, reference is now made to the following brief description, taken in conjunction with the accompanying drawings and detailed description, wherein like reference numerals represent like parts.

[0048] Figure 1 is a diagram showing an example decoder design for dynamic trellis encoding and decoding.

[0049] Figure 2 is a schematic diagram showing an example structure for a dynamic grid codec test model.

[0050] Figure 3is a block diagram illustrating an example video processing system.

[0051] Figure 4 is a block diagram of an example video processing device.

[0052] Figure 5 is a flow chart of an example method for video processing.

[0053] Figure 6 is a block diagram illustrating an example video encoding and decoding system.

[0054] Figure 7 is a block diagram illustrating an example encoder.

[0055] Figure 8 is a block diagram illustrating an example decoder.

[0056] Figure 9 is a schematic diagram of an example encoder. DETAILED DESCRIPTION

[0057] It should be understood at the outset that although exemplary implementations of one or more embodiments are provided below, the disclosed systems and / or methods may be implemented using any number of techniques, whether currently known or yet to be developed. The present disclosure should in no way be limited to the exemplary implementations, drawings, and techniques illustrated below, including the exemplary designs and implementations shown and described herein, but may be modified within the scope of the appended claims and their full scope of equivalents.

[0058] 1. Preliminary Discussion

[0059] The present disclosure relates to immersive video coding and decoding technology. Specifically, the technology of the present disclosure relates to video-based dynamic grid coding and decoding that complies with the Moving Picture Experts Group (MPEG)-I standard. The technology of the present disclosure can also be applied to other immersive video coding and decoding standards or codecs.

[0060] 2. Further Discussion

[0061] In computer graphics, three-dimensional (3D) / immersive content is typically represented by a 3D mesh and texture map. This mesh and texture data can be machine-generated or converted from images captured by multiple cameras at different angles. Similar to two-dimensional (2D) video, when this 3D content changes over time, the mesh and texture data also change, forming a dynamic mesh sequence. The amount of dynamic mesh data is often large, making it difficult to store and transmit. To meet the needs of applications using dynamic meshes, MPEG issued a request for proposals. To effectively use 2D codecs, one of the requirements is to use 2D video codec standards to compress most of the data while keeping the rest simple and low-complexity. This requirement ensures that this representation can take advantage of the strengths of 2D video hardware / software systems without the need to expend significant effort redesigning a specific system just for dynamic meshes.

[0062] MPEG received five responses to the Request for Proposals. A test model was constructed for the planned development of a dynamic trellis codec standard.

[0063] The latest test model for dynamic mesh coding as of the time of writing this disclosure can be found at the following link: http: / / mpegx.int-evry.fr / software / MPEG / dmc / mpeg-vmesh-tm / - / tags / v2.0; and the latest working draft document is Working Draft (WD) 1.0.

[0064] 2.1 Data Representation in Dynamic Mesh Encoding and Decoding

[0065] Figure 1 is a diagram showing an example decoder design for dynamic trellis encoding and decoding. Figure 1 An example decoder design is shown. As can be seen, the dynamic mesh decoder receives 3 bit streams and performs decoding to reconstruct the dynamic mesh plus texture signal. The first bit stream represents the base mesh, which is a decimated version of the original mesh. The second bit stream represents the displacement vector between the reconstructed base mesh and the original mesh. The displacement vector is arranged as a 2D video and compressed using a codec that complies with the 2D video codec standard. The third bit stream represents the texture (or attribute map). The attribute map is also arranged as a 2D video and compressed using a codec that complies with the 2D video codec standard. The design concept is to make the base mesh part small enough so that the module that processes the base mesh can be simply implemented. On the other hand, the displacement vector and attribute map account for most of the volume of the entire dynamic mesh data and can be processed by a dedicated efficient 2D video codec system. Such a design can reduce the additional effort to implement the dynamic mesh codec system and ensure high throughput and codec efficiency of dynamic mesh data.

[0066] 2.2 Test Model of Dynamic Mesh Encoding and Decoding

[0067] Figure 2 The following is a schematic diagram showing an example structure for a dynamic mesh codec test model. In this model, Draco is used to compress the base mesh, and the High Efficiency Video Codec (HEVC) Test Model (HM) is used to compress the displacement vectors and attribute maps. However, it should be noted that other meshes or video codec systems can also be used for dynamic mesh codec.

[0068] The base mesh m is generated from the original mesh using a downsampling scheme. Its quantized version m' is then encoded and decoded using Draco. The reconstructed base mesh m' is obtained by dequantizing m'. The displacement vector d' is generated by taking the difference between the original mesh and the subdivided version of m' obtained using the subdivision scheme.

[0069] 2.3 Displacement Vector Encoding and Decoding

[0070] After obtaining the displacement vector d' (i.e., the difference between the original mesh and the subdivided base mesh), a lifting-based wavelet transform is applied to further compact the energy. The wavelet transform coefficients are then traversed from low frequency to high frequency using the Morton order to form 2D coefficient blocks. Each 2D coefficient block comprises the picture to be processed by the 2D codec.

[0071] 2.4 Motion Field Codec

[0072] In the example test model, the motion field between the base grids is encoded and decoded directly using arithmetic codec. Example implementations investigated the encoding and decoding of the motion field, which was also encoded and decoded using a standard-compliant 2D codec system, and showed that the loss in codec efficiency was minimal. Therefore, it may be meaningful to further transfer the motion field encoding and decoding process to a 2D video codec.

[0073] 3. Technical problems solved by the disclosed technical solutions

[0074] In the example design of dynamic grid codec, the need to process multiple video codec bitstreams may make the solution unsuitable for deployment in video encoding / decoding systems that can only process a single video bitstream. Moreover, even if the system can process multiple video bitstreams, this issue may result in reduced throughput and / or efficiency of the entire system.

[0075] 4. List of solutions and implementation examples

[0076] Each aspect of the detailed description below should be considered as an example to explain the general concept. These examples should not be interpreted in a narrow sense. In addition, these examples can be combined in any way. Combinations between this disclosure and other disclosures are also applicable.

[0077] 1. Displacement vectors and attribute maps can be arranged in one picture to be processed by 2D image / video codecs.

[0078] a. In one example, the displacement vectors and the attribute map can be concatenated to form a picture.

[0079] b. In one example, the displacement vectors and the property map may be arranged in different parts of a picture, where the position and / or size information of each part is indicated in the video bitstream.

[0080] i. In one example, the location and / or size information of each part may be signaled in a sequence parameter set (SPS) / picture parameter set (PPS) / video parameter set (VPS) / picture header / slice header.

[0081] ii. In one example, the position and / or size information of each part can be derived from information in the SPS / PPS / VPS / picture header / slice header.

[0082] c. In one example, the displacement vectors and the attribute map can be arranged in different parts of a picture, where each part can be decoded independently.

[0083] i. In one example, the displacement vectors and attribute maps may be arranged in different slices / slices / sub-pictures of a picture.

[0084] ii. In one example, constrained intra prediction and / or motion constrained slice sets may be used to enable displacement vectors and property maps to be independently decoded.

[0085] d. In one example, the displacement vectors and attribute maps can be aligned to one bit depth and then arranged in one picture.

[0086] i. In one example, the bit depth of the displacement vector can be aligned to the property map.

[0087] ii. In one example, the bit depth of the property map may be aligned to the displacement vector.

[0088] iii. In one example, a portion of data having a lower bit depth may be aligned to another portion of data having a higher bit depth.

[0089] e. In one example, the displacement vectors and attribute maps can be aligned to a color format and then arranged in an image.

[0090] i. In one example, the color format of the displacement vector can be aligned to the property map.

[0091] ii. In one example, the color format of the property map may be aligned to the displacement vector.

[0092] iii. In one example, a portion of data with less color information may be aligned to another portion of data with more color information.

[0093] 2. The motion field and displacement vectors of the mesh can be arranged in one picture to be processed by a 2D image / video codec.

[0094] a. In one example, the motion field and displacement vectors of a mesh can be arranged in one picture, and the above items / sub-items can be applied.

[0095] 3. The motion field and property map of the mesh can be arranged in one picture to be processed by a 2D image / video codec.

[0096] a. In one example, the motion field and property map of a grid can be arranged in one picture, and the above items / sub-items can be applied.

[0097] 4. The motion field, displacement vectors and property map of the mesh can be arranged in one picture to be processed by a 2D video image / codec.

[0098] a. In one example, the motion field and property map of a grid can be arranged in one picture, and the above items / sub-items can be applied.

[0099] b. In one example, any two or all parts of the motion field, displacement vectors, and property map of a mesh can be arranged in one picture.

[0100] c. In one example, any two or all parts of the motion field, displacement vectors, and attribute map of a mesh can be aligned to one bit depth and then arranged in one picture.

[0101] i. In one example, the bit depth is the maximum bit depth of any two or all parts.

[0102] d. In one example, any two or all parts of the motion field, displacement vectors, and property map of a mesh can be aligned into a color format and then arranged into a picture.

[0103] i. In one example, the color format is the one that has the most color information for any two or all of the parts.

[0104] 5. In the codec system, multiple or all data that can be processed by a 2D image / video codec can be arranged in one picture.

[0105] a. In one example, any two or all of the occupancy map, geometry, and attributes in video-based point cloud codec can be arranged in one picture, and the above items / sub-items can be applied.

[0106] b. In one example, geometric shapes and attributes in the MPEG immersive video standard may be arranged in a picture, and the above items / sub-items may be applied.

[0107] c. In one example, data arranged in one picture may be aligned to have the same bit depth.

[0108] d. In one example, only data with the same bit depth can be arranged in one picture.

[0109] i. Alternatively, data with different bit depths may not be arranged in one picture.

[0110] e. In one example, data arranged in one picture may be aligned to have the same color format.

[0111] f. In one example, only data with the same color format can be arranged in one picture.

[0112] i. Alternatively, data with different bit depths may not be arranged in one picture.

[0113] g. In one example, only data with the same bit depth and the same color format can be arranged in one picture.

[0114] i. Alternatively, data with different bit depths or different color formats may not be arranged in one picture.

[0115] 5. References

[0116] [1] MPEG Technical Requirements, “CfP for Dynamic Mesh Coding”, ISO / IEC JTC 1 / SC 29 / WG2 Document No. N145, October 2021.

[0117] [2] K. Mammou, J. Kim, A. Tourapis, and D. Podborski, “[V-CG] Apple’s Dynamic MeshCoding CfP Response,” ISO / IEC JTC 1 / SC 29 / WG 7 Document No. m59281, April 2022.

[0118] [3]MPEG output document, “WD 1.0 of V-DMC”, ISO / IEC JTC 1 / SC 29 / WG 7 Document No. N0486, November 2022.

[0119] [4] C. Huang, X. Xu, X. Zhang, J. Tian, and S. Liu, “Investigation of video coding of motion fields,” ISO / IEC JTC 1 / SC 29 / WG 7 Document No. m61005, July 2022.

[0120] Figure 3 The block diagram of FIG4000 shows an example video processing system 4000 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 4000. System 4000 may include an input 4002 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8 or 10-bit multi-component pixel values, or may be received in a compressed or encoded format. Input 4002 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces (such as Ethernet, passive optical networks (PONs), etc.) and wireless interfaces (such as wireless fidelity (Wi-Fi) or cellular interfaces).

[0121] System 4000 may include a codec component 4004 that can implement the various codecs or encoding methods described in this disclosure. Codec component 4004 can reduce the average bit rate of the video from input 4002 to the output of codec component 4004 to generate a codec representation of the video. Codec technology is therefore sometimes referred to as video compression or video transcoding technology. The output of codec component 4004 can be stored or transmitted via a communication connection such as represented by component 4006. The bitstream (or codec) representation of the video stored or communicated received at input 4002 can be used by component 4008 to generate pixel values or playable video sent to display interface 4010. The process of generating user-visible video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tool or operation is used by the encoder, and the corresponding decoding tool or operation of the inverse codec result will be performed by the decoder.

[0122] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort interfaces, etc. Examples of storage interfaces include Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect (PCI), Integrated Drive Electronics (IDE) interfaces, etc. The technology described in this disclosure may be embodied in various electronic devices, such as mobile phones, laptop computers, smartphones, or other devices capable of performing digital data processing and / or video display.

[0123] Figure 4 4 is a block diagram of an example video processing device 4100. Device 4100 can be used to implement one or more methods described herein. Device 4100 can be embodied in a smartphone, a tablet computer, a computer, an Internet of Things (IoT) receiver, etc. Device 4100 may include one or more processors 4102, one or more memories 4104, and video processing circuitry 4106. Processor(s) 4102 can be configured to implement one or more methods described in this disclosure. Memory(s) 4104 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 4106 can be used to implement some of the techniques described in this disclosure in hardware circuitry. In some embodiments, video processing circuitry 4106 can be at least partially included in processor 4102, such as a graphics coprocessor.

[0124] Figure 5 4 is a flow chart of an example method 4200 for video processing. At step 402, a determination is made as to whether displacement vectors and a property map of three-dimensional (3D) visual media data should be arranged into a single picture for processing by a two-dimensional (2D) codec. At step 4204, conversion between the visual media data and a bitstream is performed based on the displacement vectors and the property map. According to an example, the conversion at step 4204 may include encoding at an encoder or decoding at a decoder.

[0125] It should be noted that method 4200 can be implemented in an apparatus for processing video data, such as video encoder 4400, video decoder 4500, and / or encoder 4600, that includes a processor and non-transitory memory with instructions stored thereon. In this case, the instructions, when executed by the processor, cause the processor to perform method 4200. Furthermore, method 4200 can be performed by a non-transitory computer-readable medium comprising a computer program product for use with a video codec device. This computer program product includes computer-executable instructions stored on a non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs method 4200. Furthermore, a non-transitory computer-readable recording medium may store a video bitstream generated by the video processing apparatus performing method 4200. Furthermore, method 4200 can be performed by an apparatus for processing video data, including a processor and non-transitory memory with instructions stored thereon. The instructions, when executed by the processor, cause the processor to perform method 4200.

[0126] Figure 6 4 is a block diagram illustrating an example video codec system 4300 that can utilize the techniques of this disclosure. Video codec system 4300 can include a source device 4310 and a destination device 4320. Source device 4310 generates encoded video data, where source device 4310 can be referred to as a video encoding device. Destination device 4320 can decode the encoded video data generated by source device 4310, where destination device 4320 can be referred to as a video decoding device.

[0127] Source device 4310 may include a video source 4312, a video encoder 4314, and an input / output (I / O) interface 4316. Video source 4312 may include a source such as a video capture device, an interface for receiving video data from a video content provider, and / or a computer graphics system for generating video data, or a combination of these sources. The video data may include one or more pictures. Video encoder 4314 encodes the video data from video source 4312 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are a codec representation of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 4316 may include a modulator / demodulator (modem) and / or a transmitter. The coded video data may be transmitted directly to destination device 4320 via network 4330 via I / O interface 4316. The coded video data may also be stored on storage medium / server 4340 for access by destination device 4320.

[0128] Destination device 4320 may include an I / O interface 4326, a video decoder 4324, and a display device 4322. The I / O interface 4326 may include a receiver and / or a modem. The I / O interface 4326 may obtain encoded video data from source device 4310 or storage medium / server 4340. The video decoder 4324 may decode the encoded video data. The display device 4322 may display the decoded video data to a user. The display device 4322 may be integrated with the destination device 4320, or may be external to the destination device 4320, wherein the destination device 4320 may be configured to be connected to an external display device interface.

[0129] The video encoder 4314 and the video decoder 4324 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or further standards.

[0130] Figure 7 is a block diagram illustrating an example of a video encoder 4400, which may be Figure 6 Video encoder 4314 in system 4300 is shown. Video encoder 4400 can be configured to perform any or all of the techniques of this disclosure. Video encoder 4400 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 4400. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0131] The functional components of the video encoder 4400 may include a segmentation unit 4401, a prediction unit 4402, a residual generation unit 4407, a transform processing unit 4408, a quantization unit 4409, an inverse quantization unit 4410, an inverse transform unit 4411, a reconstruction unit 4412, a cache 4413 and an entropy coding unit 4414. The prediction unit 4402 may include a mode selection unit 4403, a motion estimation unit 4404, a motion compensation unit 4405 and an intra-frame prediction unit 4406.

[0132] In other examples, the video encoder 4400 may include more, fewer, or different functional components. In one example, the prediction unit 4402 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in accordance with an IBC mode, wherein at least one reference picture is a picture in which the current video block is located.

[0133] Furthermore, some components, such as the motion estimation unit 4404 and the motion compensation unit 4405 , may be highly integrated, but are represented separately in the example of the video encoder 4400 for purposes of explanation.

[0134] The segmentation unit 4401 may segment a picture into one or more video blocks. The video encoder 4400 and the video decoder 4500 may support various video block sizes.

[0135] The mode selection unit 4403 can, for example, select one of a plurality of codec modes (intra-frame codec or inter-frame codec) based on the error result, and provide the generated intra-frame codec block or inter-frame codec block to the residual generation unit 4407 to generate residual block data, and to the reconstruction unit 4412 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 4403 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 4403 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).

[0136] To perform inter-frame prediction on the current video block, the motion estimation unit 4404 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the buffer 4413. The motion compensation unit 4405 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 4413 other than the picture associated with the current video block.

[0137] The motion estimation unit 4404 and the motion compensation unit 4405 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0138] In some examples, motion estimation unit 4404 may perform unidirectional prediction on the current video block, and motion estimation unit 4404 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 4404 may then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 4404 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0139] In other examples, motion estimation unit 4404 may perform bidirectional prediction on the current video block. Motion estimation unit 4404 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. Motion estimation unit 4404 may then generate reference indexes indicating the reference pictures in list 0 and list 1 that contain the reference video blocks, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 4404 may output the reference indexes and the motion vector of the current video block as motion information for the current video block. Motion compensation unit 4405 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0140] In some examples, motion estimation unit 4404 can output a complete set of motion information for use in the decoding process of a decoder. In some examples, motion estimation unit 4404 may not output a complete set of motion information for the current video. Instead, motion estimation unit 4404 can reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 4404 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.

[0141] In one example, the motion estimation unit 4404 may indicate to the video decoder 4500 a value in a syntax structure associated with the current video block that indicates that the current video block has the same motion information as another video block.

[0142] In another example, the motion estimation unit 4404 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 4500 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0143] As discussed above, the video encoder 4400 can signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that can be implemented by the video encoder 4400 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.

[0144] Intra-frame prediction unit 4406 can perform intra-frame prediction on the current video block. When intra-frame prediction unit 4406 performs intra-frame prediction on the current video block, intra-frame prediction unit 4406 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0145] The residual generation unit 4407 can generate residual data for the current video block by subtracting the predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0146] In other examples, such as in skip mode, there may be no residual data for the current video block and the residual generation unit 4407 may not perform a subtraction operation.

[0147] Transform processing unit 4408 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to a residual video block associated with the current video block.

[0148] After the transform processing unit 4408 generates a transform coefficient video block associated with the current video block, the quantization unit 4409 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0149] The inverse quantization unit 4410 and the inverse transform unit 4411 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 4412 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 4402 to generate a reconstructed video block associated with the current block for storage in the buffer 4413.

[0150] After the reconstruction unit 4412 reconstructs the video block, a loop filtering operation may be performed to reduce video blocking artifacts in the video block.

[0151] The entropy coding unit 4414 may receive data from other functional components of the video encoder 4400. When the entropy coding unit 4414 receives data, the entropy coding unit 4414 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0152] Figure 8 is a block diagram illustrating an example of a video decoder 4500, which may be Figure 6 Video decoder 4324 in system 4300 is shown. Video decoder 4500 can be configured to perform any or all of the techniques of this disclosure. In the example shown, video decoder 4500 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video decoder 4500. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0153] In the example shown, video decoder 4500 includes an entropy decoding unit 4501, a motion compensation unit 4502, an intra-prediction unit 4503, an inverse quantization unit 4504, an inverse transform unit 4505, a reconstruction unit 4506, and a buffer 4507. In some examples, video decoder 4500 may perform a decoding process that is generally opposite to the encoding process described with respect to video encoder 4400.

[0154] The entropy decoding unit 4501 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 4501 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 4502 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 4502 can determine this information, for example, by performing AMVP and Merge modes.

[0155] The motion compensation unit 4502 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision may be included in the syntax element.

[0156] The motion compensation unit 4502 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by the video encoder 4400 during encoding of the video block. The motion compensation unit 4502 may determine the interpolation filters used by the video encoder 4400 based on received syntax information, and the motion compensation unit 4502 may use the interpolation filters to generate a prediction block.

[0157] The motion compensation unit 4502 can use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame codec block, and other information used to decode the encoded video sequence.

[0158] The intra prediction unit 4503 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 4504 inversely quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 4501. The inverse transform unit 4505 applies an inverse transform.

[0159] The reconstruction unit 4506 can add the residual block to the corresponding prediction block generated by the motion compensation unit 4502 or the intra-frame prediction unit 4503 to form a decoded block. If necessary, a deblocking filter can also be used to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 4507, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.

[0160] Figure 9 is a schematic diagram of an example encoder 4600. The encoder 4600 is suitable for implementing techniques for VVC. The encoder 4600 includes three loop filters, namely a deblocking filter (DF) 4602, a sample adaptive offset (SAO) 4604, and an adaptive loop filter (ALF) 4606. Unlike the DF 4602, which uses a predefined filter, the SAO 4604 and the ALF 4606 use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, using codec side information to signal the offset and filter coefficients. The ALF 4606 is located at the last processing stage for each picture and can be seen as a tool that attempts to capture and repair artifacts caused by previous stages.

[0161] The encoder 4600 also includes an intra-frame prediction component 4608 and a motion estimation / compensation (ME / MC) component 4610 configured to receive input video. The intra-frame prediction component 4608 is configured to perform intra-frame prediction, while the ME / MC component 4610 is configured to perform inter-frame prediction using reference pictures obtained from a reference picture cache 4612. The residual block from the inter-frame prediction or intra-frame prediction is fed into a transform (T) component 4614 and a quantization (Q) component 4616 to generate quantized residual transform coefficients, which are fed into an entropy codec component 4618. The entropy codec component 4618 entropy codes the prediction results and the quantized transform coefficients and transmits them to a video decoder (not shown). The quantization component output from the quantization component 4616 can be fed into an inverse quantization (IQ) component 4620, an inverse transform component 4622, and a reconstruction (REC) component 4624. The REC component 4624 can output images to the DF 4602 , SAO 4604 , and ALF 4606 for filtering before these images are stored in the reference picture cache 4612 .

[0162] A list of some example preferred solutions is provided next.

[0163] The following solutions illustrate examples of the techniques discussed herein.

[0164] 1. A method for processing video or image data, comprising: determining a displacement vector and an attribute map for arranging three-dimensional (3D) visual media data into a single picture for processing by a two-dimensional (2D) codec; and performing conversion between the visual media data and a bitstream based on the displacement vector and the attribute map.

[0165] 2. The method according to solution 1, wherein the displacement vector and the attribute map are concatenated to form a picture.

[0166] 3. A method according to any one of solutions 1-2, wherein the displacement vector and the property map are arranged in different parts of a picture, and wherein the position or size information of each part is indicated in the bitstream.

[0167] 4. A method according to any one of solutions 1-3, wherein the position or size information of each part is transmitted or derived through a signal based on signaling in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a picture header or a slice header.

[0168] 5. The method according to any one of solutions 1-4, wherein the displacement vector and the property map are arranged in different parts of a picture, and wherein each part can be decoded independently.

[0169] 6. The method according to any one of solutions 1-5, wherein the displacement vector and the property map are arranged in different strips, slices or sub-pictures in one picture.

[0170] 7. The method according to any of solutions 1-6, wherein the displacement vector and the property map are encoded and decoded using constrained intra prediction or motion constrained slice sets so that each part can be decoded independently.

[0171] 8. The method according to any one of solutions 1-7, wherein the displacement vector and the property map are aligned to one bit depth and arranged in one picture.

[0172] 9. A method according to any of solutions 1-8, wherein the displacement vector bit depth is aligned to the attribute map bit depth, the attribute map bit depth is aligned to the displacement vector bit depth, or a lower bit depth is aligned to a higher bit depth.

[0173] 10. The method according to any one of solutions 1-9, wherein the displacement vector and the attribute map are aligned to a color format and arranged in a picture.

[0174] 11. A method according to any one of solutions 1-10, wherein a displacement vector color format is aligned to an attribute map color format, the attribute map color format is aligned to the displacement vector color format, or a lower information color format is aligned to a higher information color format.

[0175] 12. The method according to any of the solutions 1-11, wherein the motion field and displacement vectors of the mesh are arranged into one picture to be processed by the 2D codec.

[0176] 13. The method according to any of the solutions 1-12, wherein the motion field of the mesh and the property map are arranged into one picture to be processed by the 2D codec.

[0177] 14. The method according to any of solutions 1-13, wherein the motion field of the mesh, the property map and the displacement vectors are arranged into one picture to be processed by the 2D codec.

[0178] 15. A method according to any of solutions 1-14, wherein two or more of the motion field of the mesh, the property map and the displacement vector are aligned to a bit depth or aligned to a color format.

[0179] 16. The method according to any of solutions 1-15, wherein two or more of occupancy map, geometry and attributes are arranged into one picture to be processed by the 2D codec.

[0180] 17. The method of any of solutions 1-16, wherein two or more of the occupancy map, the geometry, and the attributes are aligned to a bit depth or to a color format.

[0181] 18. An apparatus for processing video data, comprising: a processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of solutions 1-17.

[0182] 19. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, so that when executed by a processor, the video codec device performs the method of any one of solutions 1-17.

[0183] 20. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method includes: determining to arrange displacement vectors and attribute maps of three-dimensional (3D) visual media data into a single picture for processing by a two-dimensional (2D) codec; and generating a bitstream based on the determination.

[0184] 21. A method for storing a bitstream of a video, comprising: determining to arrange displacement vectors and attribute maps of three-dimensional (3D) visual media data into a single picture for processing by a two-dimensional (2D) codec; generating a bitstream based on the determination; and storing the bitstream in a non-transitory computer-readable recording medium.

[0185] 22. A method, apparatus or system as described in the present disclosure.

[0186] In the described solution, an encoder can conform to the format rules by generating a codec representation according to the format rules. In the described solution, a decoder can parse syntax elements in the codec representation according to the format rules using known information about the presence and absence of syntax elements to produce decoded video.

[0187] In the present disclosure, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. For example, the bitstream representation of a current video block may correspond to bits spread at the same position in the bitstream or at different positions as defined by the syntax. For example, a macroblock may be encoded based on an error residual value after transformation and encoding and decoding, and bits in the header and other fields in the bitstream may also be used. Furthermore, during conversion, the decoder may parse the bitstream based on this determination, knowing that some fields may or may not be present, as described in the above solution. Similarly, the encoder may determine whether to include or not include specific syntax fields, and generate the codec representation accordingly by including the syntax fields or excluding the syntax fields from the codec representation.

[0188] The disclosed and other solutions, examples, embodiments, modules, and functional operations described in this disclosure may be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this disclosure and their structural equivalents, or in a combination of one or more thereof. The disclosed and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or to control the operation of a data processing apparatus. The computer-readable medium may be a machine-readable storage device, a machine-readable storage substrate, a storage device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" includes all devices, equipment, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, an apparatus may also include code that creates an execution environment for an associated computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0189] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including stand-alone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that preserves other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to a related program, or in multiple collaborative files (e.g., files storing one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers that are located at a site or are distributed across multiple sites and interconnected by a communication network.

[0190] The processes and logic flows described in this disclosure may be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows may also be performed by, and the apparatus may be implemented as, special-purpose logic circuitry, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).

[0191] Processors suitable for executing computer programs include, for example, general-purpose and special-purpose microprocessors, and any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal or removable hard disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated into, special-purpose logic circuitry.

[0192] Although this disclosure contains many details, these details should not be interpreted as limitations on any subject matter or the scope of what may be claimed, but rather as descriptions of features that are unique to particular embodiments of particular technologies. In this disclosure, certain features described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable subcombination. In addition, although features may function in certain combinations as described above, and may even be initially claimed in this manner, in some cases, one or more features in a claimed combination may be omitted from that combination, and a claimed combination may be directed to a subcombination or a variant of a subcombination.

[0193] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed sequentially in the particular order or sequence shown, or that all illustrated operations be performed to achieve desired results. Furthermore, the partitioning of various system components in the embodiments described in this disclosure should not be understood as requiring such partitioning in all embodiments.

[0194] Although this disclosure describes only a few implementations and examples, other implementations, improvements, and variations may be made based on what is described and shown in this disclosure.

[0195] A first component is directly coupled to a second component when there are no intervening components other than a line, trace, or other medium between the first and second components. A first component is indirectly coupled to a second component when there are intervening components other than a line, trace, or other medium between the first and second components. The term "coupled" and its variations encompass both direct and indirect couplings. The use of the term "about" is intended to encompass a range of ±10% of the subsequent figure unless otherwise indicated.

[0196] Although several embodiments are provided in this disclosure, it should be understood that the disclosed systems and methods can be embodied in many other specific forms without departing from the spirit or scope of the present disclosure. The present examples are to be considered illustrative rather than restrictive, and are not intended to be limited to the details given herein. For example, various elements or components can be combined or integrated into another system, or certain features can be omitted or not implemented.

[0197] In addition, the techniques, systems, subsystems, and methods described and illustrated as discrete or separate in various embodiments may be combined or integrated with other systems, modules, techniques, or methods without departing from the scope of this disclosure. Other items illustrated or discussed as coupled may be directly connected, or may be indirectly coupled or communicated through some interface, device, or intermediate component (whether electrically, mechanically, or otherwise). Other examples of changes, substitutions, and variations that may be determined by those skilled in the art may be made without departing from the spirit and scope disclosed herein.

Claims

1. A method for processing video or image data, comprising: Determining the arrangement of displacement vectors and property maps of three-dimensional (3D) visual media data into a single picture for processing by a two-dimensional (2D) codec; as well as Conversion between visual media data and a bitstream is performed based on the displacement vector and the property map. 2 . The method of claim 1 , further comprising concatenating the displacement vectors and the property map to form the single picture.

3. The method according to claim 1 further includes determining to arrange the displacement vector and the property map in different parts of the single picture, and indicating position information or size information of each of the different parts in the bitstream.

4. The method according to claim 3, further comprising indicating the position information or the size information in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a picture header or a slice header of the bitstream.

5. The method according to claim 3 further includes deriving the position information or the size information from information in a sequence parameter set (SPS), a picture parameter set (PPS), a video parameter set (VPS), a picture header or a slice header of the bitstream.

6. The method of claim 3, further comprising determining to arrange the displacement vector and the property map in different parts of the single picture, wherein each of the different independent parts can be independently decoded. 7 . The method according to claim 6 , further comprising determining to arrange the displacement vector and the property map in different slices, different slices, or different sub-pictures of the single picture.

8. The method of claim 6, further comprising using constrained intra prediction and / or motion constrained slice sets so that the displacement vector and the property map can be independently decoded.

9. The method of claim 1, further comprising aligning the displacement vectors and the property map to a single bit depth before determining to arrange the displacement vectors and the property map into the single picture.

10. The method of claim 9, further comprising aligning the bit depth of the displacement vector to the property map such that the displacement vector and the property map are aligned to the single bit depth.

11. The method of claim 9, further comprising aligning a bit depth of the property map to the displacement vector such that the displacement vector and the property map are aligned to the single bit depth. 12 . The method of claim 9 , further comprising aligning a portion of data having a lower bit depth to a portion of data having a higher bit depth such that the displacement vectors and the property map are aligned to the single bit depth.

13. The method of claim 1, further comprising aligning the displacement vectors and the property map to a single color format before determining to arrange the displacement vectors and the property map into the single picture.

14. The method of claim 13, further comprising aligning a color format of the displacement vectors to the property map such that the displacement vectors and the property map are aligned to the single color format.

15. The method of claim 13, further comprising aligning a color format of the property map to the displacement vector such that the displacement vector and the property map are aligned to the single color format. 16 . The method of claim 13 , further comprising aligning a portion of data having less color format information to a portion of data having more color format information so that the displacement vector and the attribute map are aligned to the single color format.

17. The method of any one of claims 1-16, further comprising determining to arrange a motion field of a mesh together with the displacement vectors and the property map into the single picture for processing by the 2D codec.

18. The method of claim 17, further comprising determining to arrange any two portions or all portions of the motion field of the mesh, the displacement vectors, and the property map into the single picture for processing by the 2D codec.

19. The method of claim 17, further comprising aligning any two parts or all parts of the motion field, the displacement vectors, and the property map of the grid to a single bit depth before determining to arrange the motion field, the displacement vectors, and the property map of the grid into the single picture.

20. The method of claim 17, further comprising aligning any two portions or all portions of the motion field, the displacement vectors, and the property map of the grid to a single color format before determining to arrange the motion field, the displacement vectors, and the property map of the grid into the single picture.

21. The method of claim 20, wherein the single color format comprises two or more portions of the motion field of the mesh, the displacement vectors, and the property map having the most color information.

22. The method of claim 1, further comprising determining to arrange any data capable of being processed by the 2D codec into the single picture for processing by the 2D codec.

23. The method of any one of claims 1-22, further comprising determining to arrange two or more portions of an occupancy map, a geometry map, and a property map of a video-based point cloud codec into the single picture for processing by the 2D codec.

24. The method of any one of claims 1-22, further comprising determining an arrangement of an occupancy map and a geometry map compliant with the Moving Picture Experts Group (MPEG) immersive video standard into the single picture for processing by the 2D codec.

25. The method of any one of claims 1-22, further comprising aligning any data arranged into the single picture to the same bit depth.

26. The method of any one of claims 1-22, further comprising determining to arrange only data having the same bit depth into the single picture.

27. The method of any one of claims 1-22, wherein any data having a different bitmap is not arranged into the single picture.

28. The method of any one of claims 1-22, further comprising aligning any data arranged in the single picture to have the same color format.

29. The method according to any one of claims 1-22, further comprising determining to arrange only data having the same color format into the single picture.

30. The method of any of claims 1-22, wherein any data having a different bit depth is not arranged into the single picture.

31. The method of any one of claims 1-22, further comprising determining to arrange only data having the same bit depth and the same color format into the single picture.

32. The method of any of claims 1-22, wherein any data having a different bit depth or a different color format is not arranged into the single picture.

33. The method of any one of claims 1-32, further comprising determining any combination of a motion field of a grid arranging 3D visual media data, the displacement vectors, and the property map.

34. The method of any one of claims 1-33, wherein the converting comprises encoding the media data into a bitstream.

35. The method of any one of claims 1-33, wherein the converting comprises decoding the media data from a bitstream.

36. An apparatus for processing media data, comprising: processor; and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any one of claims 1-35.

37. A non-transitory computer-readable medium comprising a computer program product for use by a video codec device, the computer program product comprising computer-executable instructions stored on the non-transitory computer-readable medium, such that when executed by a processor, the video codec device performs the method of any one of claims 1-35.

38. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by a video processing device, wherein the method comprises the method of any one of claims 1 to 35.

39. A method for storing a bitstream of a video, comprising the method of any one of claims 1-35.

40. A method, apparatus or system as described in the present disclosure.