Encoding and decoding method and device, encoder, decoder and storage medium
Patent Information
- Application Number
- CN202280100710.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-05-13
AI Technical Summary
Existing encoding and decoding technologies require a large number of codecs when processing multi-view video and point cloud encoding, resulting in high encoding and decoding costs and making it difficult to meet the requirements of high-efficiency encoding.
By mixing and encoding visual media content in different expression formats, heterogeneous or homogeneous spliced images are generated through splicing and splicing information processing, reducing the number of codecs and adapting blocks of different expression formats at high-level parameters, thereby improving encoding efficiency.
It reduces the implementation cost of encoding and decoding, improves encoding efficiency and the quality of reconstructed multi-view video or point cloud video, and enhances the ease of use of encoding and decoding.
Smart Images

Figure CN119999217A_ABST
Abstract
Description
A coding and decoding method, device, encoder, decoder and storage medium Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a coding and decoding method, device, encoder, decoder and storage medium. Background Art
[0002] In 3D applications such as virtual reality (VR), augmented reality (AR), and mixed reality (MR), visual media objects with different expression formats may appear in the same scene. For example, in the same 3D scene, the scene background and some characters and objects may be expressed as video, while other characters may be expressed as 3D point clouds or 3D meshes.
[0003] Using multi-view video coding, point cloud coding, and grid coding respectively during compression coding can better maintain the effective information of the original expression format than projecting all into multi-view video coding, improve the quality of the rendered viewing window during viewing, and improve the overall efficiency of bit rate-quality.
[0004] However, the current coding and decoding technology encodes and decodes multi-view videos, point cloud coding, and grid-grid separately, which requires calling a large number of codecs, making the coding and decoding costly.
[0005] Summary of the Invention
[0006] The embodiments of the present application provide a coding and decoding method, apparatus, encoder, decoder, and storage medium.
[0007] In a first aspect, the present application provides a decoding method, comprising: decoding a bitstream to obtain a mosaic graph and mosaic graph information; when the mosaic graph is a heterogeneous mixed mosaic graph, obtaining at least two isomorphic blocks and isomorphic block information based on the mosaic graph and the mosaic graph information; wherein different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information; when the mosaic graph is a homogeneous mosaic graph, obtaining a type of isomorphic block and isomorphic block information based on the mosaic graph and the mosaic graph information; and obtaining visual media content in at least two expression formats based on the isomorphic blocks and the isomorphic block information.
[0008] In a second aspect, the present application provides an encoding method, comprising: processing visual media content in at least two expression formats to obtain at least two isomorphic blocks; splicing the at least two isomorphic blocks to obtain a splicing graph and splicing graph information, wherein, when the splicing graph is a heterogeneous mixed splicing graph, the splicing graph includes at least two isomorphic blocks, and different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information; encoding the splicing graph and splicing graph information to obtain a code stream.
[0009] In a third aspect, the present application provides a decoding device, comprising:
[0010] a decoding unit configured to decode the code stream to obtain a splicing graph and splicing graph information;
[0011] a splitting unit configured to obtain at least two types of isomorphic blocks and isomorphic block information based on the mosaic and the mosaic information when the mosaic is a heterogeneous mixed mosaic; wherein different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information;
[0012] The splitting unit is configured to obtain an isomorphic block and isomorphic block information according to the splicing graph and the splicing graph information when the splicing graph is a isomorphic splicing graph;
[0013] The processing unit is configured to obtain visual media content in at least two expression formats according to the isomorphic blocks and the isomorphic block information.
[0014] In a fourth aspect, the present application provides an encoding device, applied to an encoder, comprising:
[0015] a processing unit configured to process visual media content in at least two expression formats to obtain at least two isomorphic blocks;
[0016] a splicing unit configured to splice the at least two isomorphic blocks to obtain a spliced graph and spliced graph information, wherein when the spliced graph is a heterogeneous mixed spliced graph, the spliced graph includes at least two isomorphic blocks, and different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information;
[0017] The encoding unit is configured to encode the splicing graph and the splicing graph information to obtain a code stream.
[0018] In a fifth aspect, a decoder is provided, comprising a first memory and a first processor; the first memory stores a computer program that can be run on the first processor to execute the method in the above-mentioned first aspect or its various implementations.
[0019] In a sixth aspect, an encoder is provided, comprising a second memory and a second processor; the second memory stores a computer program that can be run on the second processor to execute the method in the above-mentioned second aspect or its various implementations.
[0020] In a seventh aspect, a coding and decoding system is provided, comprising an encoder and a decoder. The encoder is configured to execute the method of the second aspect or its respective implementations, and the decoder is configured to execute the method of the first aspect or its respective implementations.
[0021] In an eighth aspect, a chip is provided for implementing the method described in any one of the first and second aspects above, or their respective implementations. Specifically, the chip includes a processor configured to load and execute a computer program from a memory, causing a device equipped with the chip to perform the method described in any one of the first and second aspects above, or their respective implementations.
[0022] In a ninth aspect, a computer-readable storage medium is provided for storing a computer program, which enables a computer to execute the method of any one of the first to second aspects or their respective implementations.
[0023] In a tenth aspect, a computer program product is provided, comprising computer program instructions, which enable a computer to execute the method of any one of the first to second aspects or their respective implementations.
[0024] In an eleventh aspect, a computer program is provided, which, when executed on a computer, enables the computer to execute the method in any one of the first to second aspects or their respective implementations.
[0025] In a twelfth aspect, a code stream is provided, which is generated based on the encoding method of the second aspect.
[0026] Based on the above technical solution, for application scenarios involving visual media content in one or more expression formats, homogeneous blocks of different expression formats are spliced into a heterogeneous hybrid mosaic graph, and data of different expression formats is mixed and encoded. This can reduce the number of encoders and decoders called, lower implementation costs, and improve usability. Furthermore, in the heterogeneous hybrid mosaic graph, certain high-level parameters of blocks of different expression formats can be unequal, thereby providing more appropriate high-level parameters for heterogeneous data, effectively improving encoding efficiency, namely reducing bitrate or improving the quality of reconstructed multi-view video or point cloud video. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] FIG1 is a schematic block diagram of a video encoding and decoding system according to an embodiment of the present application;
[0028] FIG2A is a schematic block diagram of a video encoder according to an embodiment of the present application;
[0029] FIG2B is a schematic block diagram of a video decoder according to an embodiment of the present application;
[0030] FIG3A is a diagram showing the organization and expression framework of multi-view video data;
[0031] FIG3B is a schematic diagram of generating a stitched image from multi-view video data;
[0032] FIG3C is a diagram showing the organization and expression framework of point cloud data;
[0033] 3D to 3F are schematic diagrams of different types of point cloud data;
[0034] FIG4 is a schematic diagram of encoding of a multi-view video;
[0035] FIG5 is a schematic diagram of decoding of multi-view video;
[0036] FIG6 is a schematic diagram of a coding method flow chart provided in an embodiment of the present application;
[0037] FIG7 is a schematic diagram of a heterogeneous mixed splicing graph provided in one embodiment of the present application;
[0038] FIG8 is a schematic diagram of an isomorphic splicing graph provided by an embodiment of the present application;
[0039] FIG9 is a schematic flow chart of a decoding method provided in an embodiment of the present application;
[0040] FIG10 is a schematic diagram of a V3C bitstream structure provided in an embodiment of the present application;
[0041] FIG11 is a schematic block diagram of an encoding device provided in an embodiment of the present application;
[0042] FIG12 is a schematic block diagram of a decoding device provided in an embodiment of the present application;
[0043] FIG13 is a schematic block diagram of an encoder provided in an embodiment of the present application;
[0044] FIG14 is a schematic block diagram of a decoder provided in accordance with an embodiment of the present application;
[0045] FIG15 is a schematic diagram of the composition structure of a coding and decoding system provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The present application can be applied to the field of image coding and decoding, the field of video coding and decoding, the field of hardware video coding and decoding, the field of dedicated circuit video coding and decoding, the field of real-time video coding and decoding, etc. For example, the solution of the present application can be combined with an audio and video coding standard (AVS), such as the H.264 / audio and video coding (AVC) standard, the H.265 / high efficiency video coding (HEVC) standard, and the H.266 / versatile video coding (VVC) standard. Alternatively, the solution of the present application can be combined with other proprietary or industry standards and operated, and the standards include ITU-TH.261, ISO / IEC MPEG-1 Visual, ITU-TH.262 or ISO / IEC MPEG-2 Visual, ITU-TH.263, ISO / IEC MPEG-4 Visual, ITU-TH.264 (also known as ISO / IEC MPEG-4 AVC), including scalable video coding (SVC) and multi-view video coding (MVC) extensions. It should be understood that the technology of this application is not limited to any specific coding standard or technology.
[0047] The high-degree-of-freedom immersive coding system can be roughly divided into the following links according to the task line: data acquisition, data organization and expression, data encoding and compression, data decoding and reconstruction, data synthesis and rendering, and finally presenting the target data to the user.
[0048] The encoding involved in the embodiment of the present application is mainly video encoding and decoding. For ease of understanding, the video encoding and decoding system involved in the embodiment of the present application is first introduced in conjunction with Figure 1.
[0049] FIG1 is a schematic block diagram of a video encoding and decoding system involved in an embodiment of the present application. It should be noted that FIG1 is only an example, and the video encoding and decoding system of the embodiment of the present application includes but is not limited to that shown in FIG1. As shown in FIG1, the video encoding and decoding system 100 includes an encoding device 110 and a decoding device 120. The encoding device is used to encode (which can be understood as compressing) the video data to generate a code stream, and transmit the code stream to the decoding device. The decoding device decodes the code stream generated by the encoding device to obtain decoded video data.
[0050] The encoding device 110 of the embodiment of the present application can be understood as a device with a video encoding function, and the decoding device 120 can be understood as a device with a video decoding function, that is, the embodiment of the present application includes a wider range of devices for the encoding device 110 and the decoding device 120, such as smartphones, desktop computers, mobile computing devices, notebook (e.g., laptop) computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, car computers, etc.
[0051] In some embodiments, the encoding device 110 may transmit the encoded video data (eg, a code stream) to the decoding device 120 via a channel 130. The channel 130 may include one or more media and / or devices capable of transmitting the encoded video data from the encoding device 110 to the decoding device 120.
[0052] In one example, the channel 130 includes one or more communication media that enable the encoding device 110 to transmit the encoded video data directly to the decoding device 120 in real time. In this example, the encoding device 110 can modulate the encoded video data according to a communication standard and transmit the modulated video data to the decoding device 120. The communication media includes wireless communication media, such as radio frequency spectrum. Optionally, the communication media may also include wired communication media, such as one or more physical transmission lines.
[0053] In another example, channel 130 includes a storage medium that can store the video data encoded by encoding device 110. The storage medium includes various locally accessible data storage media, such as optical disks, DVDs, and flash memories. In this example, decoding device 120 can retrieve the encoded video data from the storage medium.
[0054] In another example, the channel 130 may include a storage server that can store the video data encoded by the encoding device 110. In this example, the decoding device 120 can download the stored encoded video data from the storage server. Alternatively, the storage server can store the encoded video data and transmit the encoded video data to the decoding device 120, such as a web server (e.g., for a website), a file transfer protocol (FTP) server, etc.
[0055] In some embodiments, the encoding device 110 includes a video encoder 112 and an output interface 113. The output interface 113 may include a modulator / demodulator (modem) and / or a transmitter.
[0056] In some embodiments, the encoding device 110 may further include a video source 111 in addition to the video encoder 112 and the input interface 113 .
[0057] The video source 111 may include at least one of a video acquisition device (eg, a video camera), a video archive, a video input interface, and a computer graphics system, wherein the video input interface is used to receive video data from a video content provider, and the computer graphics system is used to generate video data.
[0058] The video encoder 112 encodes the video data from the video source 111 to generate a bitstream. The video data may include one or more pictures or a sequence of pictures. The bitstream contains the coding information of the picture or picture sequence in the form of a bitstream. The coding information may include the coded picture data and associated data. The associated data may include a sequence parameter set (SPS), a picture parameter set (PPS), and other syntax structures. The SPS may contain parameters that apply to one or more sequences. The PPS may contain parameters that apply to one or more pictures. The syntax structure refers to a set of zero or more syntax elements arranged in a specified order in the bitstream.
[0059] The video encoder 112 transmits the encoded video data directly to the decoding device 120 via the output interface 113. The encoded video data may also be stored in a storage medium or a storage server for subsequent reading by the decoding device 120.
[0060] In some embodiments, the decoding device 120 includes an input interface 121 and a video decoder 122. In some embodiments, the decoding device 120 may include a display device 123 in addition to the input interface 121 and the video decoder 122.
[0061] The input interface 121 includes a receiver and / or a modem and can receive the encoded video data via the channel 130 .
[0062] The video decoder 122 is configured to decode the encoded video data to obtain decoded video data, and transmit the decoded video data to the display device 123 .
[0063] The decoded video data is displayed on the display device 123. The display device 123 may be integrated with the decoding device 120 or external to the decoding device 120. The display device 123 may include various display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0064] In addition, Figure 1 is only an example, and the technical solution of the embodiment of the present application is not limited to Figure 1. For example, the technology of the present application can also be applied to unilateral video encoding or unilateral video decoding.
[0065] The following is an introduction to the video encoding framework involved in the embodiments of the present application.
[0066] FIG2A is a schematic block diagram of a video encoder according to an embodiment of the present application. It should be understood that the video encoder 200 can be used to perform lossy compression or lossless compression on an image. The lossless compression can be visually lossless or mathematically lossless.
[0067] The video encoder 200 can be applied to image data in a luminance and chrominance (YCbCr, YUV) format. For example, the YUV ratio can be 4:2:0, 4:2:2, or 4:4:4, where Y represents luminance (Luma), Cb (U) represents blue chrominance, Cr (V) represents red chrominance, and U and V represent chrominance (Chroma) used to describe color and saturation. For example, in terms of color format, 4:2:0 means that every 4 pixels have 4 luminance components and 2 chrominance components (YYYYCbCr), 4:2:2 means that every 4 pixels have 4 luminance components and 4 chrominance components (YYYYCbCrCbCr), and 4:4:4 represents a full pixel display (YYYYCbCrCbCrCbCrCbCr).
[0068] For example, the video encoder 200 reads video data, and for each frame of the video data, divides the frame into a number of coding tree units (CTUs). In some examples, CTB may be referred to as a "tree block", "largest coding unit" (LCU) or "coding tree block" (CTB). Each CTU may be associated with a pixel block of equal size within the image. Each pixel may correspond to a luminance (luminance or luma) sample and two chrominance (chroma) samples. Therefore, each CTU may be associated with a luminance sample block and two chroma sample blocks. The size of a CTU is, for example, 128×128, 64×64, 32×32, etc. A CTU can be further divided into a number of coding units (CUs) for encoding. The CU may be a rectangular block or a square block. A CU can be further divided into prediction units (PUs) and transform units (TUs), allowing for separation of coding, prediction, and transform, and greater flexibility in processing. In one example, a CTU is divided into CUs using a quadtree, and a CU is divided into TUs and PUs using a quadtree.
[0069] The video encoder and video decoder can support various PU sizes. Assuming that the size of a particular CU is 2N×2N, the video encoder and video decoder can support PU sizes of 2N×2N or N×N for intra-frame prediction, and support symmetric PUs of 2N×2N, 2N×N, N×2N, N×N, or similar sizes for inter-frame prediction. The video encoder and video decoder can also support asymmetric PUs of 2N×nU, 2N×nD, nL×2N, and nR×2N for inter-frame prediction.
[0070] In some embodiments, as shown in FIG2A , the video encoder 200 may include a prediction unit 210, a residual unit 220, a transform / quantization unit 230, an inverse transform / quantization unit 240, a reconstruction unit 250, a loop filter unit 260, a decoded image buffer 270, and an entropy coding unit 280. It should be noted that the video encoder 200 may include more, fewer, or different functional components.
[0071] Optionally, in this application, the current block may be referred to as the current coding unit (CU) or the current prediction unit (PU), etc. The prediction block may also be referred to as a predicted image block or an image prediction block, and the reconstructed image block may also be referred to as a reconstructed block or an image reconstructed image block.
[0072] In some embodiments, the prediction unit 210 includes an inter-frame prediction unit 211 and an intra-frame estimation unit 212. Because there is a strong correlation between adjacent pixels in a video frame, intra-frame prediction is used in video coding and decoding technologies to eliminate spatial redundancy between adjacent pixels. Because there is a strong similarity between adjacent frames in a video, inter-frame prediction is used in video coding and decoding technologies to eliminate temporal redundancy between adjacent frames, thereby improving coding efficiency.
[0073] The inter-frame prediction unit 211 is used for inter-frame prediction. Inter-frame prediction includes motion estimation and motion compensation. It can refer to image information from different frames. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block to eliminate temporal redundancy. The frames used for inter-frame prediction can be P-frames and / or B-frames. P-frames refer to forward-predicted frames, while B-frames refer to bidirectionally predicted frames. Inter-frame prediction uses motion information to find a reference block from a reference frame and generates a prediction block based on the reference block. Motion information includes the reference frame list in which the reference frame is located, the reference frame index, and the motion vector. The motion vector can be integer pixel or fractional pixel. If the motion vector is fractional pixel, interpolation filtering is required to generate the required fractional pixel block in the reference frame. Here, the integer pixel or fractional pixel block in the reference frame found based on the motion vector is called a reference block. Some technologies directly use the reference block as the prediction block, while others further process the reference block to generate a prediction block. Reprocessing the reference block to generate a prediction block can also be understood as using the reference block as the prediction block and then processing the prediction block to generate a new prediction block.
[0074] The intra-frame estimation unit 212 only refers to the information of the same frame image to predict the pixel information in the current code image block to eliminate spatial redundancy. The frame used for intra-frame prediction can be an I frame.
[0075] Intra-frame prediction has multiple prediction modes. For example, the H-series international digital video coding standard H.264 / AVC has eight angular prediction modes and one non-angular prediction mode. H.265 / HEVC expands this to 33 angular prediction modes and two non-angular prediction modes. HEVC uses planar, DC, and 33 angular modes for a total of 35 intra-frame prediction modes. VVC uses planar, DC, and 65 angular modes for a total of 67 intra-frame prediction modes.
[0076] It should be noted that with the increase of angle modes, intra-frame prediction will be more accurate and more in line with the needs of the development of high-definition and ultra-high-definition digital videos.
[0077] The residual unit 220 may generate a residual block for the CU based on the pixel blocks of the CU and the prediction blocks of the PUs of the CU. For example, the residual unit 220 may generate the residual block for the CU such that each sample in the residual block has a value equal to the difference between the sample in the pixel blocks of the CU and the corresponding sample in the prediction blocks of the PUs of the CU.
[0078] The transform / quantization unit 230 may quantize the transform coefficients. The transform / quantization unit 230 may quantize the transform coefficients associated with the TUs of a CU based on a quantization parameter (QP) value associated with the CU. The video encoder 200 may adjust the degree of quantization applied to the transform coefficients associated with the CU by adjusting the QP value associated with the CU.
[0079] The inverse transform / quantization unit 240 may apply inverse quantization and inverse transform, respectively, to the quantized transform coefficients to reconstruct a residual block from the quantized transform coefficients.
[0080] Reconstruction unit 250 may add samples of the reconstructed residual block to corresponding samples of one or more prediction blocks generated by prediction unit 210 to generate a reconstructed image block associated with the TU. By reconstructing the sample blocks of each TU of a CU in this manner, video encoder 200 can reconstruct the pixel blocks of the CU.
[0081] The loop filter unit 260 is used to process the inverse transformed and inverse quantized pixels to compensate for distortion information and provide a better reference for subsequent coded pixels. For example, it can perform a deblocking filtering operation to reduce the blocking effect of pixel blocks associated with the CU.
[0082] In some embodiments, the loop filtering unit 260 includes a deblocking filtering unit and a sample adaptive offset / adaptive loop filtering (SAO / ALF) unit, wherein the deblocking filtering unit is used to remove blocking effects, and the SAO / ALF unit is used to remove ringing effects.
[0083] The decoded image buffer 270 may store the reconstructed pixel blocks. The inter prediction unit 211 may use a reference image containing the reconstructed pixel blocks to perform inter prediction on PUs of other images. In addition, the intra estimation unit 212 may use the reconstructed pixel blocks in the decoded image buffer 270 to perform intra prediction on other PUs in the same image as the CU.
[0084] The entropy coding unit 280 may receive the quantized transform coefficients from the transform / quantization unit 230. The entropy coding unit 280 may perform one or more entropy coding operations on the quantized transform coefficients to generate entropy-coded data.
[0085] FIG2B is a schematic block diagram of a video decoder according to an embodiment of the present application.
[0086] 2B , the video decoder 300 includes an entropy decoding unit 310, a prediction unit 320, an inverse quantization / transformation unit 330, a reconstruction unit 340, a loop filter unit 350, and a decoded picture buffer 360. It should be noted that the video decoder 300 may include more, fewer, or different functional components.
[0087] The video decoder 300 may receive a bitstream. The entropy decoding unit 310 may parse the bitstream to extract syntax elements from the bitstream. As part of parsing the bitstream, the entropy decoding unit 310 may parse the entropy-encoded syntax elements in the bitstream. The prediction unit 320, the inverse quantization / transform unit 330, the reconstruction unit 340, and the loop filter unit 350 may decode the video data based on the syntax elements extracted from the bitstream, thereby generating decoded video data.
[0088] In some embodiments, the prediction unit 320 includes an inter-frame prediction unit 321 and an intra-frame estimation unit 322 .
[0089] Intra estimation unit 322 may perform intra prediction to generate a prediction block for a PU. Intra estimation unit 322 may use an intra prediction mode to generate a prediction block for the PU based on pixel blocks of spatially neighboring PUs. Intra estimation unit 322 may also determine the intra prediction mode for the PU based on one or more syntax elements parsed from a codestream.
[0090] The inter-frame prediction unit 321 may construct a first reference picture list (List 0) and a second reference picture list (List 1) based on syntax elements parsed from the codestream. In addition, if a PU is encoded using inter-frame prediction, the entropy decoding unit 310 may parse the motion information of the PU. The inter-frame prediction unit 321 may determine one or more reference blocks for the PU based on the motion information of the PU. The inter-frame prediction unit 321 may generate a prediction block for the PU based on the one or more reference blocks of the PU.
[0091] The inverse quantization / transform unit 330 may inversely quantize (ie, dequantize) the transform coefficients associated with the TU. The inverse quantization / transform unit 330 may use the QP value associated with the CU of the TU to determine the degree of quantization.
[0092] After inverse quantizing the transform coefficients, the inverse quantization / transform unit 330 may apply one or more inverse transforms to the inverse quantized transform coefficients in order to generate a residual block associated with the TU.
[0093] The reconstruction unit 340 uses the residual block associated with the TU of the CU and the prediction block of the PU of the CU to reconstruct the pixel block of the CU. For example, the reconstruction unit 340 can add samples of the residual block to corresponding samples of the prediction block to reconstruct the pixel block of the CU to obtain a reconstructed image block.
[0094] The loop filtering unit 350 may perform a deblocking filtering operation to reduce blocking artifacts of pixel blocks associated with a CU.
[0095] The video decoder 300 may store the reconstructed image of the CU in the decoded image buffer 360. The video decoder 300 may use the reconstructed image in the decoded image buffer 360 as a reference image for subsequent prediction, or transmit the reconstructed image to a display device for presentation.
[0096] The basic process of video encoding and decoding is as follows: At the encoder end, a frame of image is divided into blocks. For the current block, the prediction unit 210 uses intra-frame prediction or inter-frame prediction to generate a prediction block for the current block. The residual unit 220 calculates a residual block based on the predicted block and the original block of the current block. This residual block is the difference between the predicted block and the original block of the current block. This residual block is also called residual information. This residual block undergoes transformation and quantization by the transform / quantization unit 230, removing information that is insensitive to the human eye and eliminating visual redundancy. Optionally, the residual block before transformation and quantization by the transform / quantization unit 230 is called a time-domain residual block, and the time-domain residual block after transformation and quantization by the transform / quantization unit 230 is called a frequency residual block or a frequency-domain residual block. The entropy coding unit 280 receives the quantized change coefficients output by the change quantization unit 230 and performs entropy coding on these quantized change coefficients to output a bitstream. For example, the entropy coding unit 280 can eliminate character redundancy based on the target context model and probability information of the binary bitstream.
[0097] At the decoding end, the entropy decoding unit 310 can parse the code stream to obtain the prediction information, quantization coefficient matrix, etc. of the current block. The prediction unit 320 uses intra-frame prediction or inter-frame prediction for the current block based on the prediction information to generate a prediction block for the current block. The inverse quantization / transformation unit 330 uses the quantization coefficient matrix obtained from the code stream to inverse quantize and inverse transform the quantization coefficient matrix to obtain a residual block. The reconstruction unit 340 adds the prediction block and the residual block to obtain a reconstructed block. The reconstructed blocks constitute a reconstructed image, and the loop filtering unit 350 performs loop filtering on the reconstructed image based on the image or block to obtain a decoded image. The encoding end also requires similar operations as the decoding end to obtain a decoded image. The decoded image can also be called a reconstructed image, and the reconstructed image can be used as a reference frame for inter-frame prediction for subsequent frames.
[0098] It should be noted that the block division information determined by the encoder, as well as mode information or parameter information such as prediction, transform, quantization, entropy coding, and loop filtering, etc., are carried in the bitstream when necessary. The decoder parses the bitstream and analyzes the existing information to determine the same block division information, prediction, transform, quantization, entropy coding, loop filtering, etc. mode information or parameter information as the encoder, thereby ensuring that the decoded image obtained by the encoder and the decoder are identical.
[0099] The above is the basic process of the video codec under the block-based hybrid coding framework. With the development of technology, some modules or steps of the framework or process may be optimized. This application is applicable to the basic process of the video codec under the block-based hybrid coding framework, but is not limited to the framework and process.
[0100] In some application scenarios, multiple heterogeneous contents appear simultaneously in the same 3D scene, such as multi-view video and point cloud. For this situation, current encoding and decoding methods include at least the following two:
[0101] Method 1: MPEG (Moving Picture Experts Group) immersive video (MIV) technology is used for encoding and decoding multi-view videos, and point cloud video compression (VPCC) technology is used for encoding and decoding point clouds.
[0102] The following introduces MIV technology and VPCC technology.
[0103] MIV technology: To reduce the transmission pixel rate while preserving as much scene information as possible to ensure sufficient information for rendering the target view, MPEG-I uses a solution, as shown in Figure 3A. A limited number of viewpoints are selected as base viewpoints, and the scene's visible range is expressed as much as possible. The base viewpoints are transmitted as complete images, and redundant pixels between the remaining non-base viewpoints and the base viewpoint are removed, retaining only the valid information that is not repeated. This valid information is then extracted into sub-block images, which are reorganized with the base viewpoint images to form a larger rectangular image, called a spliced image. Figures 3A and 3B illustrate the generation process of the spliced image. The spliced image is fed into the codec for compression and reconstruction, and auxiliary data related to the splicing information of the sub-block images is also fed into the encoder to form the bitstream.
[0104] The VPCC encoding method is to project the point cloud into a two-dimensional image or video, converting the three-dimensional information into a two-dimensional information code. Figure 3C is the encoding block diagram of VPCC. The code stream is roughly divided into four parts. The geometry code stream is the code stream generated by the geometric depth map encoding, which is used to represent the geometric information of the point cloud; the attribute code stream is the code stream generated by the texture map encoding, which is used to represent the attribute information of the point cloud; the occupancy code stream is the code stream generated by the occupancy map encoding, which is used to indicate the valid area in the depth map and texture map. All three types of videos are encoded and decoded using a video encoder, as shown in Figures 3D to 3F. The auxiliary information code stream is the code stream generated by the encoding of the auxiliary information of the sub-block image, that is, the part related to the patch data unit in the V3C standard, which indicates information such as the position and size of each sub-block image.
[0105] In the second method, both multi-view videos and point clouds are encoded and decoded using the frame packing technology in Visual Volumetric Video-based Coding (V3C).
[0106] The following is an introduction to frame packing technology.
[0107] Taking multi-view video as an example, as shown in FIG4 , the encoding end includes the following steps:
[0108] In step 1, when encoding the acquired multi-view video, multi-view video sub-blocks (patches) are generated after some pre-processing. Then, the multi-view video sub-blocks are organized to generate a multi-view video mosaic.
[0109] For example, as shown in Figure 4, a multi-view video is input into TIMV for packaging, and a multi-view video mosaic is output. TIMV is a reference software for MIV. The packaging in the embodiment of the present application can be understood as splicing.
[0110] The multi-view video mosaic map includes a multi-view video texture mosaic map and a multi-view video geometry mosaic map, that is, it only contains multi-view video sub-blocks.
[0111] Step 2: Input the multi-view video mosaic image into the frame packer, and output the multi-view video mixed mosaic image.
[0112] Among them, the multi-view video mixed mosaic map includes a multi-view video texture mixed mosaic map, a multi-view video geometry mixed mosaic map, and a multi-view video texture and geometry mixed mosaic map.
[0113] Specifically, as shown in Figure 4, the multi-view video mosaic is frame-packed to generate a multi-view video hybrid mosaic. Each multi-view video mosaic occupies a region of the multi-view video hybrid mosaic. Accordingly, a flag, pin_region_type_id_minus2, is transmitted for each region in the bitstream. This flag records whether the current region belongs to a multi-view video texture mosaic or a multi-view video geometry mosaic. This information is required by the decoder.
[0114] Step 3: Use a video encoder to encode the multi-view video mixed mosaic image to obtain a bit stream.
[0115] Exemplarily, as shown in FIG5 , the decoding end includes the following steps:
[0116] Step 1: During multi-view video decoding, the acquired code stream is input into a video decoder for decoding to obtain a reconstructed multi-view video mixed splicing image.
[0117] Step 2: input the reconstructed multi-view video mixed mosaic image into the frame depacketizer, and output the reconstructed multi-view video mosaic image.
[0118] Specifically, first, obtain the flag pin_region_type_id_minus2 from the bitstream. If it is determined that pin_region_type_id_minus2 is V3C_AVD, it means that the current region is a multi-view video texture mosaic. Then, the current region is split and output as a reconstructed multi-view video texture mosaic.
[0119] If it is determined that the pin_region_type_id_minus2 is V3C_GVD, it means that the current region is a multi-view video geometric mosaic, and the current region is split and output as a reconstructed multi-view video geometric mosaic.
[0120] Step 3: decode the reconstructed multi-view video mosaic to obtain the reconstructed multi-view video.
[0121] Specifically, the multi-view video texture mosaic map and the multi-view video geometric mosaic map are decoded to obtain the reconstructed multi-view video.
[0122] The above uses multi-view video as an example to analyze and introduce the frame packing technology. The frame packing encoding and decoding method for point cloud is basically the same as that for the above multi-view video. For example, TMC (a reference software for VPCC) is used to pack the point cloud to obtain a point cloud mosaic map, and the point cloud mosaic map is input into the frame packer for frame packing to obtain a point cloud mixed mosaic map. The point cloud mixed mosaic map is stitched to obtain a point cloud code stream. I will not go into details here.
[0123] At present, if multiple visual media contents in different expression formats appear simultaneously in the same three-dimensional scene, the visual media contents in the multiple expression formats are encoded and decoded separately. For example, in the case where point cloud and multi-view video appear simultaneously in the same three-dimensional scene, the current packaging technology is to compress the point cloud to form a point cloud compressed code stream (i.e., a V3C code stream), compress the multi-view video information to obtain a multi-view video compressed code stream (i.e., another V3C code stream), and then the system layer multiplexes the compressed code stream to obtain a fused three-dimensional scene multiplexed code stream. During decoding, the point cloud compressed code stream and the multi-view video compressed code stream are decoded separately. However, when encoding and decoding visual media contents in multiple expression formats, the existing technology uses many codecs and the encoding and decoding cost is high.
[0124] To address these technical issues, data in different formats is mixed and encoded. This reduces the number of encoders and decoders required, lowers implementation costs, and improves usability. Furthermore, in heterogeneous hybrid mosaics, certain high-level parameters of blocks in different formats can be unequal. This allows for more appropriate high-level parameters for heterogeneous data, effectively improving encoding efficiency—reducing bitrate or improving the quality of reconstructed multi-view or point cloud video.
[0125] 6 , the video encoding method provided in the embodiment of the present application is introduced by taking the encoding end as an example.
[0126] FIG6 is a flow chart of an encoding method provided in an embodiment of the present application. As shown in FIG6 , the encoding method includes:
[0127] Step 601: Processing visual media content in at least two expression formats to obtain at least two isomorphic blocks;
[0128] In three-dimensional application scenarios, such as virtual reality (VR), augmented reality (AR), and mixed reality (MR), visual media objects with different expression formats may appear in the same scene. For example, in the same three-dimensional scene, the scene background and some characters and objects may be expressed as video, while other characters may be expressed as three-dimensional point clouds or three-dimensional meshes.
[0129] In some embodiments, the visual media content includes visual media content in at least two expression formats, including multi-view video, point cloud, and mesh. Multi-view video may include multiple viewpoint videos and / or single viewpoint videos. Each type of homogeneous block corresponds to one expression format. Different types of homogeneous blocks correspond to different expression formats. Exemplarily, the expression formats corresponding to the at least two homogeneous blocks include at least two of the following: multi-view video, point cloud, and mesh.
[0130] It should be noted that in the embodiments of the present application, each homogeneous block may include at least one homogeneous block with the same expression format. For example, a homogeneous block in a point cloud format may include one or more point cloud blocks, a homogeneous region in a multi-view video format may include one or more multi-view video blocks, and a homogeneous block in a grid format may include one or more grid blocks.
[0131] Specifically, visual media content in a first expression format is processed to obtain homogeneous blocks in the first expression format, and visual media content in a second expression format is processed to obtain homogeneous blocks in the second expression format. The first expression format is one of multi-view video, point cloud, and mesh, and the second expression format is one of multi-view video, point cloud, and mesh. The first expression format and the second expression format are different expression formats.
[0132] It should be noted that a block can be a mosaic with a specific shape, such as a mosaic of a rectangular area with a specific length and / or height. For example, a block includes at least one sub-block, and at least one sub-block is spliced in order, such as from large to small according to the area of the sub-block, or from large to small according to the length and / or height of the sub-block, to obtain a block corresponding to the visual media content. Optionally, a block can be accurately mapped to a tile in a mosaic (atlas). In practical applications, a block can also be called a strip (tile), that is, a point cloud block can also be called a point cloud strip, a multi-view video block can also be called a multi-view video strip, and a grid block can also be called a grid strip.
[0133] In some embodiments, each sub-tile in a block may have a patch ID (patch ID) to distinguish different sub-tiles in the same block. For example, the same block may include sub-tile 1 (patch 1), sub-tile 2 (patch 2), and sub-tile 3 (patch 3).
[0134] Furthermore, a homogeneous block refers to a block where each sub-block has the same expression format. For example, each sub-block in a homogeneous block may be a multi-view video sub-block, or a point cloud sub-block, or other sub-blocks with the same expression format. The expression format corresponding to each sub-block in a homogeneous block is the expression format corresponding to the homogeneous block.
[0135] In some embodiments, homogeneous tiles may have tile identifiers (tileIDs) to distinguish different tiles of the same expression format. For example, a point cloud tile may include point cloud tile 1 or point cloud tile 2. For example, multiple visual media contents may include point clouds and multi-view videos. The point clouds may be processed to obtain point cloud tiles, where point cloud tile 1 includes point cloud sub-tiles 1 to 3. The multi-view videos may be processed to obtain multi-view video tiles, where the multi-view video tiles include multi-view video sub-tiles 1 to 4.
[0136] When it is necessary to process visual media content in an expression format, an isomorphic block in an expression format is obtained. When it is necessary to process at least two visual media contents, at least two isomorphic blocks in expression formats are obtained. In order to improve compression efficiency, the embodiment of the present application processes the at least two visual media contents, such as packaging (also known as splicing), to obtain blocks corresponding to each visual media content in the at least two visual media contents. For example, sub-patches corresponding to at least two visual media contents can be spliced to obtain blocks. It should be noted that the embodiment of the present application processes at least two visual media contents separately, and the method of obtaining blocks is not limited.
[0137] In one possible implementation, the visual media content includes visual media content in two expression formats, namely, multi-view video and point cloud. The visual media content in at least two expression formats is processed to obtain at least two isomorphic blocks, including: after projecting and de-redundancy processing on the acquired multi-view video, non-repeated pixels are connected into multi-view video sub-blocks, and the multi-view video sub-blocks are spliced into multi-view video blocks; and parallel projection is performed on the acquired point cloud, connected points in the projection plane are formed into point cloud sub-blocks, and the point cloud sub-blocks are spliced into point cloud blocks.
[0138] Specifically, for multi-viewpoint video, taking MPEG-I as an example, a limited number of viewpoints are selected as basic viewpoints and the visible range of the scene is expressed as much as possible. The basic viewpoints are transmitted as complete images, and the redundant pixels between the remaining non-basic viewpoints and the basic viewpoints are removed, that is, only the valid information expressed in a non-repeated manner is retained, and then the valid information is extracted into sub-block images and reorganized with the basic viewpoint images to form a larger strip-shaped image. The strip-shaped image is called a multi-viewpoint video block.
[0139] In some embodiments, the visual media content is media content presented simultaneously in the same three-dimensional space. In some embodiments, the visual media content is media content presented at different times in the same three-dimensional space. In some embodiments, the visual media content can also be media content in different three-dimensional spaces. That is, in the embodiments of the present application, there is no specific limitation on the at least two visual media contents.
[0140] Step 602: Splicing the at least two isomorphic blocks to obtain a spliced graph and spliced graph information. When the spliced graph is a heterogeneous mixed spliced graph, the spliced graph includes at least two isomorphic blocks, and different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information.
[0141] When the mosaic graph is a homogeneous mosaic graph, the mosaic graph includes a homogeneous block, and a homogeneous block corresponds to a visual media content expression format.
[0142] Specifically, heterogeneous splicing is performed on homogeneous blocks in at least two expression formats to generate a heterogeneous mixed splicing graph and splicing graph information; homogeneous splicing is performed on homogeneous blocks in the same expression format to generate an homogeneous splicing graph and splicing graph information. A heterogeneous mixed splicing graph is formed by splicing homogeneous blocks in at least two expression formats, while a homogeneous splicing graph is formed by splicing homogeneous blocks in one expression format.
[0143] Exemplarily, isomorphic splicing is performed on isomorphic blocks in a first expression format to obtain a first isomorphic splicing graph and splicing graph information, and isomorphic splicing is performed on isomorphic blocks in a second expression format to obtain a second isomorphic splicing graph and splicing graph information; or, isomorphic splicing is performed on isomorphic blocks in the first expression format and isomorphic blocks in the second expression format to obtain a heterogeneous mixed splicing graph and splicing graph information; or, isomorphic splicing is performed on isomorphic blocks in the first expression format to obtain a first isomorphic splicing graph and splicing graph information, and isomorphic splicing is performed on isomorphic blocks in the first expression format and isomorphic blocks in the second expression format to obtain a heterogeneous mixed splicing graph and splicing graph information; or, isomorphic splicing is performed on isomorphic blocks in the second expression format to obtain a second isomorphic splicing graph and splicing graph information, and isomorphic splicing is performed on isomorphic blocks in the first expression format and isomorphic blocks in the second expression format to obtain a heterogeneous mixed splicing graph and splicing graph information.
[0144] That is, a homogeneous mosaic graph can include one or more homogeneous blocks in the same expression format, and a heterogeneous mixed mosaic graph includes at least two homogeneous blocks in at least two expression formats. In this embodiment of the present application, the first expression format is one of multi-view video, point cloud, and mesh, and the second expression format is one of multi-view video, point cloud, and mesh, and the first expression format and the second expression format are different expression formats. As shown in Figure 7, multi-view video block 1, multi-view video block 2, and point cloud block 1 are spliced together to form a heterogeneous mixed mosaic graph.
[0145] For example, the first expression format is a multi-view video, and the second expression format is a point cloud. A portion of the multi-view video blocks and a portion of the point cloud blocks are spliced into a heterogeneous hybrid mosaic; another portion of the multi-view video blocks is spliced into a multi-view mosaic; and another portion of the point cloud blocks is spliced into a point cloud mosaic.
[0146] The mosaic information is used to reconstruct the mosaic. Exemplarily, the mosaic information includes at least mosaic type information, mosaic information of homogeneous blocks, and homogeneous block information. In some embodiments, the mosaic information includes a first syntax element, which is used to indicate whether the mosaic is a heterogeneous mixed mosaic or a homogeneous mosaic. In some embodiments, the first syntax element is a syntax element of a mosaic sequence parameter set (ASPS) and / or a syntax element of a mosaic frame parameter set (AFPS). The ASPS and / or AFPS are parsed to determine the mosaic type.
[0147] In some embodiments, the first syntax element includes a first sub-syntax element and a second sub-syntax element; determining whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph based on the first syntax element includes: if the value of the first sub-syntax element and the value of the second sub-syntax element are equal, determining whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph based on the values.
[0148] In some embodiments, the first syntax element includes a first sub-syntax element and a second sub-syntax element; determining whether the splicing graph is a heterogeneous mixed splicing graph or a homogeneous splicing graph based on the first syntax element includes: determining whether the splicing graph is a heterogeneous mixed splicing graph or a homogeneous splicing based on the value of the first sub-syntax element; determining whether the splicing graph is a heterogeneous mixed splicing graph or a homogeneous splicing based on the value of the second sub-syntax element; when the two judgment results are consistent, determining the splicing graph type.
[0149] In other words, the codestream must ensure absolute consistency between the first and second sub-syntax elements. The mosaic type can only be determined when the two sub-syntax elements are consistent. For example, the two sub-syntax elements can be compared for consistency before the mosaic type is determined based on the value of one of the sub-syntax elements. Alternatively, the mosaic type can be determined based on each sub-syntax element and absolute consistency can be ensured by comparing the mosaic types to ensure consistency.
[0150] Exemplarily, determining whether the mosaic is a heterogeneous mixed mosaic or a homogeneous mosaic based on the values includes: if the value is a first preset value, determining the mosaic is a heterogeneous mixed mosaic; and if the value is a second preset value, determining the mosaic is a homogeneous mosaic. In other words, two values or two types of values can be set to identify heterogeneous mixed mosaics and homogeneous mosaics. Exemplarily, the first preset value is 1 and the second preset value is 0.
[0151] Exemplarily, in some embodiments, the splice information does not include the first syntax element, and the splice is determined to be a homogeneous splice. In some embodiments, the splice information does not include the first syntax element, and the value of the first syntax element is inferred to be a second preset value. Exemplarily, the splice information does not include the first sub-syntax element, and the splice is determined to be a homogeneous splice, and the value of the first sub-syntax element is inferred to be a second preset value; the splice information does not include the second sub-syntax element, and the splice is determined to be a homogeneous splice, and the value of the second sub-syntax element is inferred to be a second preset value.
[0152] Exemplarily, the determining whether the mosaic is a heterogeneous mixed mosaic or a homogeneous mosaic based on the values includes: if the value is a third preset value, the mosaic is determined to be a heterogeneous mixed mosaic including homogeneous blocks of a first expression format and a second expression format, wherein the first expression format and the second expression format are different expression formats; if the value is a fourth preset value, the mosaic is determined to be a homogeneous mosaic including homogeneous blocks of the first expression format; if the value is a fifth preset value, the mosaic is determined to be a homogeneous mosaic including homogeneous blocks of the second expression format. That is, multiple values can also be set to identify the expression formats of heterogeneous mixed mosaics and homogeneous mosaics, and even to identify which homogeneous blocks of which expression formats are included in the heterogeneous mixed mosaic. Exemplarily, the third preset value is 2, the fourth preset value is 1, and the fifth preset value is 0.
[0153] Exemplarily, the first sub-syntax element is a syntax element of an ASPS (Assembled Picture Sequence Parameter Set), and the second sub-syntax element is a syntax element of an AFPS (Assembled Picture Frame Parameter Set). In this embodiment of the present application, asps_heterogeneous_miv_extension_present_flag represents the first sub-syntax element, and afps_heterogeneous_miv_extension_present_flag represents the second sub-syntax element.
[0154] For example, in some embodiments, when the mosaic is a heterogeneous mixed mosaic, the mosaic information further includes a second syntax element; and the expression format of the homogeneous blocks in the mosaic is determined based on the second syntax element. It is understood that after determining that the mosaic is a heterogeneous mixed mosaic based on the first syntax element, the second syntax element of the homogeneous blocks in the heterogeneous mixed mosaic is further parsed to determine the homogeneous block type.
[0155] Specifically, the expression format type corresponding to the i-th block in the spliced graph can be indicated by setting different values for the second syntax element. Exemplarily, the expression format of the isomorphic blocks in the spliced graph determined according to the second syntax element includes: when the value of the second syntax element of the i-th block is the sixth preset value, the expression format of the i-th block is determined to be the first expression format; when the value of the second syntax element of the i-th block is the seventh preset value, the expression format of the i-th block is determined to be the second expression format. Take the first expression format as point cloud and the second expression format as multi-view video as an example. Optionally, the sixth preset value is 0 and the seventh preset value is 1.
[0156] Furthermore, the expression format of the i-th block is a first expression format, and the i-th block is encoded using a coding method corresponding to the first expression format; the expression format of the i-th block is a second expression format, and the i-th block is encoded using a coding method corresponding to the second expression format.
[0157] For example, in some embodiments, when the mosaic is a heterogeneous mixed mosaic, the mosaic information includes at least two types of homogeneous block information, where homogeneous blocks in different expression formats correspond to different homogeneous block information. The homogeneous block information includes reconstruction information of the homogeneous blocks and other supplementary information for decoding and reconstructing the homogeneous blocks.
[0158] Exemplarily, the isomorphic block information includes syntax elements of ASPS and syntax elements of AFPS; different isomorphic block information corresponds to different syntax elements of the ASPS and syntax elements of the AFPS. In some embodiments, for a heterogeneous mixed mosaic, the ASPS and AFPS of isomorphic blocks of different expression formats are at least partially different, that is, the ASPS and AFPS of isomorphic blocks of different expression formats are not exactly the same. When encoding a heterogeneous mixed mosaic, it can meet the situation where the high-level information (ASPS and AFPS) of blocks of different expression formats in the heterogeneous mixed mosaic are not correspondingly equal. In this way, high-level parameters that are more suitable for the heterogeneous mixed mosaic are achieved, which can effectively improve the coding efficiency, that is, reduce the bit rate or improve the quality of the reconstructed multi-view video or point cloud video.
[0159] Exemplarily, when the isomorphic block is a multi-view video block, it corresponds to the first isomorphic block information; when the isomorphic block is a point cloud block, it corresponds to the second isomorphic block information; the first isomorphic block information and the second isomorphic block information include the shared syntax elements of the ASPS parameter set and the syntax elements of the AFPS parameter set; the first isomorphic block information also includes the extended syntax elements of the ASPS parameter set and the extended syntax elements of the AFPS parameter set. Here, the extended syntax elements of the ASPS parameter set and the extended syntax elements of the AFPS parameter set are newly added for the multi-view video blocks of the heterogeneous mixed splicing image, which are used to represent the ASPS parameters and AFPS parameters that are not equal to those of the point cloud blocks, so as to improve the decoding efficiency of the multi-view video blocks. It should be noted that when decoding and reconstructing the point cloud blocks and the multi-view video blocks, these ASPS parameters and AFPS parameters may have the same functions but unequal values.
[0160] Exemplarily, the second homogeneous block information includes extended syntax elements of the ASPS parameter set and the AFPS parameter set, that is, extended syntax elements of the ASPS parameter set and the AFPS parameter set may be added for the point cloud video blocks of the heterogeneous mixed mosaic to improve the decoding efficiency of the point cloud blocks. Extended syntax elements of the first homogeneous block information and the second homogeneous block information.
[0161] Exemplarily, the extended syntax elements of the ASPS parameter set include: ashm_geometry_3d_bit_depth_minus1, which is used to indicate the bit depth of the geometric coordinates of the reconstructed geometric content. ashm_geometry_2d_bit_depth_minus1 is used to indicate the bit depth of the geometry when projected onto a 2D image. ashm_log2_max_atlas_frame_order_cnt_lsb_minus4 is used to determine the variable value used for the mosaic frame order count during the decoding process. The extended syntax elements of the AFPS parameter set include: afhm_additional_lt_afoc_lsb_len, which is used to determine the value of the variable MaxLtAtlasFrmOrderCntLsbForMiv used in the reference mosaic frame list during the decoding process. The ASPS parameters and AFPS parameters of the multi-view video block represented by these syntax elements are not completely equal to the ASPS parameters and AFPS parameters of the point cloud block.
[0162] Exemplarily, when the spliced graph is a homogeneous spliced graph, the spliced graph information includes homogeneous block information, which is used to decode and reconstruct the homogeneous blocks in the spliced graph.
[0163] In some embodiments, the heterogeneous mixed mosaic graph of the embodiments of the present application includes at least one of the following: a single-attribute heterogeneous mixed mosaic graph and a multi-attribute heterogeneous mixed mosaic graph.
[0164] A single-attribute heterogeneous mixed mosaic image refers to a heterogeneous mixed mosaic image in which all homogeneous blocks have the same attribute information. For example, a single-attribute heterogeneous mixed mosaic image only includes homogeneous blocks with attribute information, such as multi-view video texture blocks and point cloud texture blocks. Another example is a single-attribute heterogeneous mixed mosaic image only includes homogeneous blocks with geometric information, such as multi-view video geometry blocks and point cloud geometry blocks.
[0165] A multi-attribute heterogeneous hybrid mosaic image refers to a heterogeneous hybrid mosaic image comprising at least two homogeneous blocks with different attribute information. For example, a multi-attribute heterogeneous hybrid mosaic image includes homogeneous blocks with both attribute information and homogeneous blocks with geometric information. As an example, a heterogeneous hybrid mosaic image can be obtained by stitching blocks with any one attribute or any two attributes of at least two of point clouds, multi-view videos, and meshes into a single image. This application does not limit this.
[0166] In some embodiments, a homogeneous block with a single attribute in a first expression format and a block with a single attribute in a second expression format are spliced to obtain a heterogeneous hybrid spliced image, wherein the first expression format and the second expression format are each any one of a multi-view video, a point cloud, and a mesh, and the first expression format and the second expression format are different, and the attribute information of the first expression format and the second expression format are the same.
[0167] The single attribute homogeneous block of the multi-view video includes at least one of a multi-view video texture block and a multi-view video geometry block.
[0168] The single attribute homogeneous block of the point cloud includes at least one of a point cloud texture block, a point cloud geometry block, and a point cloud occupancy status block.
[0169] The single-attribute homogeneous block of the mesh includes at least one of a mesh texture block and a mesh geometry block.
[0170] For example, by splicing at least two of the multi-view video geometry blocks, point cloud geometry blocks, and mesh geometry blocks into a single image, a heterogeneous mixed mosaic image is obtained. This heterogeneous mixed mosaic image is called a single-attribute heterogeneous mixed mosaic image. For another example, by splicing at least two of the multi-view video texture blocks, point cloud texture blocks, and mesh texture blocks into a single image, a heterogeneous mixed mosaic image is obtained. This heterogeneous mixed mosaic image is called a single-attribute heterogeneous mixed mosaic image.
[0171] In some embodiments, a heterogeneous hybrid mosaic is obtained by splicing multi-attribute homogeneous blocks in a first expression format and multi-attribute homogeneous blocks in a second expression format. The first expression format and the second expression format are each any one of a multi-view video, a point cloud, and a mesh, and the first expression format and the second expression format are different, and the attribute information of the first expression format and the second expression format are not completely identical.
[0172] For example, a multi-view video texture block is spliced with at least one of a point cloud geometry block and a mesh geometry block into a single image to obtain a heterogeneous mixed spliced image. For another example, a multi-view video geometry block is spliced with at least one of a point cloud texture block and a mesh texture block into a single image to obtain a heterogeneous mixed spliced image. For another example, a point cloud texture block is spliced with at least one of a multi-view video geometry block and a mesh geometry block into a single image to obtain a heterogeneous mixed spliced image. For another example, a point cloud geometry block is spliced with at least one of a multi-view video texture block and a mesh texture block into a single image to obtain a heterogeneous mixed spliced image. For another example, a point cloud geometry block, a multi-view video texture block, and a multi-view video texture block are spliced into a single image to obtain a heterogeneous mixed spliced image. For another example, point cloud geometry blocks, point cloud texture blocks, and multi-view video texture blocks are stitched together into one image to obtain a heterogeneous hybrid stitched image. Here, the obtained heterogeneous hybrid stitched image is called a multi-attribute heterogeneous hybrid stitched image.
[0173] In some embodiments, the isomorphic mosaic graph of the present application includes at least one of the following: a single-attribute isomorphic mosaic graph and a multi-attribute isomorphic mosaic graph. In some embodiments, the isomorphic mosaic graph is obtained by splicing isomorphic blocks of a first attribute in a first expression format. Alternatively, the isomorphic mosaic graph is obtained by splicing isomorphic blocks of a first attribute and isomorphic blocks of a second attribute in a first expression format.
[0174] A single-attribute homogeneous mosaic is a homogeneous mosaic in which all homogeneous blocks have the same expression format and attribute information. For example, a single-attribute homogeneous mosaic includes only homogeneous blocks with attribute information in a certain expression format, such as a single-attribute homogeneous mosaic that includes only multi-view video texture blocks or only point cloud texture blocks. Another example is a single-attribute homogeneous mosaic that includes only homogeneous blocks with geometric information, such as only multi-view video geometry blocks or only point cloud geometry blocks.
[0175] A multi-attribute isomorphic mosaic is an isomorphic mosaic that includes at least two isomorphic blocks with the same expression format but different attribute information. For example, a multi-attribute isomorphic mosaic includes both isomorphic blocks with attribute information and isomorphic blocks with geometric information. As an example, a multi-attribute isomorphic mosaic includes a multi-view video texture block and a multi-view video collection block. For another example, a multi-attribute isomorphic mosaic includes a point cloud geometry block and a point cloud texture block. As shown in Figure 8, a multi-attribute isomorphic mosaic includes point cloud texture block 1, point cloud geometry block 1, and point cloud geometry block 2.
[0176] In some embodiments, the splicing graph information may further include syntax elements, and the splicing graph is determined to be a single-attribute heterogeneous mixed splicing graph, a multi-attribute heterogeneous mixed splicing graph, a single-attribute homogeneous splicing graph, or a multi-attribute homogeneous splicing graph based on the syntax elements.
[0177] Step 603: Encode the splicing graph and the splicing graph information to obtain a bitstream.
[0178] In some embodiments, the parameter set substream of the bitstream includes a third syntax element, and the bitstream corresponding to the visual media content including at least one expression format in the bitstream is determined based on the third syntax element. Exemplarily, the parameter set of the bitstream is V3C_VPS, and the third syntax element may be ptl_profile_toolset_idc in the V3C_VPS.
[0179] In some embodiments, the codestream corresponding to the visual media content in at least one expression format included in the codestream is indicated by setting the third syntax element to different values. Exemplarily, determining the codestream corresponding to the visual media content in at least one expression format included in the codestream based on the third syntax element includes: the third syntax element being a first value, determining that the codestream includes both the codestream corresponding to the visual media content in the first expression format and the codestream corresponding to the visual media content in the second expression format; the third syntax element being a second value, determining that the codestream includes the visual media content in the first expression format; and the third syntax element being a third value, determining that the codestream includes the visual media content in the second expression format.
[0180] Exemplarily, taking the case where the first expression format is multi-view video and the second expression format is point cloud, when the third syntax element is set to the first value, the first value is used to indicate that the code stream contains both multi-view video code stream and point cloud code stream. As a specific example, when ptl_profile_toolset_idc=X, X is 128 / 129 / 130 / 131 / 132 / 133, it means that the current code stream contains both point cloud and multi-view code streams. For another example, when the third syntax element is set to the second value, the second value is used to indicate that the code stream only contains point cloud code stream. As a specific example, when ptl_profile_toolset_idc=X, X is 0 / 1, it means that the current code stream only contains point cloud code stream. For another example, when the third syntax element is set to the third value, the third value is used to indicate that the code stream only contains multi-view video code stream. As a specific example, when ptl_profile_toolset_idc=X, X is 64 / 65 / 66, which means that the current stream only contains multi-view video streams. It should be understood that the values of the first value, the second value, and the third value are only examples, and the embodiments of the present application are not limited thereto.
[0181] In some embodiments, the bitstream includes a video compression substream and a mosaic image information substream. Encoding the mosaic image and mosaic image information to obtain the bitstream includes: encoding the mosaic image to obtain a video compression substream; encoding the mosaic image information of the mosaic image to obtain a mosaic image information substream; and combining the video compression substream and the mosaic image information substream into the bitstream. This allows support for heterogeneous source formats such as video, point cloud, and mesh in the same compressed bitstream, enabling the simultaneous presence of multi-view video mosaics, point cloud video mosaics, mesh mosaics, and heterogeneous mixed mosaics in the compressed bitstream. This reduces the number of required two-dimensional video encoders, such as HEVC, VVC, AVC, and AVS, lowering implementation costs and improving usability.
[0182] In some embodiments, encoding the mosaic image and mosaic image information to obtain a code stream includes: if the expression format of the i-th block is a first expression format, determining that the sub-block in the i-th block is encoded using the encoding standard corresponding to the first expression format, and obtaining a code stream corresponding to the visual media content in the first expression format; if the expression format of the i-th block is a second expression format, determining that the sub-block in the i-th block is encoded using the encoding standard corresponding to the second expression format, and obtaining a code stream corresponding to the visual media content in the second expression format.
[0183] Exemplarily, if the second syntax element of the i-th block is known to be 1, it is determined that the current sub-block is encoded using the multi-view video coding standard. If the second syntax element of the i-th block is known to be 0, it is determined that the current sub-block is encoded using the point cloud coding standard.
[0184] In the embodiment of the present application, the video encoder used to perform video encoding on the heterogeneous mixed mosaic and the homogeneous mosaic to obtain the video compression substream can be the video encoder shown in Figure 2A above. That is, in the embodiment of the present application, the heterogeneous mixed mosaic or the homogeneous mosaic is treated as a frame of image, firstly block-divided, then intra-frame or inter-frame prediction is used to obtain predicted values for the coding blocks, the predicted values of the coding blocks are subtracted from the original values to obtain residual values, and the residual values are transformed and quantized to obtain the video compression substream.
[0185] In this embodiment of the present application, while generating at least one mosaic, mosaic information corresponding to each mosaic is generated. The mosaic information is encoded to obtain a mosaic information substream. The mosaic information includes a first syntax element indicating the mosaic type and a second syntax element representing the format of each isomorphic block in the mosaic. This embodiment of the present application does not limit the encoding method for the mosaic information; for example, conventional data compression encoding methods such as constant-length encoding or variable-length encoding may be used for compression.
[0186] Finally, the video compression sub-stream and the mosaic information sub-stream are written into the same stream to obtain the final stream. In other words, the embodiment of the present application not only supports heterogeneous source formats such as video, point cloud, and mesh, but also homogeneous source formats in the same compressed stream.
[0187] In some embodiments, the method further includes: encoding the parameter set of the code stream to obtain a code stream parameter set sub-code stream. Specifically, the encoding end combines the video compression sub-code stream, the splicing graph information sub-code stream and the parameter set sub-code stream into a code stream. The parameter set sub-code stream of the code stream includes a third syntax element, and the code stream corresponding to the visual media content in at least one expression format is determined according to the third syntax element. That is to say, the encoding end sends the third syntax element to indicate whether the code stream contains visual media content in at least two expression formats at the same time. Exemplarily, when the third syntax element indicates that the code stream includes a code stream corresponding to visual media content in one expression format, it can be understood that the encoding end processes the visual media content in one expression format to obtain a homogeneous block, and splices the homogeneous blocks to obtain a homogeneous splicing graph. When the third syntax element indicates that the code stream includes code streams corresponding to visual media contents in at least two expression formats, it can be understood that the encoding end obtains at least two isomorphic blocks for the visual media contents in at least two expression formats, and splices the at least two isomorphic blocks to obtain a homogeneous splicing graph and / or a heterogeneous mixed splicing graph.
[0188] Exemplarily, when the third syntax element indicates that the codestream includes codestreams corresponding to visual media content in at least two expression formats, the method includes: performing isomorphic splicing on isomorphic blocks of the first expression format to obtain a first isomorphic splicing graph, and performing isomorphic splicing on isomorphic blocks of the second expression format to obtain a second isomorphic splicing graph; or, performing heterogeneous splicing on isomorphic blocks of the first expression format and isomorphic blocks of the second expression format to obtain a heterogeneous mixed splicing graph; or, performing isomorphic splicing on isomorphic blocks of the first expression format to obtain a first isomorphic splicing graph, and performing heterogeneous splicing on isomorphic blocks of the first expression format and isomorphic blocks of the second expression format to obtain a heterogeneous mixed splicing graph; or, performing isomorphic splicing on isomorphic blocks of the second expression format to obtain a second isomorphic splicing graph, and performing heterogeneous splicing on isomorphic blocks of the first expression format and isomorphic blocks of the second expression format to obtain a heterogeneous mixed splicing graph.
[0189] In an embodiment of the present application, in order to reduce the number of encoders and reduce the encoding cost, during encoding, the visual media content is first processed separately (i.e., packaged) to obtain multiple isomorphic blocks. Then, at least two isomorphic blocks with different expression formats are spliced into a heterogeneous mixed splicing graph, and at least one isomorphic block with exactly the same expression format is spliced into an isomorphic splicing graph. The heterogeneous mixed splicing graph and the isomorphic splicing graph are encoded to obtain a video compression sub-stream, and the splicing graph information is encoded to obtain a splicing information sub-stream; the video compression stream and the splicing information stream are synthesized into a compressed stream. This makes the encoding method suitable for application scenarios of visual media content in a variety of expression formats, expands the scope of application, and by mixing and encoding data in different expression formats, the video encoder can be called only once for encoding during encoding, thereby reducing the number of two-dimensional video encoders such as HEVC, VVC, AVC, AVS that need to be called, reducing the encoding cost and improving ease of use. Moreover, when encoding heterogeneous mixed mosaic images, some high-level parameters of blocks with different expression formats can be unequal, which can retain more effective information of blocks with different expression formats, improve the synthesis quality of the image, and improve the overall efficiency of bit rate-quality.
[0190] The above describes the encoding method of the present application using the encoding end as an example. The following describes the video decoding method provided in the embodiment of the present application using the decoding end as an example.
[0191] FIG9 is a schematic flow chart of a decoding method provided in an embodiment of the present application. As shown in FIG9 , the decoding method in an embodiment of the present application includes:
[0192] Step 901: Decode the code stream to obtain a splicing graph and splicing graph information;
[0193] Exemplarily, the bitstream includes a video compression substream and a mosaic graph information substream, and decoding the bitstream to obtain the mosaic graph and the mosaic graph information includes: extracting the mosaic graph information substream and the video compression substream respectively; decoding the video compression substream to obtain the mosaic graph; and decoding the mosaic graph information substream to obtain the mosaic graph information.
[0194] Exemplarily, the video compression sub-stream is decoded to obtain a heterogeneous mixed mosaic map, a multi-view mosaic map and a point cloud mosaic map; and the mosaic map information sub-stream is decoded to obtain heterogeneous mixed mosaic map information, multi-view mosaic map information and point cloud mosaic map information.
[0195] Exemplarily, in some embodiments, the bitstream further includes a parameter set sub-bitstream; the parameter set sub-bitstream includes a third syntax element; and the bitstream corresponding to the visual media content in at least one expression format is determined in the bitstream based on the third syntax element. That is, during the decoding process, the bitstream is first determined based on the third syntax element to determine the bitstream corresponding to the visual media content in several expression formats contained in the bitstream. When the bitstream is determined to contain visual media content in one expression format based on the third syntax element of the V3C bitstream layer, it is determined that all the splices are isomorphic splices; when the bitstream is determined to contain visual media content in two expression formats based on the third syntax element, it is determined that the splice may contain heterogeneous mixed splices, and it is necessary to further determine whether the splice is an isomorphic splice or a heterogeneous mixed splice. Furthermore, the type of the splice is determined based on the first syntax element, and when the splice is determined to be a heterogeneous mixed splice, the type of the isomorphic block is determined based on the second syntax element.
[0196] In some embodiments, the codestream corresponding to the visual media content in at least one expression format included in the codestream is indicated by setting the third syntax element to different values. Exemplarily, determining the codestream corresponding to the visual media content in at least one expression format included in the codestream based on the third syntax element includes: the third syntax element being a first value, determining that the codestream includes both the codestream corresponding to the visual media content in the first expression format and the codestream corresponding to the visual media content in the second expression format; the third syntax element being a second value, determining that the codestream includes the visual media content in the first expression format; and the third syntax element being a third value, determining that the codestream includes the visual media content in the second expression format.
[0197] Exemplarily, taking the case where the first expression format is multi-view video and the second expression format is point cloud, when the third syntax element is set to the first value, the first value is used to indicate that the code stream contains both multi-view video code stream and point cloud code stream. As a specific example, when ptl_profile_toolset_idc=X, X is 128 / 129 / 130 / 131 / 132 / 133, it means that the current code stream contains both point cloud and multi-view code streams. For another example, when the third syntax element is set to the second value, the second value is used to indicate that the code stream only contains point cloud code stream. As a specific example, when ptl_profile_toolset_idc=X, X is 0 / 1, it means that the current code stream only contains point cloud code stream. For another example, when the third syntax element is set to the third value, the third value is used to indicate that the code stream only contains multi-view video code stream. As a specific example, when ptl_profile_toolset_idc=X, X is 64 / 65 / 66, which means that the current stream only contains multi-view video streams. It should be understood that the values of the first value, the second value, and the third value are only examples, and the embodiments of the present application are not limited thereto.
[0198] Step 902: When the mosaic is a heterogeneous mixed mosaic, obtaining at least two types of homogeneous blocks and homogeneous block information based on the mosaic and the mosaic information; wherein different types of homogeneous blocks correspond to different visual media content expression formats and different homogeneous block information;
[0199] The mosaic information is used to reconstruct the mosaic. Exemplarily, the mosaic information includes at least mosaic type information, mosaic information of homogeneous blocks, and homogeneous block information. In some embodiments, the mosaic information includes a first syntax element, and whether the mosaic is a heterogeneous mixed mosaic or a homogeneous mosaic is determined based on the first syntax element. In some embodiments, the first syntax element is a syntax element of an ASPS mosaic sequence parameter set and a syntax element of an AFPS mosaic frame parameter set. The ASPS and AFPS are parsed to determine the mosaic type.
[0200] Exemplarily, when the mosaic is a heterogeneous mixed mosaic, the mosaic is split to obtain at least two isomorphic blocks; and based on the expression formats of the at least two isomorphic blocks, isomorphic block information corresponding to the at least two isomorphic blocks is obtained from the mosaic information. Exemplarily, the heterogeneous mixed mosaic is split based on the heterogeneous mixed mosaic information, and reconstructed multi-view video isomorphic blocks and isomorphic block information, as well as reconstructed point cloud isomorphic blocks and isomorphic block information, are output.
[0201] In some embodiments, the first syntax element includes a first sub-syntax element and a second sub-syntax element; determining whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph based on the first syntax element includes: if the value of the first sub-syntax element and the value of the second sub-syntax element are equal, determining whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph based on the values.
[0202] In other words, the codestream must ensure absolute consistency between the first and second sub-syntax elements. The mosaic type can only be determined when the two sub-syntax elements are consistent. For example, the two sub-syntax elements can be compared for consistency before the mosaic type is determined based on the value of one of the sub-syntax elements. Alternatively, the mosaic type can be determined based on each sub-syntax element and absolute consistency can be ensured by comparing the mosaic types to ensure consistency.
[0203] Exemplarily, determining whether the mosaic is a heterogeneous mixed mosaic or a homogeneous mosaic based on the values includes: if the value is a first preset value, determining the mosaic is a heterogeneous mixed mosaic; and if the value is a second preset value, determining the mosaic is a homogeneous mosaic. In other words, two values or two types of values can be set to identify heterogeneous mixed mosaics and homogeneous mosaics. Exemplarily, the first preset value is 1 and the second preset value is 0.
[0204] Exemplarily, in some embodiments, the splice information does not include the first syntax element, and the splice is determined to be a homogeneous splice. In some embodiments, the splice information does not include the first syntax element, and the value of the first syntax element is inferred to be a second preset value. Exemplarily, the splice information does not include the first sub-syntax element, and the splice is determined to be a homogeneous splice, and the value of the first sub-syntax element is inferred to be a second preset value; the splice information does not include the second sub-syntax element, and the splice is determined to be a homogeneous splice, and the value of the second sub-syntax element is inferred to be a second preset value.
[0205] Exemplarily, the determining whether the mosaic is a heterogeneous mixed mosaic or a homogeneous mosaic based on the values includes: if the value is a third preset value, the mosaic is determined to be a heterogeneous mixed mosaic including homogeneous blocks of a first expression format and a second expression format, wherein the first expression format and the second expression format are different expression formats; if the value is a fourth preset value, the mosaic is determined to be a homogeneous mosaic including homogeneous blocks of the first expression format; if the value is a fifth preset value, the mosaic is determined to be a homogeneous mosaic including homogeneous blocks of the second expression format. That is, multiple values can also be set to identify the expression formats of heterogeneous mixed mosaics and homogeneous mosaics, and even to identify which homogeneous blocks of which expression formats are included in the heterogeneous mixed mosaic. Exemplarily, the third preset value is 2, the fourth preset value is 1, and the fifth preset value is 0.
[0206] Exemplarily, the first sub-syntax element is a syntax element of an ASPS (Assembled Picture Sequence Parameter Set), and the second sub-syntax element is a syntax element of an AFPS (Assembled Picture Frame Parameter Set). In this embodiment of the present application, asps_heterogeneous_miv_extension_present_flag represents the first sub-syntax element, and afps_heterogeneous_miv_extension_present_flag represents the second sub-syntax element.
[0207] Exemplarily, in some embodiments, when the mosaic is a heterogeneous mixed mosaic, the mosaic information further includes a second syntax element; the expression format of the homogeneous blocks in the mosaic is determined based on the second syntax element. It is understood that after determining that the mosaic is a heterogeneous mixed mosaic based on the first syntax element, the second syntax element of the homogeneous blocks is further parsed to determine the homogeneous block type. Exemplarily, the second syntax element is a syntax element of the mosaic frame parameter set (AFPS).
[0208] Specifically, the expression format type corresponding to the i-th block in the spliced graph can be indicated by setting different values for the second syntax element. Exemplarily, the expression format of the isomorphic blocks in the spliced graph determined according to the second syntax element includes: when the value of the second syntax element of the i-th block is the sixth preset value, the expression format of the i-th block is determined to be the first expression format; when the value of the second syntax element of the i-th block is the seventh preset value, the expression format of the i-th block is determined to be the second expression format. Take the first expression format as point cloud and the second expression format as multi-view video as an example. Optionally, the sixth preset value is 0 and the seventh preset value is 1.
[0209] Furthermore, the expression format of the i-th block is a first expression format, and the i-th block is decoded using a decoding method corresponding to the first expression format; the expression format of the i-th block is a second expression format, and the i-th block is decoded using a decoding method corresponding to the second expression format.
[0210] For example, in some embodiments, when the mosaic is a heterogeneous mixed mosaic, the mosaic information includes at least two types of homogeneous block information, where homogeneous blocks in different expression formats correspond to different homogeneous block information. The homogeneous block information includes reconstruction information of the homogeneous blocks and other supplementary information for decoding and reconstructing the homogeneous blocks.
[0211] Exemplarily, the isomorphic block information includes syntax elements of ASPS and syntax elements of AFPS; different isomorphic block information corresponds to different syntax elements of the ASPS and syntax elements of the AFPS. In some embodiments, for a heterogeneous mixed mosaic, the ASPS and AFPS of isomorphic blocks of different expression formats are at least partially different, that is, the ASPS and AFPS of isomorphic blocks of different expression formats are not exactly the same. When encoding a heterogeneous mixed mosaic, it can meet the situation where the high-level information (ASPS and AFPS) of blocks of different expression formats in the heterogeneous mixed mosaic are not correspondingly equal. In this way, high-level parameters that are more suitable for the heterogeneous mixed mosaic are achieved, which can effectively improve the coding efficiency, that is, reduce the bit rate or improve the quality of the reconstructed multi-view video or point cloud video.
[0212] Exemplarily, when the isomorphic block is a multi-view video block, it corresponds to the first isomorphic block information; when the isomorphic block is a point cloud block, it corresponds to the second isomorphic block information; the first isomorphic block information and the second isomorphic block information include the shared syntax elements of the ASPS parameter set and the syntax elements of the AFPS parameter set; the first isomorphic block information also includes the extended syntax elements of the ASPS parameter set and the extended syntax elements of the AFPS parameter set.
[0213] In some embodiments, the expression format is multi-view video, point cloud or grid. One isomorphic block corresponds to one expression format. Different isomorphic blocks correspond to different expression formats. Exemplarily, the expression formats corresponding to at least two isomorphic blocks include at least two of the following: multi-view video, point cloud, grid. It should be noted that in the embodiment of the present application, each isomorphic block may include at least one isomorphic block with the same expression format. Exemplarily, the isomorphic block in the point cloud format includes one or more point cloud blocks, the isomorphic area in the multi-view video format includes one or more multi-view video blocks, and the isomorphic block in the grid format includes one or more grid blocks.
[0214] Step 903: when the mosaic is a homogeneous mosaic, obtaining a homogeneous block and homogeneous block information according to the mosaic and the mosaic information;
[0215] Exemplarily, when the mosaic is a homogeneous mosaic, the mosaic is split to obtain homogeneous blocks; and homogeneous block information is obtained from the mosaic information. Exemplarily, the homogeneous mosaic is split based on the homogeneous mosaic information of the multi-view video, and reconstructed multi-view video homogeneous blocks and homogeneous block information are output. Exemplarily, the homogeneous mosaic is split based on the homogeneous mosaic information of the point cloud, and reconstructed point cloud homogeneous blocks and homogeneous block information are output.
[0216] Exemplarily, when the spliced graph is a homogeneous spliced graph, the spliced graph information includes homogeneous block information, which is used to decode and reconstruct the homogeneous blocks in the spliced graph.
[0217] In some embodiments, the heterogeneous mixed mosaic graph of the embodiments of the present application includes at least one of the following: a single-attribute heterogeneous mixed mosaic graph and a multi-attribute heterogeneous mixed mosaic graph.
[0218] Step 904: Obtain visual media content in at least two expression formats according to the isomorphic blocks and the isomorphic block information.
[0219] Exemplarily, the method of obtaining visual media content in at least two expression formats based on the isomorphic blocks and the isomorphic block information includes: if the expression format of the i-th block is a first expression format, determining that the sub-block in the i-th block is decoded and reconstructed using a decoding method corresponding to the first expression format to obtain visual media content in the first expression format; if the expression format of the i-th block is a second expression format, determining that the sub-block in the i-th block is decoded and reconstructed using a decoding method corresponding to the second expression format to obtain visual media content in the second expression format.
[0220] By combining homogeneous blocks of different representation formats into a heterogeneous hybrid mosaic for encoding, this technical solution reduces the number of required 2D video codecs, such as HEVC, VVC, AVC, and AVS, lowering implementation costs and improving usability. Furthermore, when decoding the heterogeneous hybrid mosaic, certain high-level parameters of blocks of different representation formats can be unequal, preserving more effective information from these blocks, improving image synthesis quality and overall bitrate-quality efficiency.
[0221] The decoding method provided in the embodiments of the present application is further illustrated below with examples.
[0222] Figure 10 is a schematic diagram of the V3C bitstream structure provided by an embodiment of the present application. The V3C parameter set (V3C_parameter_set()) of the V3C_VPS may include a third syntax element (ptl_profile_toolset_idc). A value of ptl_profile_toolset_idc of 128 to 133 indicates that the current bitstream contains both a point cloud bitstream (such as VPCC Basic or VPCC Extended) and a multi-view video bitstream (such as MIV Main, MIV Extended, or MIV Geometry Absent).
[0223] The ASPS parameter set may include a first sub-syntax element (asps_heterogeneous_miv_extension_present_flag). When ptl_profile_toolset_idc is 128 to 133, the current splice type is determined according to asps_heterogeneous_miv_extension_present_flag.
[0224] The AFPS parameter set may include a second sub-syntax element (afps_heterogeneous_miv_extension_present_flag). When ptl_profile_toolset_idc is 128 to 133, the current mosaic type is determined based on afps_heterogeneous_miv_extension_present_flag. The AFPS parameter set also includes a second syntax element (afps_heterogeneous_frame_tile_toolset_miv_present_flag) for determining the slice type, thereby ensuring that during parsing and decoding, the current slice should belong to multi-view or point cloud.
[0225] The code stream must ensure absolute consistency between afps_heterogeneous_miv_extension_present_flag and asps_heterogeneous_miv_extension_present_flag.
[0226] Decoding Case 1
[0227] 1. During decoding, the VPS is parsed from the V3C code stream, and the ptl_profile_toolset_idc (third syntax element) is parsed from the VPS. If ptl_profile_toolset_idc=0 / 1, it means that only the point cloud code stream exists in the current code stream;
[0228] 2. Implement the point cloud decoding standard for the current code stream.
[0229] Decoding Case 2
[0230] 1. During decoding, the VPS is parsed from the V3C stream, and ptl_profile_toolset_idc is parsed from the VPS. If ptl_profile_toolset_idc = 64 / 65 / 66, it means that only multi-view stream exists in the current stream.
[0231] 2. Implement multiple viewpoint decoding standards for the current code stream.
[0232] Decoding Case 3
[0233] 1. During parsing, the VPS is obtained from the V3C code stream, and ptl_profile_toolset_idc is obtained from the VPS. If ptl_profile_toolset_idc=128-133, it means that the current code stream contains both point cloud and multi-view code streams.
[0234] 2. Analyze high-level syntax such as ASPS and AFPS for each spliced image:
[0235] a) Parse asps_heterogeneous_miv_extension_present_flag (first sub-syntax element) in ASPS:
[0236] i. When asps_heterogeneous_miv_extension_present_flag is not present (i.e., asps_extension_6bits = 0), the current mosaic is homogeneous content (all strips are point cloud or multi-view). ii. When asps_heterogeneous_miv_extension_present_flag is present and equal to 0, the current mosaic is homogeneous content (all strips are point cloud or multi-view). iii. When asps_heterogeneous_miv_extension_present_flag is present and equal to 1, the current mosaic is heterogeneous content, with both point cloud and multi-view strips. Therefore, the current mosaic's ASPS auxiliary high-layer information is split into two sub-information sets: one sub-set for decoding multi-view strips, and the other sub-set for decoding point cloud strips. The auxiliary information required for point cloud strips can be obtained by parsing part 8 of standard 23090-5; the auxiliary information required for multi-view strips can be obtained by parsing part 8 of standard 23090-5 and the newly added asps_heterogeneous_miv_extension and part 8 of standard 23090-12.
[0237] b) Parse afps_heterogeneous_miv_extension_present_flag (second sub-syntax element) in AFPS:
[0238] i. When afps_heterogeneous_miv_extension_present_flag is not present, that is, when afps_extension_7bits = 0, it indicates that the current mosaic is homogeneous content (all strips are point cloud type or multi-view type); ii. When afps_heterogeneous_miv_extension_present_flag is present and equal to 0, it indicates that the current mosaic is homogeneous content (all strips are point cloud type or multi-view type); iii. When afps_heterogeneous_miv_extension_present_flag is present and equal to 1, it indicates that the current mosaic is heterogeneous content, with both point cloud strips and multi-view strips. Therefore, the AFPS auxiliary high-layer information of the current mosaic is split into two sub-information sets, namely, one sub-set is used for decoding multi-view strips, and the other sub-set is used for decoding point cloud strips. The auxiliary information required for point cloud strips can be obtained by parsing part 8 of standard 23090-5; the auxiliary information required for multi-view strips can be obtained by parsing part 8 of standard 23090-5 and the newly added afps_heterogeneous_miv_extension and part 8 of standard 23090-12.
[0239] c) Parse afps_heterogeneous_frame_tile_toolset_miv_present_flag (second syntax element) in AFPS:
[0240] i. Determine if afps_heterogeneous_frame_tile_toolset_miv_present_flag does not exist, indicating that all strips of the current mosaic are of the same type; ii. Traverse all tiles and determine if afps_heterogeneous_frame_tile_toolset_miv_present_flag[i] of the i-th strip is 0, indicating that the current strip is a point cloud strip; determine if afps_heterogeneous_frame_tile_toolset_miv_present_flag[i] of the i-th strip is 1, indicating that the current strip is a multi-view strip.
[0241] Among them, the code stream must ensure the absolute consistency of afps_heterogeneous_miv_extension_present_flag and asps_heterogeneous_miv_extension_present_flag.
[0242] 3. Then parse the patch_data_unit information of each sub-block in each strip. Under the premise that it is known that the current strip adopts multi-view auxiliary information, determine whether the current sub-block adopts the multi-view video decoding standard; under the premise that it is known that the current strip adopts point cloud auxiliary information, determine whether the current sub-block adopts the point cloud video decoding standard.
[0243] The presence of a heterogeneous hybrid mosaic is indicated by the ptl_profile_toolset_idc number. The asps_heterogeneous_miv_extension_present_flag and afps_heterogeneous_miv_extension_present_flag fields are added to determine whether each mosaic should be classified as point cloud, multi-view, or point cloud + multi-view. To maintain compatibility with previous standards, the following new syntax and semantics, as well as constraints on legacy semantics, have been implemented. Table 9-1-1 and Table 9-1-2 respectively describe the syntax restrictions for toolbox-level components for multi-view and heterogeneous data in the integrated codestream.
[0244] Table 9-1-1 and Table 9-1-2 respectively show the restrictions on the relevant syntax of toolbox-level components for multi-viewpoints under integrated code streams and the restrictions on the relevant syntax of toolbox-level components for heterogeneous data.
[0245] In the embodiment of the present application, at the high-level part (mosaic level and mosaic sequence level), the type of each slice in the current mosaic is described through the newly added syntax element afps_heterogeneous_frame_tile_toolset_miv_present_flag, thereby ensuring that during parsing and decoding, it is realized whether the current slice belongs to multi-view or point cloud.
[0246] This scheme can ensure that no matter in multi-view parsing or point cloud parsing, there is only one usable mosaic level parameter (AFPS) and mosaic sequence level parameter (ASPS), and it can achieve that the AFPS and ASPS of multi-view are not completely equal to the AFPS and ASPS of point cloud.
[0247] Table 1 shows an example of available toolset profile components. Table 1 provides a list of toolset profile components defined for V3C and their corresponding identification syntax element values, such as ptl_profile_toolset_idc and ptc_one_v3c_frame_only_flag, which can be used only in this document. The syntax element ptl_profile_toolset_idc provides the main definition of the toolset profile. Additional syntax elements such as ptc_one_v3c_frame_only_flag can specify additional features or restrictions of the defined profile. ptc_one_v3c_frame_only_flag can be used to support only a single V3C frame. It should be noted that 2..63, 67..127, 134..255 in ptl_profile_toolset_idc are reserved and temporarily undefined. Standards organizations may make further provisions in future standards. The profile types defined in Table 1 can include dynamic or static.
[0248] Table 1 Available toolset configuration file components
[0249]
[0250]
[0251]
[0252] Table 2 shows the RBSP syntax for the General Atlas Sequence Parameter Set (GAS), which can be used in ISO / IEC 23090-5. The extended syntax element asps_heterogeneous_miv_extension_present_flag in the GAS Sequence Parameter Set indicates the type of mosaic. Specifically, the value of this syntax element determines whether the mosaic should be classified as point cloud, multi-view, or point cloud + multi-view.
[0253] Table 2 RBSP syntax of the general mosaic sequence parameter set
[0254]
[0255]
[0256] Table 3 shows the ASPS heterogeneous multiview extension syntax elements (Atlas sequence parameter set heterogeneous MIV extension syntax), which can be used by ISO / IEC 23090-5. ashm_geometry_3d_bit_depth_minus1 is used to indicate the bit depth of the geometry coordinates used to reconstruct the geometry content. ashm_geometry_2d_bit_depth_minus1 is used to indicate the bit depth of the geometry when projected onto a 2D image. ashm_log2_max_atlas_frame_order_cnt_lsb_minus4 is used to determine the value of the variable used for the mosaic frame order count during decoding.
[0257] Table 3 ASPS heterogeneous multi-view extended syntax elements
[0258]
[0259] Table 4 shows the RBSP syntax for the General Atlas Frame Parameter Set (GAP), which is used by ISO / IEC 23090-5. The extended syntax element afps_heterogeneous_miv_extension_present_flag in the GAP is used to indicate the mosaic type. Specifically, the value of this syntax element determines whether the mosaic should be classified as point cloud, multi-view, or point cloud + multi-view.
[0260] Table 4 Syntax of mosaic frame parameter set RBSP
[0261]
[0262]
[0263] Table 5 shows the AFPS heterogeneous MIV extension syntax elements (Atlas frame parameter set heterogeneous MIV extension syntax), which can be used by ISO / IEC 23090-5. afhm_additional_lt_afoc_lsb_len is used to determine the value of the variable MaxLtAtlasFrmOrderCntLsbForMiv used in the decoding process of the reference splicing frame list.
[0264] Table 5 shows the AFPS heterogeneous MIV extended syntax elements
[0265] afps_heterogeneous_miv_extension(){Descriptorafhm_additional_lt_afoc_lsb_lenue(v)}
[0266] The semantics of the ASPS syntax elements and the semantics of the AFPS syntax elements are explained below.
[0267] 1. Semantics of ASPS syntax elements:
[0268] asps_extension_6bits equal to 0 indicates that the asps_extension_data_flag is not present in the ASPS RBSP syntax structure. If present, the value of asps_extension_6bits is 0 or 1 in this standard. Values other than 0 and 1 are reserved for future use by ISO / IEC. Decoders should allow values of asps_extension_6bits other than 0 or 1 and should ignore all asps_extension_data_flag syntax elements in asps. When not present, the value of asps_extension_6bits is inferred to be equal to 0.
[0269] asps_heterogeneous_miv_extension_present_flag is 1 if the asps_heterogeneous_miv_extension() syntax structure is present in the syntax structure. asps_heterogeneous_miv_extension_present_flag is 0 if the syntax structure is not present. When asps_heterogeneous_miv_extension_present_flag is not present, the value of asps_heterogeneous_miv_extension_present_flag is inferred to be 0.
[0270] ASPS extended syntax element semantics: ashm_geometry_3d_bit_depth_minus1 plus 1 indicates the bit depth of the geometric coordinates of the reconstructed volume content. ashm_geometry3d_bitdepth_minus1 should be between 0 and 31, inclusive.
[0271] ashm_geometry_2d_bit_depth_minus1 plus 1 represents the bit depth of the geometry when projected to a 2d image. ashm_geometry2d_bit_depth_minus1 should be in the range of 0 to 31, inclusive.
[0272] ashm_log2_max_atlas_frame_order_cnt_lsb_minus4 plus 4 specifies the value of the variables Log2MaxAtlasFrmOrderCntLsbForMiv and MaxAtlasFlmOrderCNTLsbForMIv used for the mosaic frame order count during decoding, as shown below:
[0273] Log2MaxAtlasFrmOrderCntLsbForMiv=ashm_log2_max_atlas_frame_order_cnt_lsb_minus4+4
[0274] MaxAtlasFrmOrderCntLsbForMiv=2Log2MaxAtlasFrmOrderCntLsbForMiv
[0275] Among them, ashm_log2_max_atlas_frame_order_cnt_lsb_minus4 takes a value from 0 to 12, including 0 and 31.
[0276] 2. Semantics of AFPS syntax elements:
[0277] afps_extension_7bits equal to 0 specifies that the afps_extension_data_flag syntax element is not present in the AFPS RBSP syntax structure. If present, afps_extension_7bits shall be equal to a value of 0 or 1 in the codestream conforming to this version of this document. Values of afps_extension_7bits other than 0 and 1 are reserved by ISO / IEC for future use. Decoders shall allow values of afps_extension_7bits other than 0 or 1 and shall ignore the afps_extension_data_flag syntax element in the AFPS. When afps_extension_7bits is not present, the value of afps_extension_7bits is inferred to be equal to 0.
[0278] afps_heterogeneous_miv_extension_present_flag is equal to 1 to indicate the presence of the afps_hyterogeneous_miv_extension() syntax structure in the AFPS syntax structure. afps_heterogeneous_miv_extension_present_flag is equal to 0 to indicate the absence of this syntax structure specified by afps_heterogeneous_miv_extension_present_flag. When afps_heterogeneous_miv_extension_present_flag is not present, the value of afps_heterogeneous_miv_extension_present_flag is inferred to be equal to 0.
[0279] For bitstreams conforming to this version of this document, afps_heterogeneous_miv_extension_present_flag and asps_heterigeneous_miv_extenson_present_flag shall be consistent, that is, they shall both be present and have the same value.
[0280] afps_heterogeneous_frame_tile_toolset_miv_present_flag[i] is equal to 1 to indicate that the i-th slice in the heterogeneous mixed mosaic is a mosaic slice belonging to miv (i.e., a multi-view video slice). afps_heterogeneous_frame_tile_toolset_miv_present_flag[i] is equal to 0 to specify that the i-th tile in the heterogeneous mixed mosaic is a mosaic slice belonging to vpcc (i.e., a point cloud slice). When afps_heterogeneous_frame_tile_toolset_miv_present_flag[i] is not present, the value of afps_heterogeneous_frame_tile_toolset_miv_present_flag[i] is inferred to be equal to 0.
[0281] AFPS extended syntax element semantics:
[0282] afhm_additional_lt_afoc_lsb_len represents the value of the variable MaxLtAtlasFrmOrderCntLsbForMiv used in the reference mosaic frame list decoding process, as shown below:
[0283] MaxLtAtlasFrmOrderCntLsbForMiv=2*(Log2MaxAtlasFrmOrderCntLsbForMiv+afhm_additional_lt_afoc_lsb_len)
[0284] The value of afhm_additional_lt_afoc_lsb_len shall be between 0 and 32–Log2MaxAtlasFrmOrderCntLsbForMiv, inclusive.
[0285] When asps_long_term_ref_atlas_frames_flag is equal to 0, the value of afhm_additional_lt_afoc_lsb_len is inferred to be 0.
[0286] Semantics of mosaic strip data unit header,
[0287] ath_atlas_frm_order_cnt_lsb indicates that the current mosaic slice specifies the mosaic frame order count modulo MaxAtlasFrmOrderCntLsb. If the afps_heterogeneous_frame_tile_toolset_miv_present_flag for the current mosaic slice is equal to 0, the length of the ath_atlas_frm_order_cnt_lsb syntax element is equal to Log2MaxAtlasFrmOrderCntLsb bits. The value of ath_atlas_frm_order_cnt_lsb shall be in the range of 0 to MaxAtlasFrmOrderCntLsb-1, inclusive. If the afps_heterogeneous_frame_tile_toolset_miv_present_flag for the current mosaic slice is equal to 1, the length of the ath_atlas_frm_order_cnt_lsb syntax element is equal to Log2MaxAtlasFrmOrderCntLsbForMiv bits. The value of ath_atlas_frm_order_cnt_lsb shall be in the range of 0 to MaxAtlasFrmOrderCntLsbForMiv-1, inclusive.
[0288] ath_additional_afoc_lsb_val[j] specifies the value of FullAtlasFrmOrderCntLsbLt[RlsIdx][j] for the current mosaic strip. If afps_heterogeneous_frame_tile_toolset_miv_present_flag of the current mosaic strip is equal to 0, then
[0289] FullAtlasFrmOrderCntLsbLt[RlsIdx][j]=ath_additional_afoc_lsb_val[j]*MaxAtlasFrmOrderCntLsb+afoc_lsb_lt[RlsIdx][j]
[0290] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic strip, then
[0291] FullAtlasFrmOrderCntLsbLt[RlsIdx][j]=ath_additional_afoc_lsb_val[j]*MaxAtlasFrmOrderCntLsbForMiv+afoc_lsb_lt[RlsIdx][j]
[0292] ath_additional_afoc_lsb_val[j] is represented by the afps_additional_lt_afoc_lsb_len bits. When afps_additional_lt_afoc_lsb_len is not present, the value of ath_additional_afoc_lsb_val[j] is inferred to be equal to 0.
[0293] ath_raw_3d_offset_axis_bit_count_minus1 plus 1 indicates the fixed bit width size of the values of the three syntax elements rpdu_3d_offset_u[tileID][p], rpdu_3d_offset_v[tileID][p] and rpdu_3e_offset_d[tileID][p], where p indicates that the sub-tile index is p and tileID indicates that the sub-tile is located in the slice with the slice ID equal to tileID.
[0294] When present, and if afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1, then the length of the syntax element representing ath_raw_3d_offset_axis_bit_count_minus1 is equal to Floor(Log2(asps_geometry_3d_bit_depth_minus1+1)).
[0295] When not present, the value of the ath_raw_3d_offset_axis_bit_count_minus1 syntax element is inferred to be
[0296] Max(0,asps_geometry_3d_bit_depth_minus1-asps_geometry_2d_bit_depth_minus1)-1.
[0297] The variable RawShift is defined as follows:
[0298] If afps_raw_3d_offset_bit_count_explicit_mode_flag=1,
[0299] RawShift=asps_geometry_3d_bit_depth_minus1-ath_raw_3d_offset_axis_bit_count_minus1
[0300] otherwise
[0301] RawShift=asps_geometry_2d_bit_depth_minus1+1
[0302] When present, and if afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1, then the length of the syntax element representing ath_raw_3d_offset_axis_bit_count_minus1 is equal to Floor(Log2(ashm_geometry_3d_bit_depth_minus1+1)).
[0303] When not present, the value of the ath_raw_3d_offset_axis_bit_count_minus1 syntax element is inferred to be
[0304] Max(0,ashm_geometry_3d_bit_depth_minus1-ashm_geometry_2d_bit_depth_minus1)-1.
[0305] The variable RawShift is defined as follows:
[0306] If afps_raw_3d_offset_bit_count_explicit_mode_flag=1,
[0307] RawShift=ashm_geometry_3d_bit_depth_minus1-ath_raw_3d_offset_axis_bit_count_minus1
[0308] otherwise
[0309] RawShift=ashm_geometry_2d_bit_depth_minus1+1
[0310] Patch data unit semantics
[0311] pdu_3d_offset_u[tileID][p] represents the offset of the reconstructed sub-tile along the tangent axis. The current sub-tile belongs to the sub-tile with sub-tile index p in the strip with strip index tileID.
[0312] If afps_heterogeneous_frame_tile_toolset_miv_present_flag for the current mosaic strip is equal to 0, then the value of pdu_3d_offset_u[tileID][p] shall be in the range of 0 to 2asps_geometry_3d_bit_depth_minus1+1-1 (inclusive). The value of the number of bits used to represent pdu_3d_offset_u[tileID][p] shall be asps_geometry_3d_bit_depth_minus1+1.
[0313] If afps_heterogeneous_frame_tile_toolset_miv_present_flag for the current mosaic slice is equal to 1, then the value of pdu_3d_offset_u[tileID][p] shall be in the range of 0 to 2ashm_geometry_3d_bit_depth_minus1+1-1, inclusive. The value of the number of bits used to represent pdu_3d_offset_u[tileID][p] shall be ashm_geometry_3d_bit_depth_minus1+1.
[0314] pdu_3d_offset_v[tileID][p] represents the offset along the bitangent axis for reconstructing the sub-tile, where the current sub-tile belongs to the sub-tile index p in the strip with the strip index tileID.
[0315] If afps_heterogeneous_frame_tile_toolset_miv_present_flag for the current mosaic strip is equal to 0, then the value of pdu_3d_offset_v[tileID][p] shall be in the range of 0 to 2asps_geometry_3d_bit_depth_minus1+1-1 (inclusive). The value of the number of bits used to represent pdu_3d_offset_v[tileID][p] shall be asps_geometry_3d_bit_depth_minus1+1.
[0316] If afps_heterogeneous_frame_tile_toolset_miv_present_flag for the current mosaic slice is equal to 1, then the value of pdu_3d_offset_v[tileID][p] shall be in the range of 0 to 2ashm_geometry_3d_bit_depth_minus1+1-1, inclusive. The value of the number of bits used to represent pdu_3d_offset_v[tileID][p] shall be ashm_geometry_3d_bit_depth_minus1+1.
[0317] pdu_3d_offset_d[tileID][p] represents the offset of the reconstructed sub-tile along the normal axis. The current sub-tile belongs to the sub-tile index p in the strip with the strip index tileID. Pdu3dOffsetD[tileID][p] is defined as follows:
[0318] Pdu3dOffsetD[tileID][p]=pdu_3d_offset_d[tileID][p]< <ath_pos_min_d_quantizer
[0319] If afps_heterogeneous_frame_tile_toolset_miv_present_flag for the current mosaic strip is equal to 0, then the value of Pdu3dOffsetD[tileID][p] shall be in the range of 0 to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive. The value used to represent the number of bits of pdu_3d_offset_v[tileID][p] shall be (asps_geometry_3d_bit_depth_minus1–ath_pos_min_d_quantizer+1).
[0320] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic strip, the value of Pdu3dOffsetD[tileID][p] shall be in the range of 0 to 2ashm_geometry_3d_bit_depth_minus1+1-1, inclusive. The value used to represent the number of bits of pdu_3d_offset_v[tileID][p] shall be (ashm_geometry_3d_bit_depth_minus1–ath_pos_min_d_quantizer+1).
[0321] pdu_3d_range_d[tileID][p] (if present) indicates the nominal maximum offset expected in the reconstructed bit depth sub-tile geometry samples along the normal axis after conversion to the nominal representation. The current sub-tile belongs to the sub-tile with sub-tile index p in the strip with strip index tileID. Pdu3dRangeD[tileID][p] is defined as follows:
[0322]
[0323] If the afps_heterogeneous_frame_tile_toolset_miv_present_flag of the current mosaic strip is equal to 0, the variable rangeDBitDepth takes the following value:
[0324] rangeDBitDepth=Min(ashm_geometry_2d_bit_depth_minus1,ashm_geometry_3d_bit_depth_minus1)+1
[0325] If pdu_3d_range_d[tileID][p] is not present, the value of Pdu3dRangeD[tileID][p] is inferred to be 2rangeDBitDepth – 1. If present, the value of Pdu3dRangeD[tileID][p] shall be in the range of 0 to 2rangeDBitDepth – 1, inclusive.
[0326] The number of bits representing pdu_3d_range_d[tileID][p] is equal to (rangeDBitDepth–ath_pos_delta_max_d_quantizer).
[0327] Merge Patch data unit semantics
[0328] mpdu_3d_offset_u[tileID][p] represents the offset difference along the tangent axis to be applied to reconstruct the two sub-tiles, where the two sub-tiles are the sub-tile with the current mosaic strip index tileID and the sub-tile index p, and the sub-tile with the current mosaic strip index tileID and the sub-tile index RefPatchIdx.
[0329] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of mpdu_3d_offset_u[tileID][p] shall be in the range of (-2asps_geometry_3d_bit_depth_minus1+1+1) to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0330] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic slice, the value of mpdu_3d_offset_u[tileID][p] shall be in the range of (-2ashm_geometry_3d_bit_depth_minus1+1+1) to (2ashm_geometry_3d_bit_depth_minus1+1–1), inclusive.
[0331] If mpdu_3d_offset_u[tileID][p] is not present, the value is inferred to be 0.
[0332] mpdu_3d_offset_v[tileID][p] represents the offset difference along the bitangent axis to be applied to reconstruct the two sub-tiles, where the two sub-tiles are the sub-tile with the current mosaic strip index tileID and the sub-tile index p, and the sub-tile with the current mosaic strip index tileID and the sub-tile index RefPatchIdx.
[0333] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of mpdu_3d_offset_v[tileID][p] shall be in the range of (-2asps_geometry_3d_bit_depth_minus1+1+1) to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0334] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic slice, the value of mpdu_3d_offset_v[tileID][p] shall be in the range of (-2ashm_geometry_3d_bit_depth_minus1+1+1) to (2ashm_geometry_3d_bit_depth_minus1+1–1), inclusive.
[0335] If mpdu_3d_offset_v[tileID][p] is not present, this value is inferred to be 0.
[0336] mpdu_3d_offset_d[tileID][p] represents the offset difference along the normal axis to be applied to reconstruct the two sub-tiles, where the two sub-tiles are the sub-tile with the current mosaic strip index tileID and the sub-tile index p and the sub-tile with the current mosaic strip index tileID and the sub-tile index RefPatchIdx. Mpdu3dOffsetD[tileID][p] is defined as follows:
[0337] Mpdu3dOffsetD[tileID][p]=mpdu_3d_offset_d[tileID][p]< <ath_pos_min_d_quantizer
[0338] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of mpdu_3d_offset_d[tileID][p] shall be in the range of (-2asps_geometry_3d_bit_depth_minus1+1+1) to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0339] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic slice, the value of mpdu_3d_offset_d[tileID][p] shall be in the range of (-2ashm_geometry_3d_bit_depth_minus1+1+1) to (2ashm_geometry_3d_bit_depth_minus1+1–1), inclusive.
[0340] If mpdu_3d_offset_d[tileID][p] is not present, the value is inferred to be 0.
[0341] Inter Merge Patch data unit semantics
[0342] ipdu_3d_offset_v[tileID][p] represents the offset difference along the bitangent axis to be applied to reconstruct the two sub-tiles, where the two sub-tiles are the sub-tile with the strip index tileID and the sub-tile index p in the current mosaic and the sub-tile with the strip index tileID and the sub-tile index RefPatchIdx in the current mosaic.
[0343] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of ipdu_3d_offset_v[tileID][p] shall be in the range of (-2asps_geometry_3d_bit_depth_minus1+1+1) to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0344] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic slice, the value of ipdu_3d_offset_v[tileID][p] shall be in the range of (-2ashm_geometry_3d_bit_depth_minus1+1+1) to (2ashm_geometry_3d_bit_depth_minus1+1–1), inclusive.
[0345] If ipdu_3d_offset_v[tileID][p] is not present, the value is inferred to be 0.
[0346] ipdu_3d_offset_d[tileID][p] represents the offset difference along the normal axis to be applied to reconstruct the two sub-tiles, where the two sub-tiles are the sub-tile with the strip index tileID and the sub-tile index p in the current mosaic and the sub-tile with the strip index tileID and the sub-tile index RefPatchIdx in the current mosaic. Mpdu3dOffsetD[tileID][p] is defined as follows:
[0347] Ipdu3dOffsetD[tileID][p]=ipdu_3d_offset_d[tileID][p]< <ath_pos_min_d_quantizer
[0348] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of ipdu_3d_offset_d[tileID][p] shall be in the range of (-2asps_geometry_3d_bit_depth_minus1+1+1) to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0349] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic slice, the value of ipdu_3d_offset_d[tileID][p] shall be in the range of (-2ashm_geometry_3d_bit_depth_minus1+1+1) to (2ashm_geometry_3d_bit_depth_minus1+1–1), inclusive.
[0350] If ipdu_3d_offset_d[tileID][p] is not present, the value is inferred to be 0.
[0351] The specifications in ISO / IEC 23090-5:2022 / Amd 1:- subclause 8.4 apply with the following additions.
[0352] Codestream conformance requires that asps_geometry_3d_bit_depth_minus1 and asps_geometry_2d_bit-depth_minus1 be equal to gi_geometroy_3d_coordinates_bit_depth_minus1 and gi_geametry_2d_bit_depth_minus1, respectively. However, in the special case, if asps_heterogeneous_miv_extension_present_flag is equal to 1, gi_geometry_3d_coordinates_bit_depth_minus1 and gi_geametry_2d_bit_depth_minus1 are specifically referred to as those in ISO / IEC 23090-5. ashm_geometroy_3d_bit-depth_minus1 and ashm_geometry_2d_bit_depth_minus1 are not necessarily equal to gi_geominatory_3d_coordinates_bit__depth_nus1 and gi_geometroy_2d_pth_minos1.
[0353] Sub-tile data unit multi-view extended syntax and semantics
[0354] pdu_depth_occ_threshold[tileID][p] indicates that in the stripe with the stripe index equal to tileID, for the sub-tile with the index equal to p, the occupancy value is set to unoccupied when it is lower than the threshold.
[0355] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the number of bits of pdu_depth_occ_threshold[tileID][p] is equal to asps_geometry_2d_bit_depth_minus1+1.
[0356] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic slice, the number of bits of pdu_depth_occ_threshold[tileID][p] is equal to ashm_geometry_2d_bit_depth_minus1+1.
[0357] If not present, pdu_depth_occ_threshold[tileID][p] is inferred to be dq_depth_occ_threshold_default[pdu_projection_id[tileID][p]]. Note that pdu_projection_id[tileID][p] corresponds to the view ID of the sub-tile with index equal to p, in the slice indexed by tileID.
[0358] 3. Standard-related decoding design
[0359] (1) Sequential counting of mosaic frames during decoding
[0360] The output of this process is AtlasFrmOrderCntVal, the mosaic frame order count for the current mosaic slice. The mosaic frame order count is used to identify the output order of mosaic frames and for decoder consistency checking. Each encoded mosaic frame is associated with a mosaic frame order count variable, denoted as AtlasFrmOrderCntVal.
[0361] When the current mosaic frame is not an IRAP coded mosaic with NoOutputBeforeRecoveryFlag equal to 1, the variables prevAtlasFrmOrderCntLsb and prevAtrasFrmOrderCntMsb are derived as follows:
[0362] Let prevAtlasFrm be the previous mosaic frame in decoding order whose TemporalID is equal to 0 and which is not a RASL, RADL or SLNR coded mosaic frame.
[0363] The variable prevAtlasFrmOrderCntLsb is set equal to the mosaic frame order count LSB value of prevAtrasFrm ath_atlas_frm_order_cnt_LSB.
[0364] The variable prevAtlasFrmOrderCntMsb is set equal to prevAtrasFrm's AtlasFrmaOrderCNTMsb.
[0365] The variable AtlasFrmOrderCntMsb of the current mosaic frame is derived as follows:
[0366] If the current mosaic is an IRAP-encoded mosaic and NoOutputBeforeRecoveryFlag is equal to 1, then AtlasFrmaOrderCNTMsb is set to 0.
[0367] Otherwise, AtlasFrmOrderCntMsb is derived as follows:
[0368]
[0369]
[0370] The derivation process of AtlasFrmOrderCntVal is as follows:
[0371] AtlasFrmOrderCntVal=AtlasFrmOrderCntMsb+ath_atlas_frm_order_cnt_lsb
[0372] The value range of AtlasFrmOrderCntVal is -2 31 to 2 31 –1 (inclusive). In one CAS, any two mosaic frames with the same nal_layer_id value have different AtlasFrmOrderCntVal.
[0373] The AtlasFrmOrderCnt(aFrmX) function is defined as follows:
[0374] AtlasFrmOrderCnt(aFrmX)=AtlasFrmOrderCntVal of the atlas frame aFrmX
[0375] The DiffAtlasFrmOrderCnt(aFrmA, aFrmB) function is defined as follows:
[0376] DiffAtlasFrmOrderCnt(aFrmA,aFrmB)=AtlasFrmOrderCnt(aFrmA)–AtlasFrmOrderCnt(aFrmB)
[0377] The bitstream shall not contain any value that would cause the DiffAtlasFrmOrderCnt(aFrmA, aFrmB) used during decoding to have a value other than -2. 15 to 2 15 Data in the range -1 (inclusive).
[0378] Note 1: Assume that X is the current mosaic frame, Y and Z are the other two mosaic frames in the same CAS, when DiffAtlasFrmOrderCnt(X, Y) and DiffAtlasFrmOrderCnt(X, Z) are both positive or both negative, Y and Z are considered to be in the same output order direction as X.
[0379] (2) Reference mosaic frame list processing process
[0380] This procedure is called at the beginning of the decoding process for each mosaic slice of a mosaic frame.
[0381] Reference mosaic frames are handled using a reference index. The reference index is an index into a reference mosaic frame list (RAFL). When decoding an I_TILE mosaic strip, RAFL is not used to decode the mosaic strip data. When decoding a SKIP_TILE or P_TILE mosaic strip, a single reference mosaic frame list RefAtlasFrmList is used to decode the mosaic strip data.
[0382] At the beginning of the decoding process of each mosaic slice, a RAFL RefAtlasFrmList is derived. RAFL is used for reference mosaic frame marking or mosaic slice data decoding as specified in subclause 9.2.4.4.
[0383] NOTE 1: For I_TILE slices of a mosaic frame, RefAtlasFrmList may be used for bitstream conformance checking, but its derivation is not required for decoding of the current mosaic frame or mosaics that follow the current mosaic frame in decoding order.
[0384] The reference mosaic frame list RefAtlasFrmList is constructed as follows:
[0385]
[0386] The first NumRefIdxActive entry in RefAtlasFrmList is called the active entry in RefAtlasFrmList, and the other entries in RefAtlasFlmList are called inactive entries in RefAtrasFrmLists.
[0387] If the current strip is a SKIP_tile, the array RefAtduTotalNumPatches is set to the array AtduToTotalNumPatches corresponding to the first entry in RefAtlasFrmList, RefAtlasFlmList[0].
[0388] Bitstream conformance requires that the following constraints apply:
[0389] –num_ref_entries[RlsIdx] must not be less than NumRefIdxActive.
[0390] – The mosaic frame referenced by each active entry in RefAtlasFrmList shall exist in the DAB and its temporal ID shall be less than or equal to the temporal ID of the current mosaic frame.
[0391] – Each entry in RefAtlasFrmList shall reference a mosaic frame that is not the current mosaic frame.
[0392] – The short-term reference mosaic frame entry and the long-term reference mosaic frame entry in the mosaic strip RefAtlasFrmList shall not refer to the same mosaic frame.
[0393] – The difference between the AtlasFrmaOrderCntVal of the current mosaic strip and the AtlasFlmOrderCNTVal of the mosaic frame pointed to by the entry should not be greater than or equal to 2 24 The long-term reference mosaic frame entry.
[0394] – Let setOfRefAtlasFrms be the only mosaic frame referenced by all entries in RefAtlasFlmList that have the same nal_layer_id as the current mosaic frame. The number of mosaic frames in setOfRefAtlasFrms should be less than or equal to asps_max_dec_atlas_frame_buffering_minus1 and setOfrefAtlasFms should be the same for all mosaic strips of the mosaic frame.
[0395] – The mosaic frame referenced by each active entry in RefAtlasFrmList should have exactly the same number of strips as the current mosaic frame.
[0396] – The RefAtlasFrmList of all slices in the current mosaic frame should contain the same reference mosaic frame, but there is no restriction on the order of the reference mosaic frames.
[0397] – If the current mosaic frame (nal_layer_id equals a particular value layerID) is an IRAP-coded mosaic, the mosaic referenced by the entry in RefAtlasFrmList shall not precede any previous IRAP-coded mosaic (in decoding order, when nal_layer_id equals layerID) in output order or decoding order.
[0398] – When the current mosaic frame is not a RASL coded mosaic associated with a CRA coded mosaic with NoOutputBeforeRecoveryFlag equal to 1, the mosaic referenced by the active entry in RefAtlasFrmList, which was generated by the decoding process to generate an unusable reference mosaic frame for the CRA coded mosaic associated with the current mosaic, shall not exist.
[0399] – When the current mosaic frame follows an IRAP coded mosaic with the same nal_layer_id value in both decoding order and output order, the mosaic frame referenced by the active entry in RefAtlasFrmList shall not precede the IRAP coded mosaic in either output order or decoding order.
[0400] – When the current mosaic frame follows an IRAP coded mosaic with the same nal_layer_id value and all preceding mosaic frames (if any) associated with that IRAP coded mosaic in decoding order and output order, entries in RefAtlasFrmList shall not reference mosaics preceding the IRAP coded mosaic in output order or decoding order.
[0401] – When the current mosaic frame is a RADL coded mosaic, there shall be no active entry in the RefAtlasFrmList that is a mosaic frame that precedes the RADL coded mosaic’s associated IRAP coded mosaic in decoding order.
[0402] (3) General decoding process of sub-block data unit
[0403] TilePatch3dOffsetU[tileID][p] represents the offset along the tangent axis for reconstructing the sub-tile, where the current sub-tile belongs to the sub-tile index p in the strip with the strip index tileID.
[0404] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of TilePatch3dOffsetU[tileID][p] shall be in the range of 0 to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0405] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic strip, the value of TilePatch3dOffsetU[tileID][p] shall be in the range of 0 to (2ashm_geometry_3d_bit_depth_minus1+1-1), inclusive.
[0406] TilePatch3dOffsetV[tileID][p] represents the offset along the bitangent axis to reconstruct the sub-tile with sub-tile index p in the strip with strip index tileID.
[0407] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of TilePatch3dOffsetV[tileID][p] shall be in the range of 0 to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0408] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic strip, the value of TilePatch3dOffsetV[tileID][p] shall be in the range of 0 to 2ashm_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0409] TilePatch3dOffsetD[tileID][p] represents the offset along the normal axis for reconstructing the sub-tile, where the current sub-tile belongs to the sub-tile index p in the strip with the strip index tileID.
[0410] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 0 for the current mosaic strip, the value of TilePatch3dOffsetD[tileID][p] shall be in the range of 0 to 2asps_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0411] If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic strip, the value of TilePatch3dOffsetD[tileID][p] shall be in the range of 0 to 2ashm_geometry_3d_bit_depth_minus1+1-1, inclusive.
[0412] TilePatch3dRangeD[tileID][p], if present, represents the nominal maximum offset expected in the reconstructed bit depth sub-tile geometry samples along the normal axis after conversion to the nominal representation. The current sub-tile belongs to the sub-tile with sub-tile index p in the strip with strip index tileID.
[0413] If the afps_heterogeneous_frame_tile_toolset_miv_present_flag of the current mosaic strip is equal to 0, the variable
[0414] rangeDBitDepth=Min(ashm_geometry_2d_bit_depth_minus1,ashm_geometry_3d_bit_depth_minus1)+1
[0415] If the current mosaic strip's afps_heterogeneous_frame_tile_toolset_miv_present_flag is 1, the variable rangeDBitDepth = Min(ashm_geometry_2d_bit_depth_minus1, ashm_geometry_3d_bit_depth_minus1) + 1. TilePatch3dRangeD[tileID][p] takes a value between 0 and 2rangeDBitDepth – 1 (inclusive).
[0416] 4. Standard-related syntax restrictions.
[0417] Table 6. Maximum allowed syntax element values for the V-PCC toolset profile components.
[0418] Table 6
[0419]
[0420] Table 7 Max allowed syntax element values for the heterogeneous toolset profile components Extended
[0421] Table 7
[0422]
[0423]
[0424] Table 8 shows the maximum allowed syntax element values for the MIV toolset profile components.
[0425] Table 8
[0426]
[0427]
[0428] Table 9-1-1 Allowable values of syntax element values for the heterogeneous toolset profile components Extended
[0429] Table 9-1-1
[0430]
[0431]
[0432]
[0433]
[0434] Table 9-1-2 Allowable values of syntax element values for the heterogeneous toolset profile components Extended
[0435] Table 9-1-2
[0436]
[0437]
[0438]
[0439] New extended syntax element design
[0440] Bitstream conformance, for each bitstream conformance test, all of the following conditions shall be met:
[0441] 1. For each coded atlas access unit n (n greater than 0) associated with the buffering period SEI message, let the variable deltaTime90k[n] specify the following:
[0442] deltaTime90k[n]=90000*(AuNominalRemovalTime[n]-AuFinalArrivalTime[n-1])
[0443] The value constraints of InitCabRemovalDelay[Htid][SchedSelIdx] are as follows:
[0444] – If hrd_cbr_flag[!NalHrdModeFlag][Htid][SchedSelIdx] is equal to 0, then the following conditions are true:
[0445] InitCabRemovalDelay[Htid][SchedSelIdx]<=Ceil(deltaTime90k[n])
[0446] – Otherwise hrd_cbr_flag[!NalHrdModeFlag][Htid][SchedSelIdx] is equal to 1, then the following conditions are true:
[0447] Floor(deltaTime90k[n])<=InitCabRemovalDelay[Htid][SchedSelIdx]<=Ceil(deltaTime90k[n])
[0448] NOTE 1 – The exact number of bits per mosaic frame deleted in the CAB depends on which buffering period SEI message is chosen to initialize the HRD. The encoder should take this into account to ensure that all specified constraints are adhered to. The HRD can be initialized in any buffering period SEI message.
[0449] 2.CAB overflow is specified as the situation where the total number of bits in the CAB is greater than the CAB size. A CAB shall not overflow.
[0450] 3. When hrd_low_delay_flag[Htid] is equal to 0, CAB will never underflow. CAB underflow is defined as follows:
[0451] – CAB underflow is specified as the condition that the nominal CAB removal time AuNominalRemovalTime[n] of coded splice map access unit n is less than the final CAB arrival time of coded splice map access unit n AuFinalArrivalTime[n], for at least one value n.
[0452] 4. The nominal removal time of a mosaic frame from a CAB (starting from the second mosaic in decoding order) shall satisfy the constraints on AuNominalRemovalTime[n] and AuCabRemovalTime[n] in Annex A.
[0453] 5. For each current mosaic frame, the number of mosaic frames decoded in the DAB after the procedure for removing mosaics from the DAB has been called as specified, including all mosaic frames n marked as "used for reference" or for which AtlasFrameOutputFlag is equal to 1 and AuCabRemovalTime[n] is less than AuCabRemovalTime[currAtlasFrame], where currAtlasFrame is the current mosaic frame and shall be less than or equal to asps_max_dec_alas_frame_buffering_minus1.
[0454] 6. When prediction is required, all reference mosaic frames shall be present in the DAB. Each mosaic frame with AtlasFrameOutputFlag equal to 1 shall be present in the DAB at the time of DAB output, unless it is removed from the DAB before its output time by one of the processes specified in clause
[0455] 7. For each current mosaic frame that is not IRAP coded and has NoOutputBeforeRecoveryFlag equal to 1, the value of maxAtlasFrameOrderCnt - minAtlasFrameOrderCnt shall be less than MaxAtlasFrmOrderCntLsb / 2. If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic slice, the value of maxAtlasFrameOrderCnt - minAtlasFrameOrderCnt shall be less than MaxAtlasFrmOrderCntLsb ForMiv / 2.
[0456] 8. The value of DabOutputInterval[n] is the difference between the output time of one mosaic frame and the output time of the first mosaic frame after it, and AtlasFrameOutputFlag is equal to 1, which shall satisfy the level specified in the bitstream for the profile, layer, and specified decoding process.
[0457] 9. For any two mosaic frames m and n in the same CAS, when DabOutputTime[m] is greater than DabOutput Time[n], the AtlasFrmOrderCntVal of mosaic frame m should be greater than the AtlasFrmOrderCntVal of mosaic frame n.
[0458] Recovery point SEI message semantics
[0459] recovery_afoccnt specifies the recovery point of the decoded mosaic frame in output order. If there is a mosaic frame aFrmA in the CAS that follows the current mosaic frame in decoding order (i.e., the mosaic frame associated with the current SEI message) and whose AtlasFrmOrderCntVal is equal to the Atlasfrmardercntval of the current mosaic frame plus the value of recovery_afoc_cnt, then the atlas frame aFrmA is called the recovery point mosaic frame. Otherwise, the first mosaic frame in output order whose AtlasFrmOrderCntVal is greater than the AtlasFrmOrderCNTVal of the current mosaic frame plus the value of recovery_afoc_cnt is called the recovery point mosaic frame. The recovery point mosaic frame should not precede the current mosaic frame in decoding order. Starting from the output order position of the recovery point mosaic frame, all decoded mosaic frames displayed in output order are correct or approximately correct in content. The value of recovery_afoc_cnt shall be in the range -MaxAtlasFrmOrderCntLsb / 2 to MaxAtlasFlmOrderCNTLsb / 2-1 (inclusive). If afps_heterogeneous_frame_tile_toolset_miv_present_flag is equal to 1 for the current mosaic frame slice, the value of recovery_afoc_cnt shall be in the range -MaxAtlasFrmOrderCntLsbForMiv / 2 to MaxAtlasFlmOrderCNTLsbForMIv / 2–1.
[0460] Generic ASPS-level strings
[0461] The ASPSCommonByteString(stringByte,posByte) function is defined as follows:
[0462]
[0463] VUI parameter semantics
[0464] vui_display_box_origin[d] specifies the offset along axis d relative to the origin of the coordinate system. When the element vui_ddisplay_box_origin[d] does not exist, its value shall be inferred to be 0. If the current mosaic strip's afps_heterogeneous_frame_tile_toolset_miv_present_flag is 0, the number of bits used to represent vui_display_box_origin[d] is asps_geometry_3d_bit_depth_minus1+1. If the current mosaic strip's afps_heterogeneous_frame_tile_toolset_miv_present_flag is 1, the number of bits used to represent vui_display_box_origin[d] is ashm_geometry_3d_pit_depth_minus1+1. Values of d equal to 0, 1, and 2 correspond to the X, Y, and Z axes, respectively.
[0465] vui_display_box_size[d] specifies the size of the display box, sampled along axis d. When the element of vui_display_box_size[d] does not exist, its value is unknown. If the current mosaic strip has afps_heterogeneous_frame_tile_toolset_miv_present_flag equal to 0, then the number of bits used to indicate vui_display_box_size[d] is asps_geometry_3d_bit_depth_minus1+1. If the current mosaic strip has afps_heterogeneous_frame_tile_toolset_miv_present_flag equal to 1, then the number of bits used to indicate vui_display_box_size[d] is asps_geometry_3d_bit_depth_minus1+1.
[0466] The following variables come from the display box parameters:
[0467] minOffset[d]=vui_display_box_origin[d](G-1)
[0468] maxOffset[d]=vui_display_box_origin[d]+vui_display_box_size[d]
[0469] The values of d equal to 0, 1, and 2 correspond to the X, Y, and Z axes respectively.
[0470] vui_anchor_point[d] represents the position of the anchor point along the d-axis. If the current mosaic strip has afps_heterogeneous_frame_tile_toolset_miv_present_flag set to 0, the value of vui_anchor_point[d] shall be in the range 0 to 2asps_geometry_3d_bit_depth_minus1+1-1. If vui_anchor_point[d] is not present, it shall be inferred to be equal to 0. The number of bits used to represent vui_anchor_point[d] is asps_geometry_3d_bit_depth_minus1+1. If the current mosaic strip has afps_heterogeneous_frame_tile_toolset_miv_present_flag set to 1, the value of vui_anchor_point[d] shall be in the range 0 to 2ashm_geometry_3d_pth_minos1+1-1. If vui_anchor_point[d] is not present, it shall be inferred to be equal to 0. The number of bits used to represent vui_ancor_point[d] is ashm_geometry_3d_bit_depth_minus1 + 1. d values of 0, 1, and 2 correspond to the X, Y, and Z axes, respectively.
[0471] Multi-viewpoint standard, can be used by ISO / IEC 23090-12
[0472] Deep expansion process
[0473] This process expands the integer depth values of the mosaic into floating-point depth values in scene coordinates (e.g. meters).
[0474] Integer depth values may be scaled to an implementation-defined bit depth and in the range 0…maxSampleD. Otherwise, maxSampleD is set to 2asps_geometry_2d_bit_depth_minus1+1–1.
[0475] In the special case, if asps_heterogeneous_miv_extension_present_flag is equal to 1, maxSampleD is set to 2ashm_geometriy_2d_pit_deptho_minus1+1–1
[0476] Rebuilding the MPI process
[0477] This process decodes the reconstructed volume frames and reconstructs the MPI frames from the bitstream where ptc_restricted_geometry_flag is equal to 1.
[0478] NOTE – The reconstruction process described will reconstruct the entire MPI frame. Implementations may form the viewport before projecting it into the viewport without buffering the entire set of textures and transparency layers.
[0479] Inputs to this process include:
[0480] - a view parameter list containing the intrinsic and extrinsic parameters of the (unique) source view with index viewIdx;
[0481] -For each mosaic:
[0482] - variable atlasID, which is the mosaic ID;
[0483] -Variables AspsFrameHeight[atlasID] and AspsFrameWidth[atlasID] represent the number of rows and columns of the mosaic frame respectively;
[0484] -2D array AtlasBlockToPatchMap;
[0485] - variable PatchPackingBlockSize;
[0486] -The three-dimensional array size of texFrame is 3×AspsFrameHeight[atlasID]×AspsFrameworkWidth[atlasID];
[0487] -transpFrame 2D data size is AspsFrameHeight[atlasID]×AspsFrameWidth[atlasID];
[0488] NOTE – A transparency level of 0 corresponds to a fully transparent sample, while a maximum transparency level of 2(ai_attribute_2d_bit_depth_minus1[atlasID][attrIdx]+1)–1, where attrIdx is the index of the transparency attribute, corresponds to a fully opaque sample. The encoding rule is linear between the minimum and maximum transparency levels.
[0489] - extrinsic and intrinsic parameters of the target view;
[0490] The variable maxDepthSampleValue indicates the maximum value of the coded geometry sample and is set to 2(asps_geometry_3d_bit_depth_minus1+1)–1. In the special case, if asps_heterogeneous_miv_extension_present_flag is equal to 1, the variable maxDepthSampleValue indicates the maximum value of the coded geometry sample and is set to 2(ashm_geometry_3d_bit_depth_minus1+1)–1.
[0491] -The constant maxNbLayers indicates the maximum number of depth layers of MPI, which is set to maxDepthSampleValue+1.
[0492] The embodiment of the present application is used to implement a coding and decoding scheme in which multi-viewpoint mosaics, point cloud mosaics, and heterogeneous mixed mosaics coexist in the code stream, and expands the relevant standards. It has the following advantages: 1) For application scenarios composed of data in different formats, this method can be used to provide real-time immersive video interaction services for data in different formats (such as 3D grids, 3D point clouds, multi-view images, etc.), thereby promoting the development of the VR / AR / MR industry; 2) Compared with encoding multi-viewpoint video images and point cloud format data separately and then calling their own decoders to independently demultiplex the signals, the number of decoders to be called is small, the processing pixel rate of the decoders is fully utilized, and the hardware requirements are reduced; 3) the rendering advantages of data from different formats (point clouds, etc.) are retained to improve the synthesis quality of the image; 4) the reconstruction quality and coding performance of heterogeneous data are further improved.
[0493] The present application also provides an encoding device. FIG11 is a schematic block diagram of an encoding device provided in an embodiment of the present application. The encoding device 110 is applied to an encoder. As shown in FIG11 , the encoding device 110 includes:
[0494] The processing unit 1101 is configured to process the visual media content in at least two expression formats to obtain at least two isomorphic blocks;
[0495] The splicing unit 1102 is configured to splice the at least two isomorphic blocks to obtain a spliced graph and spliced graph information, wherein when the spliced graph is a heterogeneous mixed spliced graph, the spliced graph includes at least two isomorphic blocks, and different isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information;
[0496] The encoding unit 1103 is configured to encode the splicing graph and the splicing graph information to obtain a code stream.
[0497] In some embodiments, the spliced graph information includes a first syntax element, and it is determined according to the first syntax element whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph.
[0498] In some embodiments, the first syntax element includes a first sub-syntax element and a second sub-syntax element; and determining, based on the first syntax element, whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph includes:
[0499] If the value of the first sub-syntax element is equal to the value of the second sub-syntax element, it is determined according to the values whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph.
[0500] In some embodiments, determining whether the spliced image is a heterogeneous mixed spliced image or a homogeneous spliced image based on the value includes: if the value is a first preset value, determining that the spliced image is a heterogeneous mixed spliced image; if the value is a second preset value, determining that the spliced image is a homogeneous spliced image.
[0501] In some embodiments, determining, based on the value, whether the mosaic is a heterogeneous mixed mosaic or a homogeneous mosaic includes:
[0502] If the value is a third preset value, it is determined that the spliced graph is a heterogeneous mixed spliced graph including homogeneous blocks in a first expression format and a second expression format, wherein the first expression format and the second expression format are different expression formats;
[0503] If the value is a fourth preset value, it is determined that the spliced graph is an isomorphic spliced graph including isomorphic blocks in the first expression format;
[0504] If the value is a fifth preset value, it is determined that the spliced graph is an isomorphic spliced graph including isomorphic blocks in the second expression format.
[0505] In some embodiments, the first sub-syntax element is a syntax element of a mosaic view sequence parameter set ASPS, and the second sub-syntax element is a syntax element of a mosaic view frame parameter set AFPS.
[0506] In some embodiments, the splicing graph information does not include the first syntax element, and the splicing graph is determined to be a homogeneous splicing graph.
[0507] In some embodiments, when the spliced graph is a heterogeneous mixed spliced graph, the spliced graph information further includes a second syntax element; and the expression format of the homogeneous blocks in the spliced graph is determined according to the second syntax element.
[0508] In some embodiments, determining the expression format of the isomorphic blocks in the spliced graph based on the second syntax element includes: when the value of the second syntax element of the i-th block is a sixth preset value, determining that the expression format of the i-th block is the first expression format; when the value of the second syntax element of the i-th block is a seventh preset value, determining that the expression format of the i-th block is the second expression format.
[0509] In some embodiments, when the mosaic graph is a heterogeneous mixed mosaic graph, the mosaic graph information includes at least two types of homogeneous block information, wherein homogeneous blocks in different expression formats correspond to different homogeneous block information.
[0510] In some embodiments, the homogeneous block information includes syntax elements of ASPS and syntax elements of AFPS; different homogeneous block information corresponds to different syntax elements of ASPS and syntax elements of AFPS.
[0511] In some embodiments, when the isomorphic block is a multi-view video block, it corresponds to first isomorphic block information; when the isomorphic block is a point cloud block, it corresponds to second isomorphic block information; the first isomorphic block information and the second isomorphic block information include the shared syntax elements of the ASPS parameter set and the syntax elements of the AFPS parameter set; the first isomorphic block information also includes the extended syntax elements of the ASPS parameter set and the extended syntax elements of the AFPS parameter set.
[0512] In some embodiments, the parameter set sub-codestream of the codestream includes a third syntax element, and the codestream corresponding to the visual media content including at least one expression format in the codestream is determined according to the third syntax element.
[0513] In some embodiments, determining the codestream corresponding to the visual media content in at least one expression format in the codestream based on the third syntax element includes: the third syntax element is a first value, determining that the codestream includes both the codestream corresponding to the visual media content in the first expression format and the codestream corresponding to the visual media content in the second expression format; the third syntax element is a second value, determining that the codestream includes the visual media content in the first expression format; and the third syntax element is a third value, determining that the codestream includes the visual media content in the second expression format.
[0514] In some embodiments, when the spliced graph is a heterogeneous mixed spliced graph, the bitstream corresponding to the visual media content in at least two expression formats in the bitstream is determined according to the third syntax element.
[0515] In some embodiments, the encoding unit 1103 is configured to encode the mosaic to obtain a video compression substream; encode the mosaic information to obtain a mosaic information substream; and combine the video compression substream and the mosaic information substream into the codestream.
[0516] In some embodiments, the representation format is a multi-view video, a point cloud, or a mesh.
[0517] In some embodiments, the heterogeneous mixed mosaic graph includes at least one of the following: a single-attribute heterogeneous mixed mosaic graph and a multi-attribute heterogeneous mixed mosaic graph; the homogeneous mosaic graph includes at least one of the following: a single-attribute homogeneous mosaic graph and a multi-attribute homogeneous mosaic graph.
[0518] The present application also provides a decoding device. FIG12 is a schematic block diagram of a decoding device provided by an embodiment of the present application. The decoding device 120 is applied to a decoder. As shown in FIG12 , the decoding device 120 includes:
[0519] The decoding unit 1201 is configured to decode the code stream to obtain a splicing graph and splicing graph information;
[0520] The splitting unit 1202 is configured to obtain at least two types of homogeneous blocks and homogeneous block information based on the mosaic and the mosaic information when the mosaic is a heterogeneous mixed mosaic; wherein different types of homogeneous blocks correspond to different visual media content expression formats and different homogeneous block information;
[0521] The splitting unit 1202 is configured to obtain a homogeneous block and homogeneous block information according to the mosaic and the mosaic information when the mosaic is a homogeneous mosaic;
[0522] The processing unit 1203 is configured to obtain visual media content in at least two expression formats according to the isomorphic blocks and the isomorphic block information.
[0523] In some embodiments, the spliced graph information includes a first syntax element, and it is determined according to the first syntax element whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph.
[0524] In some embodiments, the first syntax element includes a first sub-syntax element and a second sub-syntax element; determining whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph based on the first syntax element includes: if the value of the first sub-syntax element and the value of the second sub-syntax element are equal, determining whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph based on the values.
[0525] In some embodiments, determining whether the spliced image is a heterogeneous mixed spliced image or a homogeneous spliced image based on the value includes: if the value is a first preset value, determining that the spliced image is a heterogeneous mixed spliced image; if the value is a second preset value, determining that the spliced image is a homogeneous spliced image.
[0526] In some embodiments, determining whether the splicing graph is a heterogeneous mixed splicing graph or a homogeneous splicing graph based on the value includes: if the value is a third preset value, determining that the splicing graph is a heterogeneous mixed splicing graph including homogeneous blocks in a first expression format and a second expression format, wherein the first expression format and the second expression format are different expression formats; if the value is a fourth preset value, determining that the splicing graph is an isomorphic splicing graph including homogeneous blocks in the first expression format; if the value is a fifth preset value, determining that the splicing graph is an isomorphic splicing graph including homogeneous blocks in the second expression format.
[0527] In some embodiments, the first sub-syntax element is a syntax element of a mosaic view sequence parameter set ASPS, and the second sub-syntax element is a syntax element of a mosaic view frame parameter set AFPS.
[0528] In some embodiments, the splicing graph information does not include the first syntax element, and the splicing graph is determined to be a homogeneous splicing graph.
[0529] In some embodiments, when the spliced graph is a heterogeneous mixed spliced graph, the spliced graph information further includes a second syntax element; and the expression format of the homogeneous blocks in the spliced graph is determined according to the second syntax element.
[0530] In some embodiments, determining the expression format of the isomorphic blocks in the spliced graph based on the second syntax element includes: when the value of the second syntax element of the i-th block is a sixth preset value, determining that the expression format of the i-th block is the first expression format; when the value of the second syntax element of the i-th block is a seventh preset value, determining that the expression format of the i-th block is the second expression format.
[0531] In some embodiments, when the mosaic graph is a heterogeneous mixed mosaic graph, the mosaic graph information includes at least two types of homogeneous block information, wherein homogeneous blocks in different expression formats correspond to different homogeneous block information.
[0532] In some embodiments, when the mosaic is a heterogeneous mixed mosaic, obtaining at least two isomorphic blocks and isomorphic block information according to the mosaic and the mosaic information includes: when the mosaic is a heterogeneous mixed mosaic, splitting the mosaic to obtain at least two isomorphic blocks; and obtaining isomorphic block information corresponding to the at least two isomorphic blocks from the mosaic information according to the expression format of the at least two isomorphic blocks.
[0533] In some embodiments, the homogeneous block information includes syntax elements of ASPS and syntax elements of AFPS; different homogeneous block information corresponds to different syntax elements of ASPS and syntax elements of AFPS.
[0534] In some embodiments, when the isomorphic block is a multi-view video block, it corresponds to first isomorphic block information; when the isomorphic block is a point cloud block, it corresponds to second isomorphic block information; the first isomorphic block information and the second isomorphic block information include the shared syntax elements of the ASPS parameter set and the syntax elements of the AFPS parameter set; the first isomorphic block information also includes the extended syntax elements of the ASPS parameter set and the extended syntax elements of the AFPS parameter set.
[0535] In some embodiments, the parameter set sub-codestream of the codestream includes a third syntax element, and the codestream corresponding to the visual media content including at least one expression format in the codestream is determined according to the third syntax element.
[0536] In some embodiments, determining the codestream corresponding to the visual media content in at least one expression format in the codestream based on the third syntax element includes: the third syntax element is a first value, determining that the codestream includes both the codestream corresponding to the visual media content in the first expression format and the codestream corresponding to the visual media content in the second expression format; the third syntax element is a second value, determining that the codestream includes the visual media content in the first expression format; and the third syntax element is a third value, determining that the codestream includes the visual media content in the second expression format.
[0537] In some embodiments, decoding the bitstream to obtain the splicing graph and the splicing graph information includes: determining, based on the second syntax element, a bitstream corresponding to visual media content including at least two expression formats in the bitstream, and decoding the bitstream to obtain a heterogeneous mixed splicing graph and the splicing graph information.
[0538] In some embodiments, the decoding unit 1201 is configured to decode the video compression sub-stream to obtain the splicing graph; and decode the splicing graph information sub-stream to obtain the splicing graph information.
[0539] In some embodiments, the representation format is a multi-view video, a point cloud, or a mesh.
[0540] In some embodiments, the heterogeneous mixed mosaic graph includes at least one of the following: a single-attribute heterogeneous mixed mosaic graph and a multi-attribute heterogeneous mixed mosaic graph; the homogeneous mosaic graph includes at least one of the following: a single-attribute homogeneous mosaic graph and a multi-attribute homogeneous mosaic graph.
[0541] It should be understood that the device embodiment and the method embodiment may correspond to each other, and similar descriptions may refer to the method embodiment. To avoid repetition, they will not be described here.
[0542] The above describes the apparatus and system of the embodiment of the present application from the perspective of functional units in conjunction with the accompanying drawings. It should be understood that the functional unit can be implemented in the form of hardware, can be implemented by instructions in the form of software, or can be implemented by a combination of hardware and software units. Specifically, the steps of the method embodiment in the embodiment of the present application can be completed by the hardware integrated logic circuit and / or software instructions in the processor, and the steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software units in the decoding processor. Optionally, the software unit can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps in the above method embodiment in conjunction with its hardware.
[0543] In practical applications, an embodiment of the present application further provides an encoder. FIG13 is a schematic block diagram of an encoder provided in an embodiment of the present application. As shown in FIG13 , the encoder 1310 includes:
[0544] The second memory 1320 and the second processor 1330; the second memory 1320 stores a computer program that can be run on the second processor 1330, and the second processor 1330 executes the program when the encoding method on the encoder side is executed.
[0545] In practical applications, an embodiment of the present application further provides a decoder. FIG14 is a schematic block diagram of a decoder provided in an embodiment of the present application. As shown in FIG14 , a decoder 1410 includes:
[0546] The first memory 1420 and the first processor 1430; the first memory 1420 stores a computer program that can be run on the first processor 1430, and the first processor 1430 executes the decoding method on the decoder side when executing the program.
[0547] In some embodiments of the present application, the processor may include but is not limited to:
[0548] General-purpose processor, Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware components, etc.
[0549] In some embodiments of the present application, the memory includes but is not limited to:
[0550] Volatile memory and / or non-volatile memory. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus random access memory (DR RAM).
[0551] In addition, the functional modules in this embodiment can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or software functional modules.
[0552] In yet another embodiment of the present application, see Figure 15 , which shows a schematic diagram of the structure of a coding and decoding system provided in an embodiment of the present application. As shown in Figure 15 , the coding and decoding system 150 may include an encoder 1501 and a decoder 1502. The encoder 1501 may be a device integrated with the encoding device described in the previous embodiment; the decoder 1502 may be a device integrated with the decoding device described in the previous embodiment.
[0553] In an embodiment of the present application, in the encoding and decoding system 150, both the encoder 1501 and the decoder 1502 can use the color component information of adjacent reference pixels and the pixels to be predicted to calculate the corresponding weighting coefficients of the pixels to be predicted; and different reference pixels can have different weighting coefficients. Applying this weighting coefficient to the chrominance prediction of the pixels to be predicted in the current block can not only improve the accuracy of the chrominance prediction and save bit rate, but also improve the encoding and decoding performance.
[0554] The present application also provides a chip for implementing the above encoding and decoding method. Specifically, the chip includes a processor for calling and running a computer program from a memory, so that an electronic device equipped with the chip executes the above encoding and decoding method.
[0555] The present application also provides a computer storage medium storing a computer program, which, when executed by a second processor, implements the encoding method of the encoder; or, when executed by a first processor, implements the decoding method of the decoder. Alternatively, the present application also provides a computer program product comprising instructions, which, when executed by a computer, causes the computer to perform the method of the aforementioned method embodiment.
[0556] The present application also provides a code stream, which is generated according to the above encoding method. Optionally, the code stream includes the above first syntax element, or includes the second syntax element and the third syntax element.
[0557] When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid state drive (SSD)).
[0558] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0559] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the unit is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0560] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment. For example, the functional units in the various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0561] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another and are not necessarily used to describe a specific order or precedence. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Second information may appear before, after, or simultaneously with the first information.
[0562] The above content is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims. Industrial Applicability
[0563] The present application provides a coding and decoding method, apparatus, encoder, decoder, and storage medium. For application scenarios involving visual media content in one or more expression formats, the method splices homogeneous blocks of different expression formats into a heterogeneous hybrid mosaic. This splicing of homogeneous blocks of different expression formats into a heterogeneous hybrid mosaic is then performed for encoding and decoding. This method can reduce the number of encoders and decoders required, lower implementation costs, and improve usability. Furthermore, in the heterogeneous hybrid mosaic, certain high-level parameters of blocks of different expression formats can be unequal, thereby providing more appropriate high-level parameters for heterogeneous data, effectively improving coding efficiency, i.e., reducing bitrate or improving the quality of reconstructed multi-view video or point cloud video.
Claims
1. A decoding method, wherein: include: Decode the code stream to obtain the splicing graph and splicing graph information; When the mosaic is a heterogeneous mixed mosaic, at least two types of isomorphic blocks and isomorphic block information are obtained according to the mosaic and the mosaic information; wherein different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information; When the mosaic is a homogeneous mosaic, obtaining a homogeneous block and homogeneous block information according to the mosaic and the mosaic information; Visual media content in at least two expression formats is obtained according to the isomorphic blocks and the isomorphic block information.
2. The method according to claim 1, wherein The spliced graph information includes a first syntax element, and it is determined according to the first syntax element whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph.
3. The method according to claim 2, wherein: The first syntax element includes a first sub-syntax element and a second sub-syntax element; The determining, according to the first syntax element, whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph includes: If the value of the first sub-syntax element is equal to the value of the second sub-syntax element, it is determined according to the values whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph.
4. The method according to claim 3, wherein: The determining, according to the value, whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph includes: If the value is a first preset value, it is determined that the spliced graph is a heterogeneous mixed spliced graph; If the value is a second preset value, it is determined that the spliced graph is a homogeneous spliced graph.
5. The method according to claim 3, wherein The determining, according to the value, whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph includes: If the value is a third preset value, it is determined that the spliced graph is a heterogeneous mixed spliced graph including homogeneous blocks in a first expression format and a second expression format, wherein the first expression format and the second expression format are different expression formats; If the value is a fourth preset value, it is determined that the spliced graph is an isomorphic spliced graph including isomorphic blocks in the first expression format; If the value is a fifth preset value, it is determined that the spliced graph is an isomorphic spliced graph including isomorphic blocks in the second expression format.
6. The method according to claim 3, wherein: The first sub-syntax element is a syntax element of an ASPS (joint picture sequence parameter set), and the second sub-syntax element is a syntax element of an AFPS (joint picture frame parameter set).
7. The method according to any one of claims 2 to 6, wherein: The method further comprises: The splicing graph information does not include the first syntax element, and the splicing graph is determined to be a homogeneous splicing graph.
8. The method according to any one of claims 1 to 7, wherein: When the spliced graph is a heterogeneous mixed spliced graph, the spliced graph information further includes a second syntax element; and the expression format of the homogeneous blocks in the spliced graph is determined according to the second syntax element.
9. The method according to claim 8, wherein The determining, according to the second syntax element, an expression format of the isomorphic blocks in the spliced graph includes: When the value of the second syntax element of the i-th block is the sixth preset value, it is determined that the expression format of the i-th block is the first expression format; When the value of the second syntax element of the i-th block is the seventh preset value, it is determined that the expression format of the i-th block is the second expression format.
10. The method according to claim 1, wherein When the mosaic graph is a heterogeneous mixed mosaic graph, the mosaic graph information includes at least two types of isomorphic block information, wherein isomorphic blocks in different expression formats correspond to different isomorphic block information.
11. The method according to claim 10, wherein: When the mosaic is a heterogeneous mixed mosaic, at least two isomorphic blocks and isomorphic block information are obtained according to the mosaic and the mosaic information, including: When the mosaic graph is a heterogeneous mixed mosaic graph, splitting the mosaic graph to obtain at least two isomorphic blocks; According to the expression formats of the at least two isomorphic blocks, isomorphic block information corresponding to the at least two isomorphic blocks is obtained from the splicing graph information.
12. The method according to claim 10 or 11, wherein: The homogeneous block information includes syntax elements of ASPS and syntax elements of AFPS; Different isomorphic block information corresponds to different syntax elements of the ASPS and syntax elements of the AFPS.
13. The method according to claim 12, wherein: When the isomorphic block is a multi-view video block, it corresponds to first isomorphic block information; when the isomorphic block is a point cloud block, it corresponds to second isomorphic block information; The first homogeneous block information and the second homogeneous block information include shared syntax elements of the ASPS parameter set and syntax elements of the AFPS parameter set; The first homogeneous block information further includes an extended syntax element of the ASPS parameter set and an extended syntax element of the AFPS parameter set.
14. The method according to any one of claims 1 to 13, wherein: The parameter set sub-codestream of the codestream includes a third syntax element, and the codestream corresponding to the visual media content including at least one expression format in the codestream is determined according to the third syntax element.
15. The method according to claim 14, wherein The determining, according to the third syntax element, a codestream corresponding to the visual media content including at least one expression format in the codestream includes: The third syntax element is a first value, and determines that the code stream includes both a code stream corresponding to the visual media content in the first expression format and a code stream corresponding to the visual media content in the second expression format; The third syntax element is a second value, which determines that the code stream includes a code stream corresponding to the visual media content in the first expression format; The third syntax element is a third value, which determines the code stream corresponding to the visual media content in the second expression format included in the code stream.
16. The method according to claim 14, wherein The decoding bitstream to obtain the splicing graph and splicing graph information includes: A code stream corresponding to visual media content in at least two expression formats in the code stream is determined according to the second syntax element, and the code stream is decoded to obtain a heterogeneous mixed splicing graph and splicing graph information.
17. The method according to any one of claims 1 to 16, wherein: The code stream includes a video compression sub-code stream and a splicing graph information sub-code stream. The decoding of the code stream to obtain the splicing graph and the splicing graph information includes: Decoding the video compression sub-stream to obtain the splicing graph; The splicing graph information sub-code stream is decoded to obtain the splicing graph information.
18. The method according to any one of claims 1 to 17, wherein: The expression format is multi-view video, point cloud or mesh.
19. A coding method, wherein: include: Processing visual media content in at least two expression formats to obtain at least two isomorphic blocks; Splicing the at least two isomorphic blocks to obtain a spliced graph and spliced graph information, wherein when the spliced graph is a heterogeneous mixed spliced graph, the spliced graph includes at least two isomorphic blocks, and different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information; The splicing graph and the splicing graph information are encoded to obtain a code stream.
20. The method according to claim 19, wherein The spliced graph information includes a first syntax element, and it is determined according to the first syntax element whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph.
21. The method according to claim 20, wherein The first syntax element includes a first sub-syntax element and a second sub-syntax element; The determining, according to the first syntax element, whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph includes: If the value of the first sub-syntax element is equal to the value of the second sub-syntax element, it is determined according to the values whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph.
22. The method according to claim 21, wherein The determining, according to the value, whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph includes: If the value is a first preset value, it is determined that the spliced graph is a heterogeneous mixed spliced graph; If the value is a second preset value, it is determined that the spliced graph is a homogeneous spliced graph.
23. The method according to claim 21, wherein The determining, according to the value, whether the spliced graph is a heterogeneous mixed spliced graph or a homogeneous spliced graph includes: If the value is a third preset value, it is determined that the spliced graph is a heterogeneous mixed spliced graph including homogeneous blocks in a first expression format and a second expression format, wherein the first expression format and the second expression format are different expression formats; If the value is a fourth preset value, it is determined that the spliced graph is an isomorphic spliced graph including isomorphic blocks in the first expression format; If the value is a fifth preset value, it is determined that the spliced graph is an isomorphic spliced graph including isomorphic blocks in the second expression format.
24. The method according to claim 21, wherein The first sub-syntax element is a syntax element of an ASPS (joint picture sequence parameter set), and the second sub-syntax element is a syntax element of an AFPS (joint picture frame parameter set).
25. The method according to any one of claims 20 to 24, wherein: The method further comprises: The splicing graph information does not include the first syntax element, and the splicing graph is determined to be a homogeneous splicing graph.
26. The method according to any one of claims 19 to 25, wherein: When the spliced graph is a heterogeneous mixed spliced graph, the spliced graph information further includes a second syntax element; and the expression format of the homogeneous blocks in the spliced graph is determined according to the second syntax element.
27. The method according to claim 26, wherein The determining, according to the second syntax element, an expression format of the isomorphic blocks in the spliced graph includes: When the value of the second syntax element of the i-th block is the sixth preset value, it is determined that the expression format of the i-th block is the first expression format; When the value of the second syntax element of the i-th block is the seventh preset value, it is determined that the expression format of the i-th block is the second expression format.
28. The method of claim 19, wherein: When the mosaic graph is a heterogeneous mixed mosaic graph, the mosaic graph information includes at least two types of isomorphic block information, wherein isomorphic blocks in different expression formats correspond to different isomorphic block information.
29. The method according to claim 28, wherein The homogeneous block information includes syntax elements of ASPS and syntax elements of AFPS; Different isomorphic block information corresponds to different syntax elements of the ASPS and syntax elements of the AFPS.
30. The method according to claim 29, wherein When the isomorphic block is a multi-view video block, it corresponds to first isomorphic block information; when the isomorphic block is a point cloud block, it corresponds to second isomorphic block information; The first homogeneous block information and the second homogeneous block information include shared syntax elements of the ASPS parameter set and syntax elements of the AFPS parameter set; The first homogeneous block information further includes an extended syntax element of the ASPS parameter set and an extended syntax element of the AFPS parameter set.
31. The method according to any one of claims 19 to 30, wherein: The parameter set sub-codestream of the codestream includes a third syntax element, and the codestream corresponding to the visual media content including at least one expression format in the codestream is determined according to the third syntax element.
32. The method according to claim 31, wherein The determining, according to the third syntax element, a codestream corresponding to the visual media content including at least one expression format in the codestream includes: The third syntax element is a first value, and determines that the code stream includes both a code stream corresponding to the visual media content in the first expression format and a code stream corresponding to the visual media content in the second expression format; The third syntax element is a second value, which determines that the code stream includes a code stream corresponding to the visual media content in the first expression format; The third syntax element is a third value, which determines the code stream corresponding to the visual media content in the second expression format included in the code stream.
33. The method according to claim 31, wherein When the spliced graph is a heterogeneous mixed spliced graph, the bitstream corresponding to the visual media content in at least two expression formats in the bitstream is determined according to the third syntax element.
34. The method according to any one of claims 19 to 33, wherein: The encoding of the splicing graph and the splicing graph information to obtain a code stream includes: Encoding the spliced image to obtain a video compression sub-stream; Encoding the splicing image information to obtain a splicing image information sub-stream; The video compression sub-stream and the splicing graph information sub-stream are synthesized into the stream.
35. The method according to any one of claims 19 to 34, wherein: The expression format is multi-view video, point cloud or mesh.
36. A decoding device, wherein: include: a decoding unit configured to decode the code stream to obtain a splicing graph and splicing graph information; a splitting unit configured to obtain at least two types of isomorphic blocks and isomorphic block information based on the mosaic and the mosaic information when the mosaic is a heterogeneous mixed mosaic; wherein different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information; The splitting unit is configured to obtain an isomorphic block and isomorphic block information according to the splicing graph and the splicing graph information when the splicing graph is a isomorphic splicing graph; The processing unit is configured to obtain visual media content in at least two expression formats according to the isomorphic blocks and the isomorphic block information.
37. An encoding device, wherein: include: a processing unit configured to process visual media content in at least two expression formats to obtain at least two isomorphic blocks; a splicing unit configured to splice the at least two isomorphic blocks to obtain a spliced graph and spliced graph information, wherein when the spliced graph is a heterogeneous mixed spliced graph, the spliced graph includes at least two isomorphic blocks, and different types of isomorphic blocks correspond to different visual media content expression formats and different isomorphic block information; The encoding unit is configured to encode the splicing graph and the splicing graph information to obtain a code stream.
38. A decoder, wherein The decoder comprises: a first memory and a first processor; The first memory stores a computer program that can be run on the first processor, and when the first processor executes the program, the decoding method according to any one of claims 1 to 18 is implemented.
39. An encoder, wherein The encoder comprises: a second memory and a second processor; The second memory stores a computer program that can be run on the second processor, and when the second processor executes the program, the encoding method according to any one of claims 19 to 35 is implemented.
40. A computer-readable storage medium, wherein: The computer-readable storage medium stores a computer program, and when the computer program is executed by a first processor, it implements the decoding method described in any one of claims 1 to 18; or, when the computer program is executed by a second processor, it implements the encoding method described in any one of claims 19 to 35.
41. A code stream, wherein The code stream is generated based on the method according to any one of claims 19 to 35.