3D data encoder and 3D data decoder
Patent Information
- Application Number
- JP2022125075
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-08-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing 3D data encoding methods, such as HEVC and VVC, suffer from distortion in depth and color images during encoding, leading to reduced accuracy and quality of reconstructed 3D data, and deep learning post filters do not adequately address coding distortions, especially in regions with and without corresponding point clouds.
A 3D data decoding device and encoding device that utilize geometry and attribute image decoding units, along with occupancy map decoding, and filter processing units to perform filter processing on geometry and attribute images using filter parameters derived from occupancy maps, enhancing the accuracy of 3D data reconstruction.
Reduces encoding distortion and improves the quality of 3D data by using filter processing based on occupancy maps, resulting in higher quality encoding and decoding of 3D data.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] An embodiment of the present invention relates to a 3D data encoding device and a 3D data decoding device. [Background technology]
[0002] In order to efficiently transmit or record 3D data, there are 3D data encoding devices that convert 3D data into 2D images and encode them using a video encoding method to generate encoded data, and 3D data decoding devices that decode and reconstruct 2D images from the encoded data to generate 3D data. There is also a technology that performs filtering on 2D images using additional information from a deep learning post-filter.
[0003] Specific examples of 3D data encoding methods include MPEG-I V3C (Volumetric Video-based Coding) and V-PCC (Video-based Point Cloud Compression) (Non-Patent Document 1). V3C can encode and decode multi-viewpoint video in addition to point clouds consisting of point positions and attribute information. Existing video encoding methods include H.266 / VVC and H.265 / HEVC. Additional information of deep learning post-filters includes, for example, a technology that transmits neural network information using the MPEG NNC (Neural network coding) standard (Non-Patent Document 2). [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] ISO / IEC 23090-5 [Non-Patent Document 2] “Additional SEI messages for VSEI (Draft 1)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO / IEC JTC 1 / SC 29, JVET-Z2006, 2022 Summary of the Invention [Problem to be solved by the invention]
[0005] In the 3D data coding method in Non-Patent Document 1, the geometry (depth image) and attributes (color image) constituting the 3D data are coded and decoded using a video coding method such as HEVC and VVC. However, there is a problem that the accuracy and image quality of the reconstructed 3D data are reduced due to distortion of the depth image and color image caused by coding. In addition, there is a problem that the use of additional information in the deep learning post-filter in Non-Patent Document 2 is not sufficient to remove the coding distortion. Specifically, the depth image and color image have an area (valid area) where a corresponding point cloud exists and an area (invalid area) where a corresponding point cloud does not exist. In addition, there is a problem that the effect of the deep learning post-filter is reduced when the characteristics of the valid area and the invalid area are different.
[0006] An object of the present invention is to reduce coding distortion and to code and decode 3D data with high quality when encoding and decoding 3D data using a video encoding method. [Means for solving the problem]
[0007] In order to solve the above problems, a 3D data decoding device according to one embodiment of the present invention is a 3D data decoding device that decodes 3D encoded data, and includes a geometry image decoding unit that decodes a geometry image from the 3D encoded data, an attribute image decoding unit that decodes an attribute image from the 3D encoded data, an occupancy map decoding unit that decodes an occupancy map from the 3D encoded data, a geometry image filter unit that performs filtering of the geometry image, and an attribute image filter unit that performs filtering of the attribute image, wherein the geometry image filter unit and / or the attribute image filter unit decodes filter parameters including information indicating the type of V3C image, such as a geometry image or an attribute image, of an image to be filtered, and performs filtering of the geometry image and / or the attribute image using the filter parameters and the occupancy map.
[0008] In order to solve the above problems, a 3D data encoding device according to one embodiment of the present invention is a 3D data encoding device that encodes 3D data, and includes a geometry image filter parameter derivation unit that derives filter parameters for a geometry image of the 3D data, an attribute image filter parameter derivation unit that derives filter parameters for an attribute image of the 3D data, a geometry image encoding unit that encodes the geometry image, an attribute image encoding unit that encodes the attribute image, and an occupancy map encoding unit that encodes the occupancy map, and the geometry image filter parameter derivation unit and / or the attribute image filter parameter derivation unit use the occupancy map to encode filter parameters including information indicating that the filtering target is a 3D data (V3C) component. Effect of the Invention
[0009] According to one aspect of the present invention, distortion caused by encoding a depth image or a color image can be reduced, and 3D data can be encoded and decoded with high quality. [Brief description of the drawings]
[0010] [Figure 1] 1 is a schematic diagram showing the configuration of a 3D data transmission system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a hierarchical structure of data in an encoded stream. [Diagram 3] 1A to 1C are diagrams for explaining 3D data, an occupancy map, a geometry image, and an attribute image. [Figure 4] FIG. 13 is a diagram for explaining the layer structure of a geometry image and an attribute image. [Diagram 5] FIG. 13 is a diagram for explaining a process of converting a 1-channel image into a 4-channel image. [Figure 6] FIG. 2 is a functional block diagram showing a schematic configuration of a 3D data decoding device 31 according to the first embodiment. [Figure 7] FIG. 1 is a functional block diagram showing a schematic configuration of a 3D data encoding device 11 according to a first embodiment. [Figure 8] 13 is an example of a filter shape applied by the 3D-ALF filter unit 610. [Figure 9] 13 is an example of a syntax for a configuration in which filter parameters are transmitted in SEI. [Figure 10] 13 is an example of a syntax for a configuration in which filter parameters are transmitted in SEI. [Figure 11] 13 is an example of a syntax for a configuration in which filter parameters are transmitted in SEI. [Figure 12] 13 is an example of a syntax for transmitting filter parameters by ASPS. [Figure 13] 13 is an example of a syntax for transmitting filter parameters by ASPS. [Figure 14] 13 shows an example of input / output data of the 3D-NN filter unit 611 and its format. [Figure 15] FIG. 13 is a diagram for explaining a process of performing error calculation using an output image from the 3D-NN filter unit 611 and an occupancy map. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.
[0012] FIG. 1 is a schematic diagram showing the configuration of a 3D data transmission system 1 according to this embodiment.
[0013] The 3D data transmission system 1 is a system that transmits an encoded stream obtained by encoding 3D data to be encoded, decodes the transmitted encoded stream, and displays the 3D data. The 3D data transmission system 1 includes a 3D data encoding device 11, a network 21, a 3D data decoding device 31, and a 3D data display device 41.
[0014] 3D data T is input to the 3D data encoding device 11.
[0015] The network 21 transmits the encoded stream Te generated by the 3D data encoding device 11 to the 3D data decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Furthermore, the network 21 may be replaced by a storage medium on which the encoded stream Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).
[0016] The 3D data decoding device 31 decodes each of the encoded streams Te transmitted by the network 21, and generates one or more decoded 3D data Td.
[0017] The 3D data display device 41 displays all or part of one or more pieces of decoded 3D data Td generated by the 3D data decoding device 31. The 3D data display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Display forms include stationary, mobile, HMD, and the like. Furthermore, when the 3D data decoding device 31 has high processing power, it displays high quality images, and when it has only lower processing power, it displays images that do not require high processing power or display power.
[0018] <Structure of the coding stream Te> The data structure of the coded stream Te generated by the 3D data coding device 11 and decoded by the 3D data decoding device 31 will be described.
[0019] Fig. 2 is a diagram showing a hierarchical structure of data in an encoded stream Te. The encoded stream Te illustratively includes a sequence and a plurality of pictures constituting the sequence. Fig. 2 shows diagrams showing an encoded video sequence that defines a sequence SEQ, an encoded picture that defines a picture PICT, an encoded slice that defines a slice S, encoded slice data that defines slice data, an encoding tree unit included in the encoded slice data, and an encoding unit included in the encoding tree unit.
[0020] (Coded Video Sequence) In the coded video sequence, a set of data that the 3D data decoding device 31 refers to in order to decode the sequence SEQ to be processed is specified. As shown in the coded video sequence of Fig. 2, the sequence SEQ includes a video parameter set, a sequence parameter set SPS (Sequence Parameter Set), a picture parameter set PPS (Picture Parameter Set), a picture PICT, and supplemental enhancement information SEI (Supplemental Enhancement Information). Note that the 3D data may further include another parameter set.
[0021] The video parameter set VPS specifies a set of coding parameters common to multiple videos composed of multiple layers, as well as a set of coding parameters related to multiple layers and each individual layer included in the video.
[0022] The sequence parameter set SPS specifies a set of coding parameters that the 3D data decoding device 31 refers to in order to decode the target sequence. For example, the width and height of a picture are specified. Note that there may be multiple SPSs. In that case, one of the multiple SPSs is selected from the PPS.
[0023] The picture parameter set PPS specifies a set of coding parameters that the video decoding device 31 refers to in order to decode each picture in the target sequence. For example, the picture parameter set PPS includes a reference value of the quantization width (pic_init_qp_minus26) used in decoding the picture and a flag (weighted_pred_flag) indicating the application of weighted prediction. Note that there may be multiple PPSs. In that case, one of the multiple PPSs is selected for each picture in the target sequence.
[0024] (Encoded Picture) The coded picture defines a set of data to be referenced by the 3D data decoding device 31 in order to decode the picture PICT to be processed. As shown in the coded picture of Fig. 2, the picture PICT includes slices 0 to NS-1 (NS is the total number of slices included in the picture PICT).
[0025] In the following description, when there is no need to distinguish between slices 0 to NS-1, the subscripts of the symbols may be omitted. The same applies to other data that are included in the coded stream Te and that are to be described below and that are to be given subscripts.
[0026] (Coded Slice) An encoded slice specifies a set of data to be referenced by the 3D data decoding device 31 in order to decode a slice S to be processed. As shown in the encoded slice of Fig. 2, a slice includes a slice header and slice data.
[0027] The slice header includes a group of coding parameters to be referred to by the 3D data decoding device 31 in order to determine a decoding method for the target slice. Slice type designation information (slice_type) that designates the slice type is an example of a coding parameter included in the slice header.
[0028] Slice types that can be specified by the slice type specification information include (1) an I slice that uses only intra prediction when encoding, (2) a P slice that uses unidirectional prediction or intra prediction when encoding, and (3) a B slice that uses unidirectional prediction, bidirectional prediction, or intra prediction when encoding. Note that inter prediction is not limited to uni-prediction or bi-prediction, and a predicted image may be generated using more reference pictures. Hereinafter, when referring to P or B slice, it refers to a slice including a block that can use inter prediction.
[0029] In addition, the slice header may include a reference to a picture parameter set PPS (pic_parameter_set_id).
[0030] (Encoded slice data) The coded slice data specifies a set of data to be referenced by the 3D data decoding device 31 in order to decode the slice data to be processed. The slice data includes a CTU, as shown in the coded slice header in Fig. 2. A CTU is a block of a fixed size (e.g., 64x64) that constitutes a slice, and is also called a Largest Coding Unit (LCU).
[0031] (Subpicture) A picture may be further divided into rectangular sub-pictures. For example, it may be divided into four sub-pictures horizontally and four sub-pictures vertically. The size of a sub-picture may be a multiple of a CTU. A sub-picture is defined as a set of an integer number of consecutive tiles vertically and horizontally. The slice header may include sh_subpic_id, which indicates the ID of the sub-picture.
[0032] (coding tree unit) In the coding tree unit in Fig. 2, a set of data to be referred to by the 3D data decoding device 31 in order to decode the CTU to be processed is specified. The CTU is divided into coding units CU, which are basic units of the coding process, by recursive quad tree division (QT (Quad Tree) division), binary tree division (BT (Binary Tree) division), or ternary tree division (TT (Ternary Tree) division). BT division and TT division are collectively called multi tree division (MT (Multi Tree) division). A node of a tree structure obtained by recursive quad tree division is called a coding node. Intermediate nodes of the quad tree, binary tree, and ternary tree are coding nodes, and the CTU itself is also specified as the top coding node.
[0033] (Encoding Unit) As shown in the coding unit of Fig. 2, a set of data to be referenced by the 3D data decoding device 31 in order to decode the coding unit to be processed is defined. Specifically, the CU is composed of a CU header CUH, prediction parameters, transformation parameters, quantization transformation coefficients, etc. The CU header defines a prediction mode, etc.
[0034] The prediction process may be performed in units of CUs, or in units of sub-CUs obtained by further dividing a CU. When the size of a CU and a sub-CU are equal, there is one sub-CU in the CU. When the size of a CU is larger than that of a sub-CU, the CU is divided into sub-CUs. For example, when a CU is 8x8 and a sub-CU is 4x4, the CU is divided into 2 parts horizontally and 2 parts vertically, into 4 sub-CUs.
[0035] There are two types of prediction (prediction modes): intra prediction and inter prediction. Intra prediction is a prediction within the same picture, while inter prediction refers to a prediction process performed between different pictures (for example, between display times or between layer images).
[0036] The transform and quantization processes are performed in units of CUs, but the quantized transform coefficients may be entropy coded in units of sub-blocks such as 4x4.
[0037] (Prediction parameters) The predicted image is derived from prediction parameters associated with the block, which include intra-prediction and inter-prediction parameters.
[0038] (Data structure of 3D information) Three-dimensional information (3D data) is represented in the form of a point cloud, which is a group of points containing positional information and attribute information in three-dimensional space, or a mesh (or polygon) consisting of the vertices and faces of triangles.
[0039] FIG. 3 is a diagram for explaining 3D data, an occupancy map, a geometry image, and an attribute image. The point cloud and mesh constituting the 3D data are divided into a plurality of parts (areas) by the 3D data encoding device 11, and the point cloud included in each part is projected onto any plane of a 3D bounding box (FIG. 3(a)) set in a 3D space. The 3D data encoding device 11 generates a plurality of patches from the projected point cloud. Information on the 3D bounding box (coordinates, size, etc.) and information on mapping onto the projection plane (projection plane, coordinates, size, rotation presence / absence, etc. of each patch) are called atlas information. The occupancy map is an image showing the valid area of each patch (area where the point cloud and mesh exist) as a 2D binary image (e.g., valid area is 1, invalid area is 0) (FIG. 3(b)). Note that the values of the valid area and invalid area may be values other than 0 and 1, such as 255 and 0. The geometry image is an image showing the depth value (distance) of each patch with respect to the projection plane (FIG. 3(c)). The relationship between the depth value and the pixel value may be linear, or the distance may be derived from the pixel value by a lookup table, a mathematical formula, or a relational formula based on a combination of branching by values. The attribute image is an image that indicates the attribute of a point (e.g., RGB color). The occupancy map image, geometry image, attribute image, and atlas information may be images in which partial images (patches) from different projection planes are mapped (pasted) onto a two-dimensional image. The atlas information includes information on the number of patches and the projection planes corresponding to the patches. The 3D data decoding device 12 reconstructs the coordinates and attribute information of the point cloud or mesh from the atlas information, occupancy map, geometry image, and attribute image. Here, a point refers to each point of a point cloud or a vertex of a mesh.
[0040] FIG. 4 is a diagram for explaining the layer structure of a geometry image and an attribute image. Each of the geometry image and the attribute image may be composed of a plurality of images (layers). For example, a near layer and a far layer may be provided. Here, the near layer and the far layer are images constituting geometry and attributes whose depths as viewed from a certain projection surface are different from each other. The near layer may be a collection of points whose depth is minimum for each pixel of the projection surface. The far layer may be a collection of points whose depth is maximum within a predetermined range (for example, a range of a distance d from the near layer) for each pixel of the projection surface.
[0041] The geometry image encoding unit 106 described later may encode the geometry image of the Near layer as an intra-picture (I-picture) and the geometry image of the Far layer as an inter-picture (P-picture or B-picture). Also, the Near layer may be encoded so as to be identifiable on the bit stream using LayerID (nuh_layer_id syntax) of a Network Abstraction Layer (NAL) unit, such as LayerID=0 for the Near layer and LayerID=1 for the Far layer. Also, the Near layer may be encoded so as to be identifiable on the bit stream using TemporalID of the NAL unit, such as TemporalID=0 for the Near layer and TemporalID=1 for the Far layer.
[0042] (4:2:0 3 channel image to 6 channel tensor conversion) This explains how to convert the luminance and chrominance channels in 4:2:0 format to match the difference in resolution. It can also be considered a conversion method from a format with one YUV channel each to a format with four luminance channels and one UV channel each (YYYYUV).
[0043] Figure 5 shows the process of converting a one-channel image into a four-channel image. An input image of width W and height H is Z-scanned in units of 2x2 size and numbered in order from 0, 1, 2, 3, .... It is then converted into a four-channel image of width W / 2 and height H / 2, consisting of a channel for pixels TL numbered 4N, a channel for pixels TR numbered 4N+1, a channel for pixels BL numbered 4N+2, and a channel for pixels BR numbered 4N+3.
[0044] inputTensor[0][0][y][x]= inSamplesTL = inSamples[0][y*2 ][x*2 ] inputTensor[0][1][y][x]= inSamplesTR = inSamples[0][y*2 ][x*2+1] inputTensor[0][2][y][x]= inSamplesBL = inSamples[0][y*2+1][x*2 ] inputTensor[0][3][y][x]= inSamplesBR = inSamples[0][y*2+1][x*2+1]
[0045] You can also convert a 3-channel image to a 6-channel image. Convert the 4:2:0 format image inSamples[ch][y][x] into a tensor inputTensor with multiple luminance and chrominance channels of the same size, allowing processing at the same time.
[0046] inputTensor[0][0][y][x]= inSamplesTL = inSamples[0][y*2 ][x*2 ] inputTensor[0][1][y][x]= inSamplesTR = inSamples[0][y*2 ][x*2+1] inputTensor[0][2][y][x]= inSamplesBL = inSamples[0][y*2+1][x*2 ] inputTensor[0][3][y][x]= inSamplesBR = inSamples[0][y*2+1][x*2+1] inputTensor[0][4][y][x]= inSamplesCb = inSamples[1][y ][x ] inputTensor[0][5][y][x]= inSamplesCr = inSamples[2][y ][x ]
[0047] (Configuration of 3D data decoding device according to the first embodiment) FIG. 6 is a functional block diagram showing a schematic configuration of a 3D data decoding device 31 according to the first embodiment.
[0048] The 3D data decoding device 31 is made up of a header decoding unit 301, an atlas information decoding unit 302, an occupancy map decoding unit 303, a geometry image decoding unit 304, a geometry image filter unit 305 (3D-CC filter unit 511), a geometry reconstruction unit 306, an attribute image decoding unit 307, an attribute image filter unit 308 (3D-CC filter unit 511), an attribute reconstruction unit 309, and a 3D data reconstruction unit 310. The processing of the geometry image filter unit 305 and the attribute image filter unit 308 may be post-filtering.
[0049] The header decoding unit 301 inputs encoded data multiplexed in a byte stream format, ISOBMFF (ISO Base Media File Format), etc., demultiplexes it, and outputs an atlas information encoded stream, an occupancy map encoded stream, a geometry image encoded stream, an attribute image encoded stream, and filter parameters.
[0050] The header decoding unit 301 decodes pfp_id, pfp_purpose, 3d_geometry_filter_enabled_flag, and 3d_attribute_filter_enabled_flag from the encoded data (SEI, ASPS, etc.).
[0051] The atlas information decoding unit 302 receives the atlas information encoded stream and decodes the atlas information.
[0052] The occupancy map decoding unit 303 decodes an occupancy map coded stream coded using VVC, HEVC, or the like, and outputs an occupancy map.
[0053] The geometry image decoding unit 304 decodes a geometry image coded stream coded using VVC, HEVC, or the like, and outputs a geometry image.
[0054] The geometry image filter unit 305 inputs a geometry image and filter parameters (for example, the type of target component, the target layer, the filter coefficients of the linear filter ALF, and the neural network model). The geometry image filter unit 305 includes a 3D-CC filter unit 511, performs filtering based on the geometry image and the filter parameters, and outputs a filtered image of the geometry image.
[0055] The geometry reconstruction unit 306 receives the atlas information, the occupancy map, and the (filtered) geometry image, and reconstructs the geometry (depth information) in the 3D space.
[0056] The attribute image decoding unit 307 decodes an encoded stream encoded by VVC, HEVC, or the like, receives the attribute encoded stream, and outputs an attribute image.
[0057] The attribute image and the above filter parameters are input to the attribute image filter unit 308. The attribute image filter unit 308 includes a 3D-CC filter unit 511, which performs filtering based on the attribute image and the filter parameters and outputs a filtered image of the attribute image.
[0058] The attribute reconstruction unit 309 receives the atlas information, the occupancy map, and the attribute image, and reconstructs the attributes (color information) in the 3D space.
[0059] The 3D data reconstruction unit 310 reconstructs 3D point cloud data or mesh data based on the reconstructed geometry information and attribute information.
[0060] (3D-CC filter section 511) The 3D-CC filter unit 511 performs filtering on the geometry image and the attribute image (target image inputTensor) using the filter parameters.
[0061] The 3D-CC filter unit 511 uses the occupancy map occFrame[y][x] and filter parameters to perform filtering on compFrame[y][x] and generate outFrame[cIdx][y][x]. compFrame is the geometry image geoFrame and / or the attribute image attFrame. outFrame is the geometry image outgeoFrame after filtering and / or the attribute image outattFrame after filtering. The 3D-CC filter unit 511 includes a 3D-ALF filter unit 610 or a 3D-NN filter unit 611. The 3D-ALF filter unit 610 decodes filter coefficients derived by the 3D data encoding device and performs adaptive filtering. The 3D-NN filter unit 611 performs filtering using a neural network. These will be described in detail later.
[0062] (3D-ALF filter section 610) The 3D-ALF filter unit 610 divides the filter target image compFrame into small regions of a certain size (e.g., 4x4) (x=xSb..xSb+bSW-1, y=xSb..xSb+bSH-1, where (xSb, ySb) are the top left coordinates of the small region, and bSW, bSH are the width and height of the small region). Then, for each small region, it derives the directionality (e.g., 5 directional classes) and activity (e.g., 5 stages) of the small region, and classifies the small region into one of multiple classes (e.g., 25 classes) based on these. Then, it performs filtering using a filter corresponding to the classification classId, and outputs the processed image.
[0063] The 3D-ALF filter unit 610 may switch the filter shape and filter coefficients depending on whether the image to be filtered is a geometry image or an attribute image, or depending on the components of the image to be filtered. Also, the filter coefficients may be switched for each small region depending on whether the region is a valid region or an invalid region, with reference to the occFrame. For example, a 7x7 diamond filter as shown in FIG. 8 may be applied to the luminance image (Y image) of the attribute image, and a 5x5 diamond filter may be applied to the color difference image (Cb, Cr image). These filter coefficients may be fixed filters provided in advance in the moving 3D data decoding device and the 3D data encoding device, or filter coefficients decoded from the encoded data.
[0064] For example, when classifying each 4x4 small region into 25 classes, the following processing may be performed.
[0065] First, the 1-D Laplacian is used to calculate the horizontal, vertical, and two diagonal gradients gv, gh, gd1, and gd2, respectively, as follows: gv=ΣΣ|2* compFrame[y][x] - compFrame[y-1][x] - compFrame[y+1][x]| gh=ΣΣ|2* compFrame[y][x] - compFrame[y][x-1] - compFrame[y][x+1]| gd1=ΣΣ|2* compFrame[y][x] - compFrame[y-1][x-1] - compFrame[y+1][x+1]| gd2=ΣΣ|2* compFrame[y][x] - compFrame[y+1][x-1] - compFrame[y-1][x+1]| Using each gradient gv, gh, gd1, and gd2, the maximum value ghvmax and minimum value ghvmin of the vertical and horizontal gradients, and the maximum value gd1d2max and minimum value gd1d2min of the gradients in the two diagonal directions are calculated as follows: ghvmax = max(gv, gh) ghvmin = min(gv, gh) gd1d2max = max(gd1, gd2) gd1d2min = min(gd1, gd2) Then, the directionality D is calculated using predetermined thresholds t1 and t2 as follows.
[0066] If ghvmax<=t1*ghvmin and gd1d2max<=t1*gd1d2min, D=0 If ghvmax / ghvmin>gd1d2max / gd1d2min and ghvmax<=t2*ghvmin, D=1 If ghvmax / ghvmin>gd1d2max / gd1d2min and ghvmax>t2*ghvmin, D=2 If ghvmax / ghvmin<=gd1d2max / gd1d2min and gd1d2max<=t2*gd1d2min, D=3 If ghvmax / ghvmin<=gd1d2max / gd1d2min and gd1d2max>t2*gd1d2min, D=4 Activity A is as follows: A = ΣΣ(Vxy+Hxy) Using activity A and direction D, the class index classId is calculated as follows: classId=5D+A
[0067] The 3D-ALF filter unit 610 selects a filter coefficient specified by classId from the transmitted set of filter coefficients coeff[][], and performs filtering on the inputTensor derived from compFrame using the selected filter coefficient coeff[classId][].
[0068] outFrame[nn][yy][xx] = ΣΣΣ(coeff[classId][i][j]*inputTensor[mm][yy+j-of][xx+i-of]+bias[nn]) Here, Σ denotes the sum of i=xSb..xSb+bSW-1, j=xSb..xSb+bSH-1, and mm=0..m-1 (m is the number of channels of the inputTensor).
[0069] As will be described later, the inputTensor may be an image composed of attFrame and occFrame, or an image composed of geoFrame, attFrame, and occFrame.
[0070] (3D-NN filter section 611) The 3D-NN filter unit 611 may change the neural network model or the weighting coefficients and bias coefficients constituting the model depending on whether the image to be filtered is a geometry image or an attribute image. The neural network model has, for example, the number of convolutions, the number of layers, the kernel size, and a topology such as a connection relationship.
[0071] Here, the neural network model (hereinafter, NN model) means the elements and connection relationships (topology) of the neural network, and the parameters (weights, biases) of the neural network. Note that the 3D-NN filter unit 611 may fix the topology and switch only the parameters according to the image to be filtered.
[0072] The 3D-NN filter unit 611 performs filtering by a neural network model using the input image CompFrame and input parameters (e.g., QP, bS, etc.). The input image may be an image for each component, or an image having multiple components as channels. The input parameters may be assigned to a channel different from the image.
[0073] The 3D-NN filter unit 611 may repeatedly apply the following process.
[0074] The 3D-NN filter unit 611 performs convolution (conv, convolution) on kernels k[m][i][j] as many times as the number of layers using compFrame as inputTensor, and generates an output image outFrame by adding bias, where nn=0..n-1, xx=0..width-1, and yy=0..height-1.
[0075] Here, in each layer, an outputTensor is generated from an inputTensor.
[0076] outputTensor[nn][xx][yy]=ΣΣΣ(k[mm][i][j]*inputTensor[mm][xx+i-of][yy+j-of]+bias[nn])
[0077] In the next layer, the obtained outputTensor is used as a new inputTensor, and the same process is repeated for the number of layers. An activation layer may be placed between layers. A pooling layer or skip connection may also be used. Finally, the outFrame is derived from the obtained outputTensor.
[0078] In the case of 1x1 Conv, Σ represents the sum of mm=0..m-1, i=0, j=0. In this case, of=0 is set. In the case of 3x3 Conv, Σ represents the sum of mm=0..m-1, i=0..2, j=0..2. In this case, of=1 is set. n is the number of channels of outputTensor (1), m is the number of channels of inputTensor, width is the width of inputTensor and outputTensor, and height is the height of inputTensor and outputTensor. of is the width or height of the area required around inputTensor to generate outFrame. Hereinafter, when the output of the 3D-NN filter unit 611 is a value (correction value) rather than an image, the output is represented by corrNN instead of outFrame.
[0079] Note that if you write inputTensor and outputTensor in CHW format instead of inputTensor and outFrame in CWH format, it is equivalent to the following process.
[0080] outputTensor[nn][yy][xx]=ΣΣΣ(k[mm][i][j]*inputTensor[mm][yy+j-of][xx+i-of]+bias[nn])
[0081] Alternatively, a process called Depth wise Conv may be performed, which is expressed by the following formula: Here, nn=0..n-1, xx=0..width-1, and yy=0..height-1.
[0082] outputTensor[nn][xx][yy]=ΣΣ(k[nn][i][j]*inputTensor[nn][yy+j-of][xx+i-of]+bias[nn])
[0083] Also, a nonlinear process called Activate, for example, ReLU, may be used. ReLU(x) = x >= 0 ? x : 0 Alternatively, leakyReLU shown in the following formula may be used.
[0084] leakyReLU(x) = x >= 0 ? x : a * x Here, a is a predetermined value less than 1, for example, 0.1 or 0.125. In order to perform integer arithmetic, all the above values of k, bias, and a may be set to integers, and a right shift may be performed after conv to generate the outputTensor.
[0085] With ReLU, 0 is always output for values less than 0, and the input value is output as is for values greater than or equal to 0. On the other hand, with leakyReLU, linear processing is performed for values less than 0 with the gradient set by a. With ReLU, the gradient for values less than 0 disappears, which can make it difficult for learning to progress. With leakyReLU, the gradient for values less than 0 remains, making the above problem less likely to occur. Of the above leakyReLU(x), PReLU, which uses a parameterized value of a, may also be used.
[0086] (Basic configuration of 3D-NN filter) The 3D-NN filter unit 611 may include a size conversion unit that converts the occupancy map into the same resolution as the image to be filtered, and an NN processing unit that performs filtering using NN model parameters based on the NN model.
[0087] (Size conversion section) The size conversion unit included in the 3D-NN filter unit 611 converts the resolution of the occupancy map to the same size as the size of the image to be filtered using the resolution ratios occFrameWidthRatio and occFrameHeightRatio between the occupancy map occFrame and the image to be filtered (attFrame or geoFrame).
[0088] OccFrame[y][x] = occFrame[y / occFrameWidthRatio][x / occFrameHeightRatio] Note that the following may be derived assuming that the horizontal and vertical ratios are equal (occFrameizeRatio = occFrameWidthRatio = occFrameHeightRatio).
[0089] OccFrame[y][x] = occFrame[y / occFrameizeRatio][x / occFrameizeRatio] If the sizes are the same, the following may be used:
[0090] OccFrame[y][x] = occFrame[y][x] In addition, when the image to be filtered is 4:2:0 and has channels in YYYYUV format, the size conversion unit included in the 3D-NN filter unit 611 may perform resolution conversion processing at the same resolution as the color difference image of the image to be filtered (half the luminance resolution).
[0091] OccFrame[cy][cx] = occFrame[(2*cy) / occFrameWidthRatio][(2*cx) / occFrameHeightRatio)] Alternatively, it may be derived using the following formula:
[0092] OccFrame[cy][cx] = occFrame[cy / (occFrameWidthRatio / 2)][cx / (occFrameHeightRatio / 2)] Alternatively, when deriving the resolution ratio, occFrameWidthRatio and occFrameHeightRatio may be set to 1 / 2.
[0093] If the occupancy map has a resolution lower than the image to be filtered, the occupancy map may use a nearest neighbor algorithm or the like. In addition, an interpolation may be used in which filter coefficients are derived from decimal precision positions (phases) and weighted averages are taken with pixel values. Filter coefficients such as 2-tap linear prediction, 4-tap bicubic, and 4-tap or more Lanczos may be used.
[0094] The 3D-NN filter unit 611 may derive occFrameWidthRatio and occFrameHeightRatio from the ratio between the target image and the occupancy map. For example, when the image to be filtered is a geometry image, they are derived as follows.
[0095] occFrameWidthRatio = geoFrameWidth / occFrameWidth occFrameHeightRatio = geoFrameHeight / occFrameHeight Here, geoFrameWidth, geoFrameHeight are the width and height of the geometry image. For example, if the target image is an attribute image, it is derived as follows.
[0096] attFrameWidthRatio = attFrameWidth / occFrameWidth attFrameHeightRatio = attFrameHeight / occFrameHeight Here, attFrameWidth, attFrameHeight are the width and height of the attribute image.
[0097] The ratio of the target image to the occupancy map may be derived to a fixed integer precision. occFrameWidthRatio = (geoFrameWidth * R_UNIT) / occFrameWidth occFrameHeightRatio = (geoFrameHeight * R_UNIT) / occFrameHeight Here, R_UNIT indicates the precision and is taken as a power of two, such as 1, 2, 4, ...
[0098] Furthermore, the 3D-NN filter unit 611 may derive occFrameWidthRatio and occFrameHeightRatio in accordance with a syntax value (3d_filter_occupancy_scale_idc) in the encoded data. 3d_filter_occupancy_scale_idc is an index indicating the downsampling ratio (resolution ratio) of the occupancy map. occFrameWidthRatio = 1<< 3d_filter_occupancy_scale_idc occFrameHeightRatio = 1<< 3d_filter_occupancy_scale_idc In addition, if the horizontal and vertical ratios of the target image and the occupancy map are different, 3d_filter_occupancy_width_scale_idc and 3d_filter_occupancy_height_scale_idc may be notified as syntax instead of 3d_filter_occupancy_scale_idc. 3d_filter_occupancy_width_scale_idc is the ratio of the width of the target image to the occupancy map, and 3d_filter_occupancy_height_scale_idc is the ratio of the target image to the occupancy map.
[0099] occFrameWidthRatio = 1<< 3d_filter_occupancy_width_scale_idc occFrameHeightRatio = 1<< 3d_filter_occupancy_height_scale_idc
[0100] (Operation of 3D-NN filter unit 611 using occupancy map) The 3D-NN filter unit 611 derives input data inputTensor[][][] using an occupancy map occFrame[][] (or an image OccFrame obtained by size conversion) based on the target image compFrame[][] and the horizontal and vertical chrominance subsampling values SW, SH of the target image decoded by the geometry image decoding unit 304 and / or the attribute image decoding unit 307. In other words, the 3D-NN filter unit 611 uses the occupancy map OccFrame[][] as one channel of inputTensor. where SW=SubWidthC and SG=SubHeightC indicate the subsampling of the color components, here the variables representing the ratio of chrominance resolution to luma resolution.
[0101] The header decoding unit 301 derives the following variables according to chroma_format_idc in the encoded data. SW=SubWidthC = 1, SH=SubHeghtC = 1 (chroma_format_idc == 0) SW=SubWidthC = 2, SH=SubHeghtC = 2 (chroma_format_idc == 1) SW=SubWidthC = 2, SH=SubHeghtC = 1 (chroma_format_idc == 2) SW=SubWidthC = 1, SH=SubHeghtC = 1 (chroma_format_idc == 3)
[0102] The 3D-NN filter unit 611 may derive the inputTensor using a one-channel target image compFrame and a one-channel OccFrame. In the following, compFrame may be either attFrame or geoFrame. In other words, the following compFrame may be replaced with attFrame or geoFrame.
[0103] inputTensor[0][y][x] = compFrame[0][y][x] <occ-inp0> Here, IsOccupancyUsedForFilter is a flag indicating whether to set the occupancy map to inputTensor for use. IsOccupancyUsedForFilter may be switched based on the value decoded from the encoded data, or it may be configured to always use the occupancy map by setting IsOccupancyUsedForFilter to true. If it is false, the occupancy map is not used and OccFrame is not set to inputTensor. True and false may be 1, 0, or other predetermined values.
[0104] The 3D-NN filter unit 611 may derive an inputTensor using a three-channel target image compFrame and a one-channel occupancy map OccFrame to perform filtering. inputTensor[0][y][x] = compFrame[0][y][x] <Formula OCC-INP1> inputTensor[1][y][x] = compFrame[1][y / SH][x / SW] inputTensor[2][y][x] = compFrame[2][y / SH][x / SW] inputTensor[3][y][x] = OccFrame[y][x] (if IsOccupancyUsedForFilter == true)
[0105] Furthermore, the 3D-NN filter unit 611 may expand the luminance of a 4:2:0 image in the channel direction, and derive the inputTensor with the same resolution as the color difference image (channels in YYYYUV format). inputTensor[0][cy][cx] = compFrame[0][cy*2 ][cx*2 ] <Formula OCC-INP2> inputTensor[1][cy][cx] = compFrame[0][cy*2 ][cy*2+1] inputTensor[2][cy][cx] = compFrame[0][cy*2+1][cx*2 ] inputTensor[3][cy][cx] = compFrame[0][cy*2+1][cx*2+1] inputTensor[4][cy][cx] = compFrame[1][cy][cx] inputTensor[5][cy][cx] = compFrame[2][cy][cx] inputTensor[6][y][x] = OccFrame[y][x] (if IsOccupancyUsedForFilter == true) As already explained, compFrame may be attFrame. In this case, it may be derived as follows.
[0106] inputTensor[0][cy][cx] = attFrame[0][cy*2 ][cx*2 ] <Formula OCC-INP2> inputTensor[1][cy][cx] = attFrame[0][cy*2 ][cy*2+1] inputTensor[2][cy][cx] = attFrame[0][cy*2+1][cx*2 ] inputTensor[3][cy][cx] = attFrame[0][cy*2+1][cx*2+1] inputTensor[4][cy][cx] = attFrame[1][cy][cx] inputTensor[5][cy][cx] = attFrame[2][cy][cx] inputTensor[6][y][x] = OccFrame[y][x] (if IsOccupancyUsedForFilter == true)
[0107] Furthermore, when the size of the occupancy map is the same as the luminance size, the 3D-NN filter unit 611 may expand the occupancy map in the channel direction so that the size is also the same as the chrominance size, to derive the inputTensor.
[0108] inputTensor[0][cy][cx] = compFrame[0][cy*2 ][cx*2 ] <Formula OCC-INP3> inputTensor[1][cy][cx] = compFrame[0][cy*2 ][cx*2+1] inputTensor[2][cy][cx] = compFrame[0][cy*2+1][cx*2 ] inputTensor[3][cy][cx] = compFrame[0][cy*2+1][cx*2+1] inputTensor[4][cy][cx] = compFrame[1][cy][cx] inputTensor[5][cy][cx] = compFrame[2][cy][cx] inputTensor[6][cy][cx] = OccFrame[cy*2 ][cx*2 ] (if IsOccupancyUsedForFilter == true) inputTensor[7][cy][cx] = OccFrame[cy*2 ][cx*2+1] (if IsOccupancyUsedForFilter == true) inputTensor[8][cy][cx] = OccFrame[cy*2+1][cx*2 ] (if IsOccupancyUsedForFilter == true) inputTensor[9][cy][cx] = OccFrame[cy*2+1][cx*2+1] (if IsOccupancyUsedForFilter == true) Here, x,y are the coordinates of the luminance pixel. For example, in compFrame[][y][x], the ranges of x and y are x=0..LumaWidth-1, y=0..LumaHeight-1, respectively. cx,cy are the coordinates of the chrominance pixel. The ranges of cx and cy are cx=0..ChromaWidth-1, cy=0..ChromaHeight-1, respectively.
[0109] In the above configuration, since the inputTensor for performing filtering is derived using compFrame and OccFrame, the accuracy of outFrame is improved by the filter. When compFrame is attFrame, the accuracy of attFrame is improved, and when compFrame is geoFrame, the accuracy of geoFrame is improved.
[0110] The 3D-NN filter unit 611 may derive the inputTensor using a one-channel geometry image geoFrame and a one-channel occupancy map OccFrame.
[0111] inputTensor[0][y][x] = geoFrame[0][y][x] <formula OCC-INP1> inputTensor[1][y][x] = occFrame[y][x] (if IsOccupancyUsedForFilter == true)
[0112] Furthermore, the 3D-NN filter unit 611 may expand the luminance of a 4:2:0 image in the channel direction, set it to the same resolution as the color difference image (channels in YYYYUV format), derive an inputTensor, and perform NN filtering.
[0113] inputTensor[0][cy][cx] = geoFrame[0][cy*2 ][cx*2 ] <Formula OCC-INP2> inputTensor[1][cy][cx] = geoFrame[0][cy*2 ][cy*2+1] inputTensor[2][cy][cx] = geoFrame[0][cy*2+1][cx*2 ] inputTensor[3][cy][cx] = geoFrame[0][cy*2+1][cx*2+1]
[0114] Furthermore, when the size of the occupancy map is the same as the luminance size, the 3D-NN filter unit 611 may expand the occupancy map in the channel direction so that the size is also the same as the chrominance size, to derive the inputTensor.
[0115] inputTensor[0][cy][cx] = geoFrame[0][cy*2 ][cx*2 ] <Formula OCC-INP3> inputTensor[1][cy][cx] = geoFrame[0][cy*2 ][cx*2+1] inputTensor[2][cy][cx] = geoFrame[0][cy*2+1][cx*2 ] inputTensor[3][cy][cx] = geoFrame[0][cy*2+1][cx*2+1] inputTensor[4][cy][cx] = OccFrame[cy*2 ][cx*2 ] (if IsOccupancyUsedForFilter == true) inputTensor[5][cy][cx] = OccFrame[cy*2 ][cx*2+1] (if IsOccupancyUsedForFilter == true) inputTensor[6][cy][cx] = OccFrame[cy*2+1][cx*2 ] (if IsOccupancyUsedForFilter == true) inputTensor[7][cy][cx] = OccFrame[cy*2+1][cx*2+1] (if IsOccupancyUsedForFilter == true)
[0116] In the above configuration, the inputTensor that performs filtering is derived using the geoFrame and OccFrame, which has the effect of improving the accuracy of the outFrame by the filter.
[0117] The 3D-NN filter unit 611 may derive an inputTensor using a three-channel attFrame, a one-channel geometry image geoFrame, and a one-channel occupancy map OccFrame, and perform NN filtering.
[0118] inputTensor[0][y][x] = attFrame[0][y][x] <occ-inp0> inputTensor[1][y][x] = geoFrame[1][y][x] inputTensor[2][y][x] = OccFrame[y][x] (if IsOccupancyUsedForFilter == true)
[0119] The 3D-NN filter unit 611 may derive an inputTensor using a three-channel target image attFrame and a one-channel occupancy map occFrame to perform filtering.
[0120] inputTensor[0][y][x] = attFrame[0][y][x] <Formula OCC-INP1> inputTensor[1][y][x] = attFrame[1][y / SH][x / SW] inputTensor[2][y][x] = attFrame[2][y / SH][x / SW] inputTensor[3][y][x] = attFrame[2][y / SH][x / SW] inputTensor[4][y][x] = geoFrame[2][y / SH][x / SW] inputTensor[5][y][x] = OccFrame[y][x] (if IsOccupancyUsedForFilter == true)
[0121] Furthermore, the 3D-NN filter unit 611 may expand the luminance of a 4:2:0 image in the channel direction, and derive the inputTensor with the same resolution as the color difference image (channels in YYYYUV format).
[0122] inputTensor[0][cy][cx] = attFrame[0][cy*2 ][cx*2 ] <Formula OCC-INP2> inputTensor[1][cy][cx] = attFrame[0][cy*2 ][cy*2+1] inputTensor[2][cy][cx] = attFrame[0][cy*2+1][cx*2 ] inputTensor[3][cy][cx] = attFrame[0][cy*2+1][cx*2+1] inputTensor[4][cy][cx] = attFrame[1][cx][cy] inputTensor[5][cy][cx] = attFrame[2][cx][cy] inputTensor[6][cy][cx] = OccFrame[0][cx][cy] (if IsOccupancyUsedForFilter == true)
[0123] In the above configuration, since filter processing can be performed using attFrame, geoFrame, and OccFrame, the effect of further improving accuracy is achieved. When the filter output outFrame is attFrame, the accuracy of attFrame can be improved by the filter. When the filter output is geoFrame, the accuracy of geoFrame can be improved by the filter. Furthermore, if the filter output is both attFrame and geoFrame, the accuracy of both can be improved.
[0124] The 3D-NN filter unit 611 performs NN filtering processing to derive an outputTensor from an inputTensor. Filter processing indicated by PostProcessingFilter() may be performed in units of patch size (inpPatchWidth x inpPatchHeight) as shown below.
[0125] for( cTop = 0; cTop < InpPicHeightInLumaSamples; cTop += inpPatchHeight ) for( cLeft = 0; cLeft < InpPicWidthInLumaSamples; cLeft += inpPatchWidth ) { DeriveInputTensors( ) outputTensor = PostProcessingFilter( inputTensor ) StoreOutputTensors( ) } Here, DeriveInputTensors( ) indicates input data setting, and StoreOutputTensors( ) indicates output data storage. InpPicHeightInLumaSamples and InpPicWidthInLumaSamples are the height and width of the image to be filtered, and inpPatchWeight and inpPatchHeight are the width and height of the patch.
[0126] The 3D-NN filter unit 611 derives an output image outFrame from the three-dimensional array of NN output data outputTensor[][][], which is output data of the NN filter.
[0127] (Operation of the 3D-NN filter unit 611 when the occupancy map is used variable) The 3D-NN filter unit 611 may select whether to use an occupancy map for filtering the inputTensor according to the syntax value of the encoded data. For example, as shown in FIG. 9, the 3D-NN filter unit 611 decodes occupancy map usage information 3d_filter_occupancy_enabled from the encoded data. Then, when the value of 3d_filter_occupancy_enabled is true, the occupancy map may be used for filtering, and when the value is false, the occupancy map may not be used for filtering. The 3D-NN filter unit 611 sets IsOccupancyUsedForFilter = 3d_filter_occupancy_enabled, and sets the above-mentioned inputTensor.
[0128] Furthermore, the 3D-NN filter unit 611 may decode 3d_filter_occupancy_component_idc, which indicates the format of the input and output tensors, as the occupancy map usage information, and change the image used for the input tensor as follows to perform the NN filter process. If 3d_filter_occupancy_component_idc==0, set IsOccupancyUsedForFilter = false and use the formula <occ-inp1>Derive the inputTensor using If 3d_filter_occupancy_component_idc==1, set IsOccupancyUsedForFilter = true and use the formula <occ-inp1>Derive the inputTensor using If 3d_filter_occupancy_component_idc==2, set IsOccupancyUsedForFilter = false and use the formula <occ-inp2>Derive the inputTensor using If 3d_filter_occupancy_component_idc==3, set IsOccupancyUsedForFilter = true and use the formula <occ-inp2>Derive the inputTensor using
[0129] (output) After operating the NN filter, the 3D-NN filter unit 611 derives an output image outFrame from the NN output data outputTensor[][][], which is the output data, that is, a three-dimensional array.
[0130] The 3D-NN filter unit 611 may derive the output image outFrame using the following <equation OCC-OUT0>.
[0131] outFrame[0][y][x] = outputTensor[0][y][x] <formula OCC-OUT0> Here, outFrame may be either an attFrame or a geoFrame, and so on.
[0132] The 3D-NN filter unit 611 may derive the output image outFrame using the following <equation OCC-OUT1>.
[0133] outFrame[0][y][x] = outputTensor[0][y][x] <formula OCC-OUT1> outFrame[1][y / SH][x / SW] = outputTensor[1][y][x] outFrame[2][y / SH][x / SW] = outputTensor[2][y][x]
[0134] The 3D-NN filter unit 611 may derive the output image outFrame using the following <equation OCC-OUT2>.
[0135] outFrame[0][cy*2 ][cx*2 ] = outputTensor[0][cy][cx] <formula OCC-OUT2> outFrame[0][cy*2 ][cy*2+1] = outputTensor[1][cy][cx] outFrame[0][cy*2+1][cx*2 ] = outputTensor[2][cy][cx] outFrame[0][cy*2+1][cx*2+1] = outputTensor[3][cy][cx] outFrame[1][cy][cx] = outputTensor[4][cy][cx] outFrame[2][cy][cx] = outputTensor[5][cy][cx]
[0136] When outputting in 4:4:4 format, the 3D-NN filter unit 611 may derive the output image outFrame using the following <equation OCC-OUT1>.
[0137] outFrame[0][y][x] = outputTensor[0][y][x] <formula OCC-OUT1> outFrame[1][y][x] = outputTensor[1][y][x] outFrame[2][y][x] = outputTensor[2][y][x]
[0138] When outputting in 4:4:4 format, the 3D-NN filter unit 611 may derive the output image outFrame using the following <equation OCC-OUT2>.
[0139] outFrame[0][cy*2 ][cx*2 ] = outputTensor[0][cy][cx] <formula OCC-OUT2> outFrame[0][cy*2 ][cx*2+1] = outputTensor[1][cy][cx] outFrame[0][cy*2+1][cx*2 ] = outputTensor[2][cy][cx] outFrame[0][cy*2+1][cx*2+1] = outputTensor[3][cy][cx] outFrame[1][cy*2 ][cx*2 ] = outputTensor[4][cy][cx] outFrame[1][cy*2 ][cx*2+1] = outputTensor[4][cy][cx] outFrame[1][cy*2+1][cx*2 ] = outputTensor[4][cy][cx] outFrame[1][cy*2+1][cx*2+1] = outputTensor[4][cy][cx] outFrame[2][cy*2 ][cx*2 ] = outputTensor[5][cy][cx] outFrame[2][cy*2 ][cx*2+1] = outputTensor[5][cy][cx] outFrame[2][cy*2+1][cx*2 ] = outputTensor[5][cy][cx] outFrame[2][cy*2+1][cx*2+1] = outputTensor[5][cy][cx] outFrame[0], outFrame[1], and outFrame[2] represent the luma channel, chrominance (Cb) channel, and chrominance (Cr) channel of the output image, respectively.
[0140] The 3D-NN filter unit 611 may derive the output image outFrame from the outputTensor as follows, based on the value of 3d_filter_occupancy_component_idc, which indicates the format of the input and output tensors. If 3d_filter_occupancy_component_idc==0, then the formula <occ-out0>may also be used. If 3d_filter_occupancy_component_idc==1, then the formula <occ-out1>may also be used. If 3d_filter_occupancy_component_idc==2,3, then the formula <occ-out2>may also be used.
[0141] According to the above configuration, it is possible to switch whether or not to use the occupancy map in the filtering process by transmitting the occupancy map usage information 3d_filter_occupancy_component_idc. Then, the decoding device can select an optimal trade-off between image quality and complexity according to the processing capability.
[0142] <Configuration for filtering both geometry and attribute images> The 3D-CC filter unit 511 (the 3D-ALF filter unit 610 or the 3D-NN filter unit 611) may perform filtering on both the geometry image and the attribute image.
[0143] The 3D-CC filter unit 511 sets the geometry image geoFrame[y][x] decoded by the geometry image decoding unit 304 as the target image compFrame[y][x], derives inputTensor using compFrame and the occupancy image occFrame, and performs the above-mentioned filtering process.
[0144] The 3D-CC filter unit 511 sets the attribute image attFrame[y][x] decoded by the attribute image decoding unit 307 to the target image compFrame[y][x], sets the parameter attNNModel of the NN network for the attribute to the parameter of the NN model, and performs the above-mentioned filtering process.
[0145] According to the above configuration, since the filter processing using the occupancy map can be performed on both the geometry image and the attribute image, the image quality is improved. Furthermore, by using different NN model parameters for the geometry image and the attribute image, the image quality is further improved.
[0146] The 3D-NN filter unit 611 may use different NN models in the filtering process for both the geometry image and the attribute image.
[0147] The 3D-NN filter unit 611 sets the geometry image geoFrame[y][x] decoded by the geometry image decoding unit 304 as the target image compFrame[y][x], and performs the above-mentioned NN filter process using the NN model geoNNModel for the geometry image. Since the geometry image is one component, <occ-inp0>It is preferable to derive the inputTensor in the above way and use a two-channel NN model of the geometry image and the occupancy map.
[0148] The 3D-NN filter unit 611 sets the attribute image attFrame[y][x] decoded by the attribute image decoding unit 307 to the target image compFrame[y][x], and performs the above-mentioned NN filter process using the NN model attNNModel for the attribute image.
[0149] Since attribute images often have three components, <occ-inp1> 、 <occ-inp2>It is preferable to derive the inputTensor in this way and use a 4-7 channel NN model of the attribute image and occupancy map.
[0150] According to the above configuration, by using different NN models for the geometry image and the attribute image, the image quality is further improved.
[0151] <Configuration for transmitting occupancy map usage information separately for geometry image and attribute image> The 3D-CC filter unit 511 may use different occupancy map usage information for the geometry image and the attribute image. The 3D-CC filter unit 511 may decode the geometry occupancy map usage information 3d_geometry_filter_occupancy_component_idc and the attribute occupancy map usage information 3d_attribute_filter_occupancy_component_idc from the encoded data, and change the input tensor as follows to perform filtering.
[0152] Derive the inputTensor as follows, referring to the 3d_geometry_filter_occupancy_component_idc. If 3d_geometry_filter_occupancy_component_idc==0, set IsOccupancyUsedForFilter = false and use the formula <occ-inp0>Derive the inputTensor from the geoFrame using. If 3d_geometry_filter_occupancy_component_idc==1, set IsOccupancyUsedForFilter = true and use the formula <occ-inp1>Derive the inputTensor from the geoFrame and occFrame using. If 3d_geometry_filter_occupancy_component_idc==2, set IsOccupancyUsedForFilter = false and use the formula <occ-inp2>Derive the inputTensor from the geoFrame using. If 3d_geometry_filter_occupancy_component_idc==3, set IsOccupancyUsedForFilter = true and use the formula <occ-inp3>Derive the inputTensor from the geoFrame and occFrame using.
[0153] The 3D-CC filter unit 511 and the 3D-NN filter unit 611 use ALF filtering and NN filtering to perform filtering using the inputTensor as input, and output the output image outFrame as a geometry image.
[0154] The 3D-CC filter unit 511 derives the inputTensor as follows, with reference to 3d_attribute_filter_occupancy_component_idc. If 3d_attribute_filter_occupancy_component_idc==0, set IsOccupancyUsedForFilter = false and use the formula <occ-inp0>Derive inputTensor from attFrame using. If 3d_attribute_filter_occupancy_component_idc==1, set IsOccupancyUsedForFilter = true and use the formula <occ-inp1>Derive the inputTensor from attFrame and occFrame using. If 3d_attribute_filter_occupancy_component_idc==2, set IsOccupancyUsedForFilter = false and use the formula <occ-inp2>Derive inputTensor from attFrame using. If 3d_attribute_filter_occupancy_component_idc==3, set IsOccupancyUsedForFilter = true and use the formula <occ-inp3>Derive the inputTensor from attFrame and occFrame using.
[0155] The 3D-CC filter unit 511 and the 3D-NN filter unit 611 use ALF filtering and NN filtering to perform filtering using the inputTensor as input, and output the output image outFrame as an attribute image.
[0156] The NN filter processing and outFrame processing have already been explained, so the explanation will be omitted.
[0157] According to the above configuration, by switching the filters for the geometry image and the attribute image, it is possible to select a more suitable relationship between image quality and complexity.
[0158] <Syntax configuration example> Below, an example of the syntax and processing content for performing filtering separately on the geometry image and the attribute image will be described.
[0159] FIG. 11 shows an example of additional information SEI including different usage information for a geometry image and an attribute image. In the figure, assuming additional information SEI, when pfp_purpose indicating the purpose of the post filter is a predetermined value, filter processing using an occupancy map is performed on the geometry image and / or the attribute image. For example, the header decoding unit 301 decodes the filter parameters in the format of FIG. 11 from the encoded data. When pfp_purpose==3, the geometry image filter unit 305 performs filter processing on the geometry image. When pfp_purpose==4, the attribute image filter unit 308 performs filter processing on the attribute image.
[0160] <Example of configuration for transmitting filter parameters in SEI> 9 shows an example of the syntax configuration of additional information SEI for post-filtering. When using the syntax notified by post_filter_purpose(), the 3D-CC filter unit 511 applies a post-filter to the image. pfp_id: Indicates an ID for identifying a post filter. pfp_purpose: Indicates the purpose of the post filter indicated by pfp_id. 0:Improved video quality 1: Super-resolution processing 2: Chrominance format conversion 3: V3C component (geometry image and / or attribute image) quality improvement 3d_filter_component_idc: Indicates the V3C component (type of V3C image such as geometry image, attribute image, etc.) to which the post filter is applied. The image filter unit (geometry image filter unit 305, attribute image filter unit 308) performs the following operations according to the value of 3d_filter_component_idc. 0: No post-filter is applied. 1: Apply the postfilter only to the geometry image. 2: Apply post-filter to attribute images only. 3: Apply a post-filter to the geometry and attribute images. 3d_filter_occupancy_component_idc: Indicates the input data format of the following post filters to be applied to the V3C component. 0: Input data is in the form of YUV 3-channel components and does not contain an occupancy map. 1: The input data is in YUV 3-channel component format with an occupancy map as an additional channel. 2: The input data is in YYYYUV 6-channel component format without occupancy map. 3: The input data is in a format that includes an occupancy map as an additional channel in the 6-channel components of YYYYUV. 3d_filter_occupancy_scale_idc: Indicates the downsampling ratio of the occupancy map. The size ratio occFrameSizeRatio of the occupancy map with respect to the luminance image or / and the geometry image of the attribute image is represented by the following formula.
[0161] occFrameSizeRatio = 1 << 3d_filter_occupancy_scale_idc
[0162] <Another configuration example of SEI configuration> Figure 10 shows another syntax configuration example of the additional information SEI for the post-filter. pfp_id: As already described. pfp_purpose: As already described. 3d_geometry_filter_enabled_flag: A flag indicating whether to apply a predetermined post-filter to the geometry image. When this flag is equal to 1, the image filter section (geometry image filter section 305) applies the post-filter to the geometry image. When this flag is equal to 0, the post-filter is not applied to the geometry image. 3d_geometry_filter_occupancy_component_idc: Indicates the input data format of the post-filter to be applied to the geometry image. The definition of the value is the same as that of 3d_filter_occupancy_component_idc described above. 3d_geometry_filter_occupancy_scale_idc: Indicates the downsampling ratio of the occupancy map with respect to the geometry image.
[0163] occFrameSizeRatio = 1 << 3d_geometry_filter_occupancy_scale_idc 3d_attribute_filter_enabled_flag: A flag indicating whether or not to apply a specific post-filter to the attribute image. If this flag is equal to 1, the image filter unit (attribute image filter unit 308) applies a post-filter to the attribute image. If this flag is equal to 0, the post-filter is not applied to the attribute image. 3d_attribute_filter_occupancy_component_idc: Indicates the input data format of the post filter applied to the attribute image. The value definition is the same as 3d_filter_occupancy_component_idc above. 3d_attribute_filter_occupancy_scale_idc: Downsampling ratio of the occupancy map to the attribute image.
[0164] occFrameSizeRatio = 1 << 3d_attribute_filter_occupancy_scale_idc If the downsampling ratio of the geometry image and attribute image for the occupancy map is the same, 3d_filter_occupancy_scale_idc may be signaled instead of 3d_geometry_filter_occupancy_scale_idc and 3d_attribute_filter_occupancy_scale_idc.
[0165] According to the above configuration, the post filter can be turned on and off separately for the geometry image and the attribute image, which has the effect of transmitting a suitable image quality.
[0166] <Configuration for transmitting filter parameters via ASPS> Another example of the transmission format of the filter parameters will be described. When the filter parameters are notified by ASPS, the 3D-CC filter unit 511 may apply a loop filter to the image inside the geometry image decoding unit 304, and the filtered image may be referred to by the subsequent geometry image decoding unit 304 and attribute image decoding unit 307.
[0167] Figure 12 shows an example of the syntax for transmitting filter parameters using ASPS (Atlas Sequence Parameter Set) or AFPS (Atlas Frame Parameter Set). The semantics of each field are as follows: 3d_filter_enabled_flag: A flag indicating whether to apply a filter. If this flag is equal to 1, a post filter is applied based on 3d_filter_id, 3d_filter_component_idc, 3d_filter_occupancy_component_idc, and 3d_filter_occupancy_scale_idc. If this flag is equal to 0, the filter is not applied. 3d_filter_id: Indicates the ID of the NN filter to be used. The meaning of the other syntax has already been explained.
[0168] Furthermore, according to the configuration of transmitting the data in the V3C data structure, in addition to the effect of transmitting the preferred image quality described above, it becomes easy to transmit the atlas image, geometry image, attribute image and filter parameters as a set, making it easy to use in applications.
[0169] <Separate layer selection for geometry images and attribute images> Another syntax example of the transmission format of the filter parameters will be described.
[0170] FIG. 13 shows another example syntax configuration of a parameter set for three-dimensional volumetric information. 3d_filter_enabled_flag: A flag indicating whether or not to apply a post filter. If this flag is equal to 1, a post filter is applied based on 3d_geometry_filter_enabled_flag and 3d_attribute_filter_enabled_flag. If this flag is equal to 0, a post filter is not applied.
[0171] 3d_geometry_filter_enabled_flag: A flag indicating whether to apply a post filter to the geometry image. If this flag is equal to 1, a post filter is applied based on the values of 3d_geometry_filter_id, 3d_geometry_filter_occupancy_scale_idc, and 3d_geometry_filter_occupancy_component_idc. If this flag is equal to 0, a post filter is not applied. 3d_geometry_filter_id: Indicates the ID for identifying the post filter for geometry images. 3d_geometry_filter_occupancy_component_idc: Indicates the input data format of the post filter to be applied to the geometry image. 3d_geometry_filter_occupancy_scale_idc: Specifies the downsampling ratio of the occupancy map for geometry images. 3d_attribute_filter_enabled_flag: A flag indicating whether to apply a post filter to the attribute image. If this flag is equal to 1, a post filter is applied based on the values of 3d_attribute_filter_id, 3d_attribute_filter_occupancy_scale_idc, and 3d_geometry_filter_occupancy_component_idc. If this flag is equal to 0, a post filter is not applied. 3d_attribute_filter_id: Indicates the ID for identifying the post-filter for the attribute image. 3d_attribute_filter_occupancy_component_idc: Indicates the input data format of the post filter to be applied to the attribute image. 3d_attribute_filter_occupancy_scale_idc: Specifies the downsampling ratio of the occupancy map for the attribute image.
[0172] According to the above configuration, since it is possible to select whether or not to additionally input an occupancy map for the geometry image and the attribute image separately, an effect of transmitting suitable image quality is achieved.
[0173] According to the above-mentioned configuration of transmitting the information as additional information to the video stream, in addition to the effect of transmitting the preferred image quality described above, by switching filters separately for the geometry image and the attribute image, it is possible to select an even more preferred relationship between image quality and complexity.
[0174] (Configuration of 3D data encoding device according to the first embodiment) FIG. 7 is a functional block diagram showing a schematic configuration of a 3D data encoding device 11 according to the first embodiment.
[0175] The 3D data encoding device 11 includes a patch generation unit 101, an atlas information encoding unit 102, an occupancy map generation unit 103, an occupancy map encoding unit 104, a geometry image generation unit 105, a geometry image encoding unit 106, a geometry image filter parameter derivation unit 107, an attribute image generation unit 108, an attribute image encoding unit 109, an attribute image filter parameter derivation unit 110, and a multiplexing unit 111. The 3D data encoding device 11 inputs a point cloud or a mesh as 3D data, and outputs encoded data.
[0176] The patch generation unit 101 inputs 3D data (point cloud, mesh), generates a set of patches (rectangular images in this case), and outputs it. Specifically, the 3D data is divided into multiple regions, and each region is projected onto one of the planes of a 3D bounding box (FIG. 3(a)) set in the 3D space to generate multiple patches. The patch generation unit 101 outputs information about the 3D bounding box (coordinates, size, etc.) and information about mapping onto the projection surface (projection surface, coordinates, size, rotation presence / absence, etc. of each patch) as atlas information. The atlas information encoding unit 102 encodes the atlas information output from the patch generating unit 101 and outputs an atlas information encoded stream.
[0177] The occupancy map generator 103 inputs a set of patches output from the patch generator 101, and generates an occupancy map that shows the valid area of each patch (area where a point cloud or mesh exists) as a 2D binary image (e.g., valid area is 1, invalid area is 0) (FIG. 3(b)). Note that the values of the valid area and invalid area may be other values such as 255 and 0.
[0178] The occupancy map encoding unit 104 inputs the occupancy map output from the occupancy map generation unit 103, and outputs an occupancy map encoded stream and an encoded occupancy map. As an encoding method, VVC, HEVC, or the like is used.
[0179] The geometry image generating unit 105 generates a geometry image storing a depth value for the projection surface of each patch based on the 3D data (point cloud, mesh), the occupancy map, the coded occupancy map, and the atlas information (FIG. 3(c)). The geometry image generating unit 105 derives the point with the minimum depth for the projection surface among the points projected onto the pixel g(x,y) as p_min(x,y,z). In addition, the geometry image generating unit 105 derives the point with the maximum depth among the points projected onto the pixel g(x,y) and at a predetermined distance d from p_min(x,y,z) as p_max(x,y,z). The geometry image with p_min(x,y,z) projected onto all pixels of the projection surface is set as the geometry image of the Near layer (images 0,2,4,...,2N in FIG. 4). The geometry image obtained by projecting p_max(x,y,z) for all pixels on the projection plane is set as the geometry image of the Far layer (images 1, 3, 5, ..., 2N+1 in Figure 4).
[0180] The geometry image encoding unit 106 inputs a geometry image and outputs a geometry image encoding stream and an encoded geometry image. An encoding method such as VVC or HEVC is used. The geometry image encoding unit 106 may encode the geometry image of the near layer as an intra-screen picture (I-picture) and the geometry image of the far layer as an inter-screen picture (P-picture or B-picture).
[0181] The geometry image filter parameter derivation unit 107 inputs the encoded geometry image and the original geometry image, selects or derives optimal filter parameters for processing such as a linear filter or a neural network-based filter, and outputs the filter parameters.
[0182] The attribute image generating unit 108 generates an attribute image storing color information (e.g., YUV value, RGB value, etc.) for the projection surface of each patch based on the 3D data (point cloud, mesh), the coded occupancy map, the filtered image of the coded geometry image, and the atlas information (FIG. 3(d)). The attribute image generating unit 108 obtains the value of the attribute corresponding to the point p_min(x,y,z) with the minimum depth calculated by the geometry image generating unit 106, and sets the attribute image projected with the value as the attribute image of the Near layer (images 0,2,4,...,2N in FIG. 4). The attribute image obtained in the same manner for p_max(x,y,z) is set as the attribute image of the Far layer (images 1,3,5,...,2N+1 in FIG. 4).
[0183] The attribute image encoding unit 109 inputs an attribute image, and outputs an attribute image encoding stream and an encoded attribute image. An encoding method such as VVC or HEVC is used. The attribute image encoding unit 109 may encode the attribute image of the near layer as an I picture, and the attribute image of the far layer as a P picture or a B picture.
[0184] The attribute image filter parameter derivation unit 110 inputs the encoded attribute image and the original attribute image, selects or derives optimal filter parameters for filtering, such as a linear filter or a neural network-based filter, and outputs the filter parameters.
[0185] The geometry image filter parameter derivation unit 107 and / or the attribute image filter parameter derivation unit 110 set the values of pfp_id and pfp_purpose. For example, when pfp_purpose=3 is set, the geometry image filter parameter derivation unit 107 sets the values of 3d_geometry_filter_enabled_flag, 3d_geometry_filter_occupancy_component_idc, and 3d_geometry_filter_occupancy_scale_idc according to the semantics of the above SEI. The attribute image filter parameter derivation unit 110 sets the values of 3d_attribute_filter_enabled_flag, 3d_attribute_filter_occupancy_component_idc, and 3d_attribute_filter_occupancy_scale_idc according to the above SEI.
[0186] The multiplexing unit 111 inputs the filter parameters output from the geometry image filter parameter derivation unit 107 and the filter parameters output from the attribute image filter parameter derivation unit 110, and outputs them in a predetermined format. The predetermined format is, for example, SEI, which is additional information of video data, ASPS and AFPS, which are information specifying a data structure in the V3C standard, and ISOBMFF, which is a media file format standard. The multiplexing unit 111 also multiplexes the atlas information coded stream, the occupancy map coded stream, the geometry image coded stream, the attribute image coded stream, and the above filter parameters, and outputs them as coded data. As a multiplexing method, a byte stream format, ISOBMFF, etc. are used.
[0187] Even if the occupancy map is not included in the input of the NN filter, the geometry image filter parameter derivation unit 107 and / or the attribute image filter parameter derivation unit 110 may use the occupancy map for error calculation during training of the NN filter. FIG. 14 is a diagram showing a process of calculating an error used for training of the NN filter using an output image from the 3D-NN filter unit 611 and an occupancy map. The error calculation unit 612 uses outputTensor[][][] and occFrame[][] to calculate an error lossTensor[][][] that back-propagates the error at each of c=0..C-1, y=0..H-1, x=0..W-1 as follows. Here, C is the number of channels, H is the height of the input and output images, and W is the width of the input and output images. outputTensor[][][] is a three-dimensional array output image of size C*H*W output by the 3D-NN filter unit 611. occFrame[][] is an occupancy map of H*W, a two-dimensional array.
[0188] lossTensor[c][y][x] = lossFunc(outputTensor)[c][y][x] * occFrame[y][x] Here, lossFunc is a loss function. For example, a function such as Mean Squared Error (MSE) or L1 loss may be used as the loss function. In other words, if the pixel value of the occupancy map is 1 at the position (x, y) of the output pixel of each c=0..C-1, the error at that position is backpropagated. If the pixel value of the occupancy map is 0, the error at that position is not backpropagated (i.e., the backpropagated value is 0).
[0189] The above process has the effect of preferably learning only the features of the area of the attribute image and / or geometry image where the corresponding 3D data point cloud and mesh exist.
[0190] It should be noted that some of the 3D data encoding device 11 and the 3D data decoding device 31 in the above-mentioned embodiments, such as the 3D patch generation unit 101, the atlas information encoding unit 102, the patch packing unit 103, the occupancy map encoding unit 104, the geometry image generation unit 105, the geometry image encoding unit 106, the geometry image filter parameter derivation unit 107, the attribute image generation unit 108, the attribute image encoding unit 109, the attribute image filter parameter derivation unit 110, the multiplexing unit 111, the demultiplexing / decoding unit 301, the atlas information decoding unit 302, the occupancy map decoding unit 303, the geometry image decoding unit 304, the geometry image filter unit 305, the geometry reconstruction unit 306, the attribute image decoding unit 307, the attribute image filter unit 308, the attribute reconstruction unit 309 and the 3D data reconstruction unit 310, may be realized by a computer. In this case, the control function may be realized by recording a program for realizing the control function in a computer-readable recording medium, reading the program recorded in the recording medium into a computer system, and executing the program. The term "computer system" as used herein refers to a computer system built into either the 3D data encoding device 11 or the 3D data decoding device 31, and includes hardware such as an OS and peripheral devices. The term "computer-readable recording medium" refers to portable media such as flexible disks, optical magnetic disks, ROMs, and CD-ROMs, and storage devices such as hard disks built into a computer system. The term "computer-readable recording medium" may also include a medium that dynamically holds a program for a short period of time, such as a communication line when a program is transmitted via a network such as the Internet or a communication line such as a telephone line, and a medium that holds a program for a certain period of time, such as a volatile memory inside a computer system that is a server or client in that case. The above program may be a program for realizing part of the above-mentioned functions, or may be a program that can realize the above-mentioned functions in combination with a program already recorded in the computer system.
[0191] In addition, a part or all of the 3D data encoding device 11 and the 3D data decoding device 31 in the above-mentioned embodiment may be realized as an integrated circuit such as an LSI (Large Scale Integration). Each functional block of the 3D data encoding device 11 and the 3D data decoding device 31 may be individually processed, or a part or all of them may be integrated and processed. The integrated circuit method is not limited to LSI, and may be realized by a dedicated circuit or a general-purpose processor. In addition, when an integrated circuit technology that replaces LSI appears due to the progress of semiconductor technology, an integrated circuit based on that technology may be used.
[0192] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.
[0193] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims. [Industrial Applicability]
[0194] The embodiments of the present invention can be suitably applied to a 3D data decoding device that decodes encoded data in which 3D data is encoded, and a 3D data encoding device that generates encoded data in which 3D data is encoded, and can also be suitably applied to the data structure of encoded data generated by the 3D data encoding device and referenced by the 3D data decoding device. [Explanation of symbols]
[0195] 11 3D data encoding device 101 Patch Generation Unit 102 Atlas Information Encoding Unit 103 Occupancy map generator 104 Occupancy map coding unit 105 Geometry Image Generation Unit 106 Geometry Image Encoding Unit 107 Geometry Image Filter Parameter Derivation Unit 108 Attribute Image Generator 109 Attribute Image Encoding Unit 110 Attribute image filter parameter derivation unit 111 Multiplexer 21 Network 31 3D data decoding device 301 Demultiplexing and Decoding Unit 302 Atlas Information Decoding Unit 303 Occupancy Map Decoding Unit 304 Geometry Image Decoding Unit 305 Geometry Image Filter Unit 306 Geometry Reconstruction Unit 307 Attribute Image Decoding Unit 308 Attribute Image Filter Section 309 Attribute Reconstruction Unit 310 3D Data Reconstruction Department 41 3D data display device < / occ-inp1>
Claims
1. In a 3D data decoding apparatus for decoding 3D encoded data, a geometry image decoding unit that decodes a geometry image from the 3D encoded data, an attribute image decoding unit that decodes an attribute image from the 3D encoded data, an occupancy map decoding unit that decodes an occupancy map from the 3D encoded data, and a geometry image filter unit that performs filtering processing on the geometry image, an attribute image filter unit that performs filtering processing on the attribute image, wherein the geometry image filter unit and / or the attribute image filter unit decodes filter parameters including information indicating that the filter target is a 3D data (V3C) component, and performs filtering processing on the geometry image and / or the attribute image using the filter parameters and the occupancy map. 3D data decoding apparatus.
2. The geometry image filter unit and / or the attribute image filter unit further includes information for identifying the type of 3D data (V3C) component to be filtered. The 3D data decoding apparatus according to Claim 1.
3. The geometry image filter unit and / or the attribute image filter unit further includes information for identifying the format of the 3D data (V3C) component to be filtered and the presence or absence of additional input of the occupancy map to the filter. The 3D data decoding apparatus according to Claim 1.
4. The geometry image filter unit and / or the attribute image filter unit derives the filter parameters from at least one of SEI, ASPS, and AFPS. The 3D data decoding apparatus according to Claim 1.
5. In a 3D data encoding apparatus for encoding 3D data, a geometry image filter parameter derivation unit that derives filter parameters for a geometry image of the 3D data, an attribute image filter parameter derivation unit that derives filter parameters for an attribute image of the 3D data, a geometry image encoding unit that encodes the geometry image, an attribute image encoding unit that encodes the attribute image, and an occupancy map encoding unit that encodes an occupancy map. The geometry image filter parameter derivation unit and / or the attribute image filter parameter derivation unit encodes filter parameters including information indicating that the filtering target is a 3D data (V3C) component using the occupancy map. 3D data encoding device. **Claim 6** The geometry image filter parameter derivation unit and / or the attribute image filter parameter derivation unit further includes information for identifying the type of the 3D data (V3C) component to be filtered. The 3D data encoding device according to claim 5. **Claim 7** The geometry image filter parameter derivation unit and / or the attribute image filter parameter derivation unit further includes information for identifying the format of the 3D data (V3C) component to be filtered and the presence or absence of additional input of the occupancy map to the filter. The 3D data encoding device according to claim 5. **Claim 8** The geometry image filter parameter derivation unit and / or the attribute image filter parameter derivation unit encodes the filter parameters using at least one of SEI, ASPS, and AFPS. The 3D data encoding device according to claim 5.