3D data encoding device and 3D data decoding device

JP2024151760A5Pending Publication Date: 2026-04-07SHARP KK
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing 3D data encoding and decoding technologies face challenges in applying different filters to different maps (near and far) and geometry and attributes, leading to inappropriate filter processing and encoding distortion.

Method used

A 3D data decoding device that decodes encoded data using a header decoding unit to extract post-filter characteristic information and activation information, applying linear or neural network filters to geometry and attribute images based on atlas identifiers, map numbers, and persistence information.

Benefits of technology

Reduces encoding distortion and enables high-quality encoding and decoding of 3D data by applying filters appropriately to different maps and attributes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To solve a problem in which different filters cannot be applied to 3D data with different characteristics.SOLUTION: A 3D data decoding device that decodes encoded data and decodes 3D data including attribute information includes a header decoding unit that decodes post-filter characteristic information and post-filter activation information, and a frame image filter unit that performs post-filter processing of an attribute image or a geometry image, and the header decoding unit decodes characteristic information and activation information of either a linear filter or a neural network filter as the post-filter characteristic information and activation information.SELECTED DRAWING: Figure 18
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] An embodiment of the present invention relates to a 3D data encoding device and a 3D data decoding device. [Background technology]

[0002] In order to efficiently transmit or record 3D data, there are 3D data encoding devices that convert 3D data into 2D images and encode them using a video encoding method to generate encoded data, and 3D data decoding devices that decode and reconstruct 2D images from the encoded data to generate 3D data. There is also a technology that performs filtering on 2D images using auxiliary extension information of a deep learning postfilter.

[0003] Specific examples of 3D data encoding methods include MPEG-I's ISO / IEC 23090-5 V3C (Volumetric Video-based Coding) and V-PCC (Video-based Point Cloud Compression). V3C is used for encoding and decoding point clouds consisting of point positions and attribute information. Furthermore, ISO / IEC 23090-12 (MPEG Immersive Video, MIV) and ISO / IEC 23090-29 (Video-based Dynamic Mesh Coding, V-DMC), which is currently being standardized, are also used for encoding and decoding multi-viewpoint images and mesh images. In addition, in the auxiliary extension information of the neural network postfilter in Non-Patent Document 1, model information of the neural network is transmitted in the characteristic SEI, and is called for a frame specified in the application SEI, thereby enabling adaptive filtering of point cloud data. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] K.Takada, Y.Tokumo, T.Chujoh. T.Ikai,“[V-PCC][EE2.8] V3C neural-network post-filter SEI messages,” ISO / IEC JTC 1 / SC 29 / WG7, m61805, Jan. 2023 Summary of the Invention [Problem to be solved by the invention]

[0005] In Non-Patent Document 1, there are cases where filtering is inappropriate for 3D data having different characteristics. Specifically, there is a problem that different filters cannot be applied to different maps, Near and Far. There is also a problem that different filters cannot be applied to geometry and attributes.

[0006] The present invention aims to solve the above-mentioned problems in encoding and decoding 3D data using a video encoding method, and to further reduce encoding distortion by using auxiliary information from a post filter, thereby encoding and decoding 3D data with high quality. [Means for solving the problem]

[0007] In order to solve the above problems, a 3D data decoding device according to one embodiment of the present invention is a 3D data decoding device that decodes encoded data and decodes 3D data including attribute information, and is equipped with a header decoding unit that decodes post-filter characteristic information and post-filter activation information from the encoded data, a geometry image decoding unit that decodes a geometry image from the encoded data, an attribute image decoding unit that decodes an attribute image from the encoded data, and a frame image filter unit that performs post-filter processing of the attribute image or the geometry image, and is characterized in that the header decoding unit decodes an atlas identifier (atlasID) from a V3C unit header, and the number of maps, an identifier of the post-filter characteristic information, a cancellation flag, and persistence information from the activation information.

[0008] A header decoding unit of a 3D data decoding device according to one aspect of the present invention is characterized in that it decodes linear filter characteristic information and activation information as the post filter characteristic information and activation information. Effect of the Invention

[0009] According to one aspect of the present invention, distortion caused by encoding of color images can be reduced, and 3D data can be encoded and decoded with high quality. [Brief description of the drawings]

[0010] [Figure 1] 1 is a schematic diagram showing the configuration of a 3D data transmission system according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a diagram showing a hierarchical structure of data in an encoded stream. [Diagram 3] 1A to 1C are diagrams for explaining 3D data, an occupancy map, a geometry image, and an attribute image. [Figure 4] FIG. 13 is a diagram for explaining the layer structure of a geometry image and an attribute image. [Diagram 5] FIG. 11 is a diagram showing the relationship between characteristic SEI and applied SEI according to an embodiment of the present invention. [Figure 6] 1 is a functional block diagram showing a schematic configuration of a 3D data decoding device 31 according to an embodiment of the present invention. [Figure 7] FIG. 2 is a functional block diagram showing the configuration of a frame image filter unit 308. [Figure 8] FIG. 2 is a functional block diagram showing the configuration of a frame image filter unit 308. [Figure 9] 1 is a functional block diagram showing a schematic configuration of a 3D data encoding device 11 according to an embodiment of the present invention. [Figure 10] 13 is a flowchart showing a decoding process of SEI. [Figure 11] 13 is an example of the syntax of a neural network postfilter characteristics SEI message. [Figure 12]13 is an example of the syntax of a post-filter apply SEI message with map unit configuration. [Figure 13] 3A to 3D data encoding device, a 3D data decoding device, and a relationship with an SEI message according to an embodiment of the present invention. [Figure 14] 13 is an example of the syntax of a post-filter apply SEI message with map unit configuration. [Figure 15] 13 is an example of the syntax of an apply post-filter SEI message for attribute and map unit configuration. [Figure 16] 13 is an example of syntax for a geometry post-filter apply SEI message. [Figure 17] 13 is an example of the syntax of a linear filter characteristics SEI message. [Figure 18] 13 is an example of the syntax of a characteristic SEI message common to NN filters and linear filters. [Figure 19] 13 is an example of the syntax of a linear properties SEI message and an application SEI message. [Figure 20] 13 is an example of the syntax of a linear apply SEI message. [Figure 21] 1 is an example of filter parameter syntax. [Figure 22] 13 is an example of an efficient post-filter apply SEI message syntax. [Diagram 23] 13 is an example of an efficient post-filter apply SEI message syntax. [Figure 24] 13 is an example of the syntax of an apply SEI message for switching between a NN filter and a linear filter. [Diagram 25] 13 is an example of the syntax of a neural network post-filter apply SEI message. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0012] FIG. 1 is a schematic diagram showing the configuration of a 3D data transmission system 1 according to this embodiment.

[0013] The 3D data transmission system 1 is a system that transmits an encoded stream obtained by encoding 3D data to be encoded, decodes the transmitted encoded stream, and displays the 3D data. The 3D data transmission system 1 includes a 3D data encoding device 11, a network 21, a 3D data decoding device 31, and a 3D data display device 41.

[0014] 3D data T is input to the 3D data encoding device 11.

[0015] The network 21 transmits the encoded stream Te generated by the 3D data encoding device 11 to the 3D data decoding device 31. The network 21 is the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network 21 is not necessarily limited to a bidirectional communication network, and may be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting and satellite broadcasting. Furthermore, the network 21 may be replaced by a storage medium on which the encoded stream Te is recorded, such as a DVD (Digital Versatile Disc: registered trademark) or a BD (Blu-ray Disc: registered trademark).

[0016] The 3D data decoding device 31 decodes each of the encoded streams Te transmitted by the network 21, and generates one or more decoded 3D data Td.

[0017] The 3D data display device 41 displays all or part of one or more pieces of decoded 3D data Td generated by the 3D data decoding device 31. The 3D data display device 41 includes a display device such as a liquid crystal display or an organic EL (Electro-luminescence) display. Display forms include stationary, mobile, HMD, and the like. Furthermore, when the 3D data decoding device 31 has high processing power, it displays high quality images, and when it has only lower processing power, it displays images that do not require high processing power or display power.

[0018] <operator> The operators used in this specification are listed below.

[0019] >> is a right bit shift, << is a left bit shift, & is a bitwise AND, | is a bitwise OR, |= is the OR assignment operator, and || indicates logical OR.

[0020] x?y:z is a ternary operator that takes y if x is true (non-zero) and z if x is false (zero). Let y..z denote the set of integers from y to z.

[0021] Floor(a) is a function that returns the largest integer less than or equal to a.

[0022] <Structure of the coding stream Te> The data structure of the coded stream Te generated by the 3D data coding device 11 and decoded by the 3D data decoding device 31 will be described.

[0023] 2 is a diagram showing a hierarchical structure of data in the coded stream Te. The coded stream Te has a data structure of either a V3C sample stream or a V3C unit stream. A V3C sample stream includes a sample stream header and a V3C unit, while a V3C unit stream includes a V3C unit.

[0024] A V3C unit includes a V3C unit header and a V3C unit payload. The beginning of a V3C unit (=V3C unit header) is a Unit Type, which is an ID that indicates the type of V3C unit, and takes a value indicated by a label such as V3C_VPS, V3C_AD, V3C_AVD, V3C_GVD, or V3C_OVD.

[0025] If the Unit Type is V3C_VPS (Video Parameter Set), the V3C unit payload includes a V3C parameter set.

[0026] When the Unit Type is V3C_AD (Atlas Data), the V3C unit payload includes a VPS ID, atlasID, sample stream nal header, and multiple NAL units. ID stands for Idenfication and takes an integer value of 0 or greater. This atlasID may be used as an element of the applicable SEI.

[0027] A NAL unit contains a NALUnitType, a layerID, a TemporalID, and a RBSP (Raw byte sequence payload).

[0028] NAL units are identified by NALUnitType, which includes ASPS (Atlas Sequence Parameter Set), AAPS (Atlas Adaptation Parameter Set), ATL (Atlas Tile layer), SEI (Supplemental Enhancement Information), etc.

[0029] The ATL includes an ATL header and an ATL data unit, and the ATL data unit includes information such as the position and size of a patch, such as patch information data.

[0030] The SEI includes payloadType, which indicates the type of SEI, payloadSize, which indicates the size (number of bytes) of the SEI, and sei_payload, which is the data of the SEI.

[0031] If the UnitType is V3C_AVD (Attribute Video Data), it includes the VPS ID, atlasID, attrIdx of the attribute image ID, partIdx of the partition ID, mapIdx of the map ID, auxFlag indicating whether the data is auxiliary data, and video stream. The video stream indicates data such as HEVC or VVC.

[0032] If UnitType is V3C_GVD (Geometry Video Data), it includes VPS ID, atlasID, mapIdx, auxFlag, and video stream.

[0033] If UnitType is V3C_OVD (Occumancy Video Data), it contains the VPS ID, atlasID, and video stream.

[0034] (Data structure of 3D information) In this specification, three-dimensional information (3D data) is a collection of position information (x, y, z) and attribute information in three-dimensional space. For example, 3D data is expressed in the form of a point cloud, which is a group of points representing position information and attribute information in three-dimensional space, or a mesh (or polygon) consisting of the vertices and faces of triangles.

[0035] FIG. 3 is a diagram for explaining 3D data, an occupancy map, a geometry image (position information), and an attribute image. The point cloud and mesh constituting the 3D data are divided into a plurality of parts (areas) by the 3D data encoding device 11, and the point cloud included in each part is projected onto any plane of a 3D bounding box (FIG. 3(a)) set in the 3D space. The 3D data encoding device 11 generates a plurality of patches from the projected point cloud. Information on the 3D bounding box (coordinates, size, etc.) and information on mapping onto the projection plane (projection plane, coordinates, size, rotation presence / absence, etc. of each patch) are called atlas information. The occupancy map is an image showing the valid area of ​​each patch (area where the point cloud and mesh exist) as a 2D binary image (e.g., valid area is 1, invalid area is 0) (FIG. 3(b)). Values ​​other than 0 and 1, such as 255 and 0, may be used for the values ​​of the valid area and invalid area. The geometry image is an image showing the depth value (distance) of each patch with respect to the projection plane (FIG. 3(c)). The relationship between the depth value and the pixel value may be linear, or the distance may be derived from the pixel value by a lookup table, a mathematical formula, or a relational formula based on a combination of branching by values. The attribute image is an image showing the attribute of the point (e.g., RGB color). Note that these occupancy map image, geometry image, attribute image, and atlas information may be images in which partial images (patches) from different projection planes are mapped (packed) onto a two-dimensional image. The atlas information includes information on the number of patches and the projection planes corresponding to the patches. The 3D data decoding device 12 reconstructs the coordinates and attribute information of the point cloud or mesh from the atlas information, occupancy map, geometry image, and attribute image. Here, the point refers to each point of the point cloud or the vertex of the mesh. Note that instead of the occupancy map image and the geometry image, mesh information (position information) indicating the vertices of the mesh may be encoded, decoded, and transmitted. In addition, the mesh information may be divided into a base mesh that constitutes a basic mesh, which is a subset of a mesh, and a mesh displacement that indicates the displacement from the base mesh to indicate meshes other than the base mesh, and then encoded, decoded, and transmitted.

[0036] FIG. 4 is a diagram for explaining the layer structure of a geometry image and an attribute image. The geometry image and the attribute image may each be composed of a plurality of images (layers, maps). For example, a Near layer and a Far layer may be provided. Here, the Near layer and the Far layer are images constituting geometry and attributes whose depths as viewed from a certain projection surface are different from each other. The Near layer may be a collection of points whose depth is minimum for each pixel of the projection surface. The Far layer may be a collection of points whose depth is maximum within a predetermined range (for example, a range of a distance d from the Near layer) for each pixel of the projection surface. In addition, the Near layer may be coded so that it can be identified using mapIdx, such as mapIdx=0 for the Near layer and mapIdx=1 for the Far layer. In addition, the maps are not limited to 0 and 1, and more maps may be used. As shown in the figure, the geometry image encoding unit 106 may also encode the near layer images (images 0, 2, 4, ..., 2N in Figure 4) as intra-screen pictures (I-pictures) and the far layer geometry images (images 1, 3, 5, ..., 2N+1 in Figure 4) as inter-screen pictures (P-pictures or B-pictures).

[0037] (Overview of characteristic SEI and application SEI) FIG. 5(a) is a diagram showing the relationship between a post-filter characteristic SEI (hereinafter, characteristic SEI, NNFPC_SEI, characteristic information) and a post-filter application SEI (hereinafter, application SEI, NNFPA_SEI, activation SEI, activation information) according to an embodiment of the present invention. The NNC decoding unit decodes a compressed NN (Neural Network) model transmitted in the characteristic SEI to derive an NN model. MPEG Neural Network Coding (ISO / IEC 15938-19) may be used for encoding and decoding the NN model. The NN filter unit 611 performs a filter process on the attribute image using the derived NN model. The details of the NN filter unit 611 will be described later. The characteristic SEI used for the filter process is specified by a syntax element included in the application SEI, for example, an identifier nnpfa_target_id of the characteristic information of the post-filter. The application SEI transmits persistence information using a syntax element, and specifies that the attribute image is subjected to a filter process for a duration specified by the persistence information. The NN filter unit 611 may include an NNC decoding unit. The applied SEI (activation information) is not limited to the SEI in Fig. 2, and may be data including an NAL unit or atlas data.

[0038] FIG. 5(b) is a diagram showing the relationship between a post-filter characteristic SEI (hereinafter, characteristic SEI, characteristic information) and a post-filter application SEI (hereinafter, application SEI, activation SEI, activation information) according to an embodiment of the present invention. The filter parameter decoding unit decodes the syntax of the filter parameters transmitted in the characteristic SEI to derive a linear model. The linear post-filter may be, for example, an Adaptive Loop Filter (ALF). Although described as ALF, it is not limited to a loop filter and may be a post-filter. In addition, although described as linear, it may include a nonlinear element such as a simple clip or a square component. The ALF unit 610 performs a filter process on the attribute image or the geometry image using the derived linear model. The characteristic SEI used for the filter process is specified by a syntax element included in the application SEI, for example, an identifier lmpfa_target_id of the characteristic information of the post-filter. The application SEI transmits persistence information using a syntax element, and specifies that the attribute image or the geometry image is subjected to a filter process for a duration specified by the persistence information.

[0039] In addition, the characteristic SEI and application SEI may be used when the filter is a neural network or linear, and in the following description, the initial "nn" in the syntax element name and the NN in the SEI name are not limited to a neural network. However, characteristic SEIs and application SEIs dedicated to linear models, or characteristic SEIs and application SEIs dedicated to neural networks, may be used. "lm" indicates that the filter is dedicated to a linear model.

[0040] (Configuration of 3D data decoding device) FIG. 6 is a functional block diagram showing a schematic configuration of a 3D data decoding device 31 according to an embodiment of the present invention.

[0041] The 3D data decoding device 31 is made up of a header decoding unit 301, an atlas information decoding unit 302, an occupancy map decoding unit 303, a geometry image decoding unit 304, a geometry reconstruction unit 306, an attribute image decoding unit 307, a frame image filter unit 308, an attribute reconstruction unit 309, and a 3D data reconstruction unit 310. The processing of the frame image filter unit 308 may be a post-filter.

[0042] The header decoding unit 301 inputs encoded data (bit stream) multiplexed in a byte stream format, ISOBMFF (ISO Base Media File Format), etc., demultiplexes it, and outputs an atlas information encoded stream, an occupancy map encoded stream (V3C_OVD video stream), a geometry image encoded stream (V3C_GVD video stream), an attribute image encoded stream (V3C_AVD video stream), and filter parameters.

[0043] The header decoding unit 301 decodes a characteristic SEI indicating the characteristic of the post-filter processing from the encoded data. Furthermore, the header decoding unit 301 decodes an application SEI from the encoded data. For example, the header decoding unit 301 decodes nnpfc_id and nnpfc_purpose from the characteristic SEI, and decodes nnpfa_atlas_id, nnpfa_attribute_count, and nnpfa_cancel_flag from the application SEI. The nnpfc_id is an identifier used to identify the post-filter. In addition, the header decoding unit 301 may decode nnpfa_target_id, nnpfa_weight_block_size_idx, nnpfa_weight_map_width_minus1, nnpfa_weight_map_height_minus1, nnpfa_weight_map, nnpfa_strength_block_size_idx, nnpfa_strength_map_width_minus1, nnpfa_strength_map_height_minus1, and nnpfa_strength_map on an attribute image basis or a map basis, or on an attribute image and map basis.

[0044] 25 shows an example of the syntax configuration of nn_post_filter_activation of the application SEI. i is a value for identifying an attribute image included in a target atlas indicated by an atlasID (not shown), i=attrIdx, and j is a value for identifying a map, mapIdx=j.

[0045] When nnpfa_enabled_flag[i][j] is 0, post-filter processing is not performed on the image with attrIdx=i and mapIdx=j. When nnpfa_enabled_flag[i][j] is 1, nnpfa_target_id[i][j] and the value indicating the post-filter weight, nnpfa_weight_map[i][j][yIdx][xIdx], are decoded for the image with attrIdx=i and mapIdx=j.

[0046] nnpfa_weight_block_size_idx[i][j] indicates the size of the two-dimensional map showing the filter weight coefficients of the image with attrIdx = i and mapIdx = j. The size is equal to ((1 << weightSizeShift) x (1 << weightSizeShift)). When weightSizeShift == 0, it indicates that the image is not divided within the screen and the entire screen is processed as one unit.

[0047] weightSizeShift[i] == nnpfa_weight_block_size_idx == 0? 0 : nnpfa_weight_block_size_idx[i] + 5 nnpfa_weight_map_width_minus1[i][j] indicates the width - 1 of the two-dimensional map showing the filter weight coefficients. If it does not appear in the syntax of the encoded data, it is assumed that nnpfa_weight_map_width_minus1 = 0.

[0048] nnpfa_weight_map_height_minus1[i][j] indicates the height - 1 of the two-dimensional map showing the filter weight coefficients. If it does not appear in the syntax of the encoded data, it is assumed that nnpfa_weight_map_height_minus1 = 0.

[0049] nnpfa_weight_map[i][j][yIdx][xIdx] indicates the value of the two-dimensional weight coefficient map showing the filter processing weight coefficients in the (xIdx, yIdx) block of the image with attrIdx = i and mapIdx = j. nnFilterMap[i][partIdx][j][frameIdx][yIdx][xIdx] = nnpfa_weight_map[i][j][yIdx][xIdx] == 0? 0 : nnpfa_weight_map[i][j][yIdx][xIdx] + 1 where yIdx = 0..nnpfa_weight_map_height_minus1[i][j], xIdx = nnpfa_weight_map_width_minus1[i][j]-1.

[0050] nnpfa_strength_present_flag[i][j] indicates whether strength coefficients are present for image with attrIdx=i, mapIdx=j.

[0051] nnpfa_strength[i][j] exists when nnpfa_strength_present_flag[i][j] is true, and is the strength coefficient used for the input tensor of the filtering process of the image with attrIdx=i and mapIdx=j. If it does not appear in the encoded data, it is estimated as nnpfa_strength[i][j]=0. nnStrengthVal[i][partIdx][j][frameIdx] = nnpfa_strength[i][j] In the above, we have explained an example of applying the attribute image of attrIdx=i and mapIdx=j. However, when applying it to a geometry image, etc., the loop variable of the syntax can be changed from i, j to j, and the encoded data with the syntax element [i][j] to [j] can be decoded and applied to the image of mapIdx=j. nnFilterMap[j][frameIdx][yIdx][xIdx] = nnpfa_weight_map[j][yIdx][xIdx]==0 ? 0 : nnpfa_weight_map[j][yIdx][xIdx]+1 nnStrengthVal[j][frameIdx] = nnpfa_strength[j] The atlas information decoding unit 302 receives the atlas information encoded stream and decodes the atlas information.

[0052] The occupancy map decoding unit 303 decodes an occupancy map coded stream coded using VVC, HEVC, or the like, and outputs an occupancy map.

[0053] The geometry image decoding unit 304 decodes a geometry image coded stream coded using VVC, HEVC, or the like, and outputs a geometry image.

[0054] The geometry reconstruction unit 306 inputs the atlas information, the occupancy map, and the geometry image, and reconstructs the geometry (depth information, position information) in the 3D space.

[0055] The attribute image decoding unit 307 decodes an encoded stream encoded by VVC, HEVC, or the like, receives the attribute encoded stream, and outputs an attribute image.

[0056] The frame image filter unit 308 inputs the attribute image or geometry image and filter parameters for the specified image. The frame image filter unit 308 includes an NN filter unit 611, performs filtering based on the frame image and the filter parameters, and outputs the filtered image of the attribute image to the attribute reconstruction unit 309. Also, the filtered image of the geometry image is output to the geometry reconstruction unit 306.

[0057] The attribute reconstruction unit 309 receives the atlas information, the occupancy map, and the attribute image, and reconstructs the attributes (color information) in the 3D space.

[0058] The 3D data reconstruction unit 310 reconstructs 3D point cloud data or mesh data based on the reconstructed geometry information and attribute information.

[0059] (Frame image filter unit 308) 7 is a functional block diagram showing the configuration of the frame image filter unit 308. The header decoding unit 301 decodes an identifier nnpfa_atlas_id indicating target atlas information indicating the target of post-filter application from the coded data of the nn_post_filter_activation SEI (application SEI). The nnpfa_atlas_id is an identification number used to identify each patch of the attribute image. The nnpfa_atlas_id is set to the atlasID.

[0060] The identifier indicating the target atlas information can be the atlasID of the V3C unit in Figure 2. When the UnitType is V3C_AVD (attribute data), V3C_GVD (geometry data), or V3C_OVD (occupancy data), vuh_atlas_id is signaled. This vuh_atlas_id may be set to atlasID.

[0061] The frame image filter unit 308 uses the NN filter unit 611 (or the ALF unit 610) to improve the image quality of the image decoded by the attribute image decoding unit 307 or the geometry image decoding unit 304. When there are multiple attribute images in the atlas information specified by each atlasID, the frame image filter unit 308 performs filtering by switching characteristics SEI, etc., on a frame image basis.

[0062] More specifically, when nnpfa_enabled_flag[j] is 1, a filter process may be performed on the geometry image j identified by mapIdx=j based on the characteristic SEI indicated by the id TargetId[j] of the characteristic SEI to be applied to the image j. When nnpfa_enabled_flag[j] is 0, no process is performed.

[0063] Alternatively, when nnpfa_enabled_flag[i] is 1, filter processing based on the characteristic SEI indicated by TargetId[i] may be performed on the attribute image i identified by attrIdx=i. When nnpfa_enabled_flag[i] is 0, no processing is performed.

[0064] Alternatively, when nnpfa_enabled_flag[i][j] is 1, filter processing based on the characteristic SEI indicated by TargetId[i][j] may be performed on the attribute image identified by attrIdx=i and mapIdx=j. If nnpfa_enabled_flag[i][j] is 0, no processing is performed.

[0065] The frame image filter unit 308 derives the following variables according to DecAttrChromaFormat: DecAttrChromaFormat is the chrominance format of the attribute image. SW=SubWidthC = 1, SH=SubHeghtC = 1 (DecAttrChromaFormat == 0) SW=SubWidthC = 2, SH=SubHeghtC = 2 (DecAttrChromaFormat == 1) SW=SubWidthC = 2, SH=SubHeghtC = 1 (DecAttrChromaFormat == 2) SW=SubWidthC = 1, SH=SubHeghtC = 1 (DecAttrChromaFormat == 3) Here, SW=SubWidthC and SH=SubHeightC indicate the subsampling ratio of the color components to the luminance component.

[0066] When filtering into a geometry image, the frame image filter unit 308 inputs DecGeoFrames[mapIdx][frameIdx][cIdx][y][x] for each mapIdx and frameIdx, and outputs FilteredFrame[cIdx][y][x]. The output is stored in DecGeoFrames[mapIdx][frameIdx][cIdx][y][x]. where cIdx=0, x=0..DecGeoWidth-1, y=0..DecGeoHeight-1.

[0067] The following settings may be made to perform filtering in an NN filter unit 611 (ALF unit 610) described below. PicHeight=DecGeoHeight PicWidth=DecGeoWidth DecPic[frameIdx][0] = DecGeoFrames[mapIdx][frameIdx][0] DecPic[frameIdx][1] = DecGeoFrames[mapIdx][frameIdx][1] DecPic[frameIdx][2] = DecGeoFrames[mapIdx][frameIdx][2] BitDepthY = BitDepthC = DecGeoBitDepth ChromaFormatIdc = DecGeoChromaFormat DecGeoBitDepth is the bit depth of the geometry image. nnFilterWeight[yIdx][xIdx] = nnFilterMap[mapIdx][frameIdx][yIdx][xIdx] StrengthControlVal = nnStrengthVal[mapIdx][frameIdx] When filtering into an attribute image, DecAttrFrames[attrIdx][partIdx][mapIdx][frameIdx][cIdx][y][x] is input to the frame image filter unit 308 for each attrIdx, partIdx, mapIdx, and frameIdx, and FilteredFrame[cIdx][y][x] is output. The frame image filter unit 308 stores the output in DecAttrFrames[attrIdx][partIdx][mapIdx][frameIdx][cIdx][y][x].

[0068] The following settings may be made to perform filtering in an NN filter unit 611 (ALF unit 610) described below. PicHeight=DecAttrHeight PicWidth=DecAttrWidth DecPic[i][0] = DecAttrFrames[attrIdx][partIdx][mapIdx][frameIdx-i][0] DecPic[i][1] = DecAttrFrames[attrIdx][partIdx][mapIdx][frameIdx-i][1] DecPic[i][2] = DecAttrFrames[attrIdx][partIdx][mapIdx][frameIdx-i][2] where i=0..numInputPics-1. DecFrame=DecPic[0] / / Current frame (frameIdx) BitDepthY = BitDepthC = DecAttrBitDepth ChromaFormatIdc = DecAttrChromaFormat[attrIdx][partIdx][mapIdx][frameIdx] nnFilterWeight[yIdx][xIdx] = nnFilterMap[attrIdx][partIdx][mapIdx][frameIdx][yIdx][xIdx] StrengthControlVal = nnStrengthVal[attrIdx][partIdx][mapIdx][frameIdx] DecGeoFrames[mapIdx][frameIdx][cIdx][y][x]: A 7-dimensional array showing the decoded geometry image. Each array value shows the pixel value at component cIdx and position (x, y). Here, cIdx=0..DecAttrNumComp-1, x=0..DecGeoWidth-1, y=0..DecGeoHeight-1.

[0069] DecGeoHeight[mapIdx][frameIdx]: height of the geometry image.

[0070] DecGeoWidth[mapIdx][frameIdx]: width of the geometry image.

[0071] DecAttrFrames[attrIdx][partIdx][mapIdx][frameIdx][cIdx][y][x]: A 7-dimensional array showing the attribute image after decoding. Each array value shows the pixel value at the component cIdx and position (x, y). Here, cIdx=0..DecAttrNumComp-1, x=0..DecAttrWidth-1, y=0..DecAttrHeight-1.

[0072] DecAttrHeight[attrIdx][partIdx][mapIdx][frameIdx]: Height of the attribute image.

[0073] DecAttrWidth[attrIdx][partIdx][mapIdx][frameIdx]:Width of the attribute image.

[0074] DecAttrNumComp[attrIdx][partIdx][mapIdx][frameIdx]: The number of components in the attribute image.

[0075] DecAttrChromaFormat[attrIdx][partIdx][mapIdx][frameIdx]: Chroma format of the attribute image. Here, DecAttrChromaFormat indicates the ChromaFormatIdc of the attribute image. For example, it may be 1 (YUV4:2:0) or 3 (YUV4:4:4).

[0076] nnFilterMap[attrIdx][partIdx][mapIdx][frameIdx] is an array that stores the weights of the filtered image and the unfiltered image. As described above, the decoded value of the applied SEI may be used.

[0077] nnStrengthVal[attrIdx][partIdx][mapIdx][frameIdx] is a strength coefficient used in the NN filter. This value may be a value obtained by decoding the applied SEI as described above. It may also be, for example, a slice QP value of the video coding data.

[0078] According to the above configuration, the network model (filter characteristics) to be applied to each frame is specified by the characteristics SEI, which has the effect of improving image quality.

[0079] (NN filter section 611) A neural network model has a topology, such as the number of convolutions, the number of layers, the kernel size, and the connection relationships.

[0080] Here, the neural network model (hereinafter, NN model) means the elements and connection relationships (topology) of a neural network, and the parameters (weights, biases) of the neural network. Note that the NN filter unit 611 may fix the topology and switch only the parameters according to the image to be filtered.

[0081] The frame image filter unit 308 derives an input InputTensor to the NN filter unit 611 from the input image Frame, and the NN filter unit 611 performs filtering by using the inputTensor. The neural network model used is a model corresponding to nnpfa_target_id. The input image may be an image for each component, or may be an image having multiple components as channels.

[0082] (Derivation of scaled strength input value strengthControlScaledVal) Furthermore, the NN filter unit 611 derives strengthControlScaledVal using the integer StrengthControlVal transmitted in the encoded data of the adaptive SEI as the input tensor according to nnpfc_inp_format_idc. nnpfc_inp_format_idc is a variable indicating whether the input tensor is a floating-point number (real number) or an integer. Here, normalization is performed so that the value range is 0..1 when the value is a floating-point number (nnpfc_inp_format_idc==0). The normalized strengthControlScaledVal is input to the input tensor. if(nnpfc_inp_format_idc==0) { strengthControlScaledVal = Floor(StrengthControlVal * ((1 << inpTensorBitDepthY) - 1)) } else { strengthControlScaledVal = StrengthControlVal / ((1 << inpTensorBitDepthY) - 1) } In addition, the integer StrengthControlVal transmitted in the encoded data of the adaptive SEI may be converted to a floating-point number StrengthControlVal of 0..1, and then the converted floating-point number may be converted again to a value corresponding to nnpfc_inp_format_idc using the following formula.

[0083] StrengthControlVal = nnStrengthVal[attrIdx][partIdx][mapIdx][frameIdx] StrengthControlVal = Floor(StrengthControlVal / ((1 << inpTensorBitDepthY) - 1)) The conversion process from real numbers to nnpfc_inp_format_idc is as follows. if(nnpfc_inp_format_idc==1) { strengthControlScaledVal = Floor(StrengthControlVal * ((1 << inpTensorBitDepthY) - 1)) } else { strengthControlScaledVal = StrengthControlVal } According to the above configuration, whether the input of the neural network model to which post-filtering is performed is a real number or an integer, conversion is performed according to the input tensor format information, so that the model format can be freely selected.

[0084] Note that a real StrengthControlVal may be derived from video coding data for the coding stream, and a conversion process according to nnpfc_inp_format_idc may be performed from the real number to derive strengthControlScaledVal. In particular, when filtering is performed on an attribute image, a real StrengthControlVal may be derived from video coding data for the attribute coding stream (V3C_AVD). When filtering is performed on a geometry image, a real StrengthControlVal may be derived from video coding data for the geometry coding stream (V3C_AVD). StrengthControlVal may also be derived from a luminance quantization parameter value sliceQPY in a slice header in the VVC decoding unit.

[0085] StrengthControlVal = sliceQPY ÷ 63 In the above configuration, a strength value StrengthControlVal derived from a quantization parameter of VVC encoded data is used, and strengthControlScaledVal converted to a real number or an integer is used for the input tensor. Therefore, it is possible to perform a filter process adjusted to the image of the transmitted attribute encoded stream and geometry encoded stream encoded data.

[0086] The NN filter unit 611 may repeatedly apply the following process.

[0087] The NN filter unit 611 performs a convolution operation (conv, convolution) on the inputTensor and kernel k[m][n][yy][xx] for the number of layers, and generates an output image outputTensor by adding bias. Here, m is the number of channels in the inputTensor, n is the number of channels in the outputTensor, yy is the height of kernel k, and xx is the width of kernel k.

[0088] At each layer, an outputTensor is generated from an inputTensor.

[0089] outputTensor[nn][yy][xx]=ΣΣΣ(k[mm][nn][i][j]*inputTensor[mm][yy+j-of][xx+i-of]+bias[nn]) Here, nn=0..n-1, mm=0..m-1, yy=0..height-1, xx=0..width-1, i=0..yy-1, j=0..xx-1. width is the width of inputTensor and outputTensor, and height is the height of inputTensor and outputTensor. Σ is the sum of mm=0..m-1, i=0..yy-1, j=0..xx-1, respectively. of is the width or height of the area required around the inputTensor to generate the outputTensor.

[0090] In the case of 1x1 Conv, Σ represents the sum of mm=0..m-1, i=0, j=0. In this case, set of=0. In the case of 3x3 Conv, Σ represents the sum of mm=0..m-1, i=0..2, j=0..2. In this case, set of=1. When the value of yy+j-of is less than 0 or greater than or equal to height, or when the value of xx+i-of is less than 0 or greater than or equal to width, the value of inputTensor[mm][yy+j-of][xx+i-of] may be set to 0. Alternatively, the value of inputTensor[mm][yy+j-of][xx+i-of] may be set to inputTensor[mm][yclip][xclip], where yclip is max(0, min(yy+j-of, height-1)) and xclip is (0, min(xx+i-of, width-1)).

[0091] In the next layer, the obtained outputTensor is used as a new inputTensor, and the same process is repeated for the number of layers. An activation layer may be placed between layers. A pooling layer or skip connection may also be used. Finally, the OutAttrFrame is derived from the obtained outputTensor.

[0092] In addition, a process called Depth-wise Conv, which is shown in the following formula, may be performed using a kernel k[n][yy][xx], where nn=0..n-1, xx=0..width-1, and yy=0..height-1.

[0093] outputTensor[nn][yy][xx]=ΣΣ(k[nn][i][j]*inputTensor[nn][yy+j-of][xx+i-of]+bias[nn]) Also, a nonlinear process called Activate, for example, ReLU, may be used. ReLU(x) = x >= 0 ? x : 0 Alternatively, leakyReLU shown in the following formula may be used.

[0094] leakyReLU(x) = x >= 0 ? x : a*x Here, a is a predetermined value less than 1, for example, 0.1 or 0.125. In order to perform integer arithmetic, all the above values ​​of k, bias, and a may be set to integers, and a right shift may be performed after conv to generate the outputTensor.

[0095] With ReLU, 0 is always output for values ​​less than 0, and the input value is output as is for values ​​greater than or equal to 0. On the other hand, with leakyReLU, linear processing is performed for values ​​less than 0 with the gradient set by a. With ReLU, the gradient for values ​​less than 0 disappears, which can make it difficult for learning to progress. With leakyReLU, the gradient for values ​​less than 0 remains, making the above problem less likely to occur. Of the above leakyReLU(x), PReLU, which uses a parameterized value of a, may also be used.

[0096] (Filtering process of NN filter unit 611) The NN filter unit 611 derives input data inputTensor[][][] for the NN filter unit 611 based on the input image DecPic in accordance with nnpfc_inp_order_idc. Furthermore, the above-mentioned strengthControlScaledVal may be input. for(i = 0; i < numInputPics; i++) { if (nnpfc_inp_order_idc == 0) for(yP = -overlapSize; yP < inpPatchHeight+overlapSize; yP++) for(xP = -overlapSize; xP < inpPatchWidth+overlapSize; xP++) { xIdx = (strengthSizeShift == 0) ? 0 : xP >> strengthSizeShift yIdx = (strengthSizeShift == 0) ? 0 : yP >> strengthSizeShift inpVal = InpY(InpSampleVal(cTop+yP, cLef+xP, PicHeight, PicWidth, DecPic[i][0])) if(!nnpfc_component_last_flag) inputTensor[0][i][0][yP+overlapSize][xP+overlapSize] = inpVal else inputTensor[0][i][yP+overlapSize][xP+overlapSize][0] = inpVal if(nnpfc_auxiliary_inp_idc == 1) if(!nnpfc_component_last_flag) inputTensor[0][i][1][yP+overlapSize][xP+overlapSize] = StrengthControlVal[yIdx][xIdx] else inputTensor[0][i][yP+overlapSize][xP+overlapSize][1] = StrengthControlVal[yIdx][xIdx] } else if(nnpfc_inp_order_idc == 1) for(yP = -overlapSize; yP < inpPatchHeight+overlapSize; yP++) for(xP = -overlapSize; xP < inpPatchWidth+overlapSize; xP++) { inpCbVal = InpC(InpSampleVal(cTop+yP, cLeft+xP, PicHeight / SubHeightC, PicWidth / SubWidthC, DecPic[i][1])) inpCrVal = InpC(InpSampleVal(cTop+yP, cLeft + xP, PicHeight / SubHeightC, PicWidth / SubWidthC, DecPic[i][2])) if(!nnpfc_component_last_flag) { inputTensor[0][i][0][yP+overlapSize][xP+overlapSize] = inpCbVal inputTensor[0][i][1][yP+overlapSize][xP+overlapSize] = inpCrVal } else { inputTensor[0][i][yP+overlapSize][xP+overlapSize][0] = inpCbVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][1] = inpCrVal } if(nnpfc_auxiliary_inp_idc == 1) if(!nnpfc_component_last_flag) inputTensor[0][i][2][yP+overlapSize][xP+overlapSize] = strengthControlScaledVal else inputTensor[0][i][yP+overlapSize][xP+overlapSize][2] = strengthControlScaledVal } else if(nnpfc_inp_order_idc == 2) for(yP = -overlapSize; yP < inpPatchHeight+overlapSize; yP++) for(xP = -overlapSize; xP < inpPatchWidth+overlapSize; xP++) { yY = cTop + yP xY = cLeft + xP yC = yY / SubHeightC xC = xY / SubWidthC inpYVal = InpY(InpSampleVal(yY, xY, PicHeight, PicWidth, DecPic[i][0])) inpCbVal = InpC(InpSampleVal(yC, xC, PicHeight / SubHeightC, PicWidth / SubWidthC, DecPic[i][1])) inpCrVal = InpC(InpSampleVal(yC, xC, PicHeight / SubHeightC, PicWidth / SubWidthC, DecPic[i][2])) if(!nnpfc_component_last_flag) { inputTensor[0][i][0][yP+overlapSize][xP+overlapSize] = inpYVal inputTensor[0][i][1][yP+overlapSize][xP+overlapSize] = inpCbVal inputTensor[0][i][2][yP+overlapSize][xP+overlapSize] = inpCrVal } else { inputTensor[0][i][yP+overlapSize][xP+overlapSize][0] = inpYVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][1] = inpCbVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][2] = inpCrVal } if(nnpfc_auxiliary_inp_idc == 1) if(!nnpfc_component_last_flag) inputTensor[0][i][3][yP+overlapSize][xP+overlapSize] = StrengthControlVal[yIdx][xIdx] else inputTensor[0][i][yP+overlapSize][xP+overlapSize][3] = StrengthControlVal[yIdx][xIdx] } else if(nnpfc_inp_order_idc == 3) for(yP = -overlapSize; yP < inpPatchHeight+overlapSize; yP++) for(xP = -overlapSize; xP < inpPatchWidth+overlapSize; xP++) { yTL = cTop + yP*2 xTL = cLeft + xP*2 yBR = yTL + 1 xBR = xTL + 1 yC = cTop / 2 + yP xC = cLeft / 2 + xP inpTLVal = InpY(InpSampleVal(yTL, xTL, PicHeight, PicWidth, DecPic[i][0])) inpTRVal = InpY(InpSampleVal(yTL, xBR, PicHeight, PicWidth, DecPic[i][0])) inpBLVal = InpY(InpSampleVal(yBR, xTL, PicHeight, PicWidth, DecPic[i][0])) inpBRVal = InpY(InpSampleVal(yBR, xBR, PicHeight, PicWidth, DecPic[i][0])) inpCbVal = InpC(InpSampleVal(yC, xC, PicHeight / 2, PicWidth / 2, DecPic[i][1])) inpCrVal = InpC(InpSampleVal(yC, xC, PicHeight / 2, PicWidth / 2, DecPic[i][2])) if(!nnpfc_component_last_flag) { inputTensor[0][i][0][yP+overlapSize][xP+overlapSize] = inpTLVal inputTensor[0][i][1][yP+overlapSize][xP+overlapSize] = inpTRVal inputTensor[0][i][2][yP+overlapSize][xP+overlapSize] = inpBLVal inputTensor[0][i][3][yP+overlapSize][xP+overlapSize] = inpBRVal inputTensor[0][i][4][yP+overlapSize][xP+overlapSize] = inpCbVal inputTensor[0][i][5][yP+overlapSize][xP+overlapSize] = inpCrVal } else { inputTensor[0][i][yP+overlapSize][xP+overlapSize][0] = inpTLVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][1] = inpTRVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][2] = inpBLVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][3] = inpBRVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][4] = inpCbVal inputTensor[0][i][yP+overlapSize][xP+overlapSize][5] = inpCrVal } if(nnpfc_auxiliary_inp_idc == 1) if(!nnpfc_component_last_flag) inputTensor[0][i][6][yP+overlapSize][xP+overlapSize] = StrengthControlVal[yIdx][xIdx] else inputTensor[0][i][yP+overlapSize][xP+overlapSize][6] = StrengthControlVal[yIdx][xIdx] } } The functions Reflect(x,y) and Wrap(x,y) are defined as follows: Reflect(x, y) = Min (-y, x) (if y < 0) = Min (x-(yx), x) (if y > x) = y (otherwise) Wrap(x, y) = Max (0, x+y+1) (if y < 0) = Min (x, yx-1) (if y > x) = y (otherwise) The NN filter unit 611 performs NN filtering to derive an outputTensor from an inputTensor. Filtering indicated by PostProcessingFilter() may be performed in units of patch size (inpPatchWidth x inpPatchHeight) as shown below.

[0097] According to the value of nnpfc_out_order_idc, the decoded image is filtered by a post-filtering process PostProcessingFilter() to generate Y, Cb, and Cr pixel arrays FilteredPic[0], FilteredPic[1], and FilteredPic[2] as follows.

[0098] The post-filtering process PostProcessingFilter() of the decoded image is as follows.

[0099] if(nnpfc_inp_order_idc==0) for(cTop=0;cTop <InpPicHeightInLumaSamples;cTop+=inpPatchHeight) for(cLeft=0;cLeft <InpPicWidthInLumaSamples;cLeft+=inpPatchWidth){ DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } else if(nnpfc_inp_order_idc==1) for(cTop=0;cTop <InpPicHeightInLumaSamples / SubHeightC; cTop+=inpPatchHeight) for(cLeft=0;cLeft <InpPicWidthInLumaSamples / SubWidthC; cLeft+=inpPatchWidth) { DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } else if(nnpfc_inp_order_idc==2) for(cTop=0;cTop <InpPicHeightInLumaSamples;cTop+=inpPatchHeight) for(cLeft=0;cLeft <InpPicWidthInLumaSamples;cLeft+=inpPatchWidth){ DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } else if(nnpfc_inp_order_idc==3) for(cTop=0;cTop <InpPicHeightInLumaSamples;cTop+=inpPatchHeight*2) for(cLeft=0;cLeft <InpPicWidthInLumaSamples;cLeft+=inpPatchWidth*2){ DeriveInputTensors() outputTensor=PostProcessingFilter(inputTensor) StoreOutputTensors() } Here, DeriveInputTensors( ) is a function that sets input data, and StoreOutputTensors( ) is a function that stores output data. InpWidth and InpHeight are the size of the input data, and may be DecAttrWidth, DecAttrHeight, or DecGeoWidth, DecGeoHeight. inpPatchWeight and inpPatchHeight are the width and height of the patch. for(i=0; i <numOutputPics; i++) { if(nnpfc_out_order_idc==0) for(yP=0; yP<outPatchHeight; yP++) for(xP=0; xP<outPatchWidth; xP++) { yY = cTop * outPatchHeight / inpPatchHeight + yP xY = cLeft * outPatchWidth / inpPatchWidth + xP if (yY<nnpfc_pic_height_in_luma_samples && xY<nnpfc_pic_width_in_luma_samples) if(!nnpfc_component_last_flag) FilteredPic[i][0][yY][xY] = outputTensor[0][i][0][yP][xP] else FilteredPic[i][0][yY][xY]= outputTensor[0][i][yP][xP][0] } else if(nnpfc_out_order_idc==1) for(yP=0; yP<outPatchCHeight; yP++) for(xP=0; xP<outPatchCWidth; xP++) { xSrc = cLeft * horCScaling + xP ySrc = cTop * verCScaling + yP if (ySrc<nnpfc_pic_height_in_luma_samples / outSubHeightC && xSrc<nnpfc_pic_width_in_luma_samples / outSubWidthC) if(!nnpfc_component_last_flag) { FilteredPic[i][1][ySrc][xSrc] = outputTensor[0][i][0][yP][xP] FilteredPic[i][2][ySrc][xSrc] = outputTensor[0][i][1][yP][xP] } else { FilteredPic[i][1][ySrc][xSrc] = outputTensor[0][i][yP][xP][0] FilteredPic[i][2][ySrc][xSrc] = outputTensor[0][i][yP][xP][1] } } else if(nnpfc_out_order_idc==2) for(yP=0; yP<outPatchHeight; yP++) for(xP=0; xP<outPatchWidth; xP++) { yY = cTop*outPatchHeight / inpPatchHeight + yP xY = cLeft*outPatchWidth / inpPatchWidth + xP yC = yY / outSubHeightC xC = xY / outSubWidthC yPc = (yP / outSubHeightC)*outSubHeightC xPc = (xP / outSubWidthC)*outSubWidthC if (yY<nnpfc_pic_height_in_luma_samples && xY<nnpfc_pic_width_in_luma_samples) if(nnpfc_component_last_flag==0) { FilteredPic[i][0][yY][xY] = OutY(outputTensor[i][0][0][yP][xP]) FilteredPic[i][1][yC][xC] = OutC(outputTensor[i][0][1][yPc][xPc]) FilteredPic[i][2][yC][xC] = OutC(outputTensor[i][0][2][yPc][xPc]) } else { FilteredPic[i][0][yY][xY][yY][xY] = OutY( outputTensor[i][0][yP][xP][0]) FilteredPic[i][1][yC][xC] = OutC(outputTensor[i][0][yPc][xPc][1]) FilteredPic[i][2][yC][xC] = OutC(outputTensor[i][0][yPc][xPc][2]) } } else if(nnpfc_out_order_idc==3) for(yP=0; yP<outPatchHeight; yP++) for(xP = 0; xP < outPatchWidth; xP++) { ySrc = cTop / 2*outPatchHeight / inpPatchHeight + yP xSrc = cLeft / 2*outPatchWidth / inpPatchWidth + xP if (ySrc<nnpfc_pic_height_in_luma_samples / 2 && xSrc<nnpfc_pic_width_in_luma_samples / 2) if(nnpfc_component_last_flag==0) { FilteredPic[i][0][xSrc*2][ySrc*2] = OutY(outputTensor[i][0][0][yP][xP]) FilteredPic[i][0][xSrc*2+1][ySrc*2] = OutY(outputTensor[i][0][1][yP][xP]) FilteredPic[i][0][xSrc*2][ySrc*2+1] = OutY(outputTensor[i][0][2][yP][xP]) FilteredPic[i][0][xSrc*2+1][ySrc*2+1] = OutY(outputTensor[i][0][3][yP][xP]) FilteredPic[i][1][xSrc][ySrc] = OutC(outputTensor[i][0][4][yP][xP]) FilteredPic[i][2][xSrc][ySrc] = OutC(outputTensor[i][0][5][yP][xP]) } else { FilteredPic[i][0][xSrc**2][ySrc*2] = OutY(outputTensor[i][0][yP][xP][0]) FilteredPic[i][0][xSrc*2+1][ySrc*2] = OutY(outputTensor[i][0][yP][xP][1]) FilteredPic[i][0][xSrc*2][ySrc*2+1] = OutY(outputTensor[i][0][yP][xP][2]) FilteredPic[i][0][xSrc*2+1][ySrc*2+1] = OutY(outputTensor[i][0][yP][xP][3]) FilteredPic[i][1][xSrc][ySrc] = OutC(outputTensor[i][0][yP][xP][4]) FilteredPic[i][2][xSrc][ySrc] = OutC(outputTensor[i][0][yP][xP][5]) } } } OutFrame = FilteredPic[0

[0100] (ALF unit 610) Note that the post-filter specified by the persistence information and characteristic SEI identifier in the NNPFA SEI may be a filter using a linear model (Wiener Filter). The linear filter may be filtered by the ALF unit 610. Specifically, the filter target image DecFrame is divided into small regions of a certain size (e.g., 4x4, 1x1) (x=xSb..xSb+bSW-1, y=xSb..xSb+bSH-1, (xSb, ySb) are the top left coordinates of the small region, bSW, bSH are the width and height of the small region). Then, the filter process is performed on a small region basis. The selected filter coefficient coeff[] is used to derive from DecFrame.

[0101] outframe[cIdx][y][x] = Σ(coeff[i] * DecFrame[cIdx][y+ofy][x+ofx] + offset) >> shift Here, ofx and ofy are the offsets of the reference position determined according to the filter position i. offset=1<<(shift-1). shift corresponds to the precision of the filter coefficients and is a constant such as 6, 7, 8, etc.

[0102] Note that the image may be classified in units of small regions, and the filter coefficient coeff[classId][] may be selected according to the classId (=classId[y][x]) of the derived small region to perform the filtering process.

[0103] outframe[cIdx][y][x] = (Σ(coeff[classId][i] * DecFrame[cIdx][y+ofy][x+ofx] + offset) >> shift Note that classId may be derived using the activity or directionality of a block (pixel in the case of 1x1). For example, classId may be derived using activity Act derived from the sum of absolute difference as shown below.

[0104] classId=Act Act may be derived as follows:

[0105] Act = Σ|x-xi| Here, i=0..7 and xi indicates the adjacent pixels of the target pixel x. The directions of the adjacent pixels may be eight directions, including up, down, left, right, and four directions at 45 degrees with respect to i=0..7. Act = Σ|-xl+2*x-xr|+Σ|-ya+2*x-yb| may be used. Note that l, r, a, and b are the opposites of left, right, above, and below, respectively, and indicate the pixels to the left, right, top, and bottom of x. Act=Min(Num-1, Act>>shift) It is also possible to quantize the signal by the shift value as shown above, and then clip the resulting signal to NumA-1. Furthermore, classId may be derived using the directionality D using the following formula.

[0106] classId=Act+D*NumA Here, for example, Act=0..NumA-1, D=0..4 etc.

[0107] When the filtered image synthesis unit 612 is not applied, the OutFrame obtained by the NN filter unit 611 (ALF unit 610) is output as it is as the FilteredFrame.

[0108] (Configuration with synthesis unit) Fig. 8 is a functional block diagram showing another configuration of the frame image filter unit 308. Here, a weighted average of the filtered image OutFrame and the unfiltered image DecFrame is calculated using nnFilterWeight, enabling more adaptive filtering. Explanation of the means already explained in Fig. 7 will be omitted.

[0109] The filter image synthesis unit 612 synthesizes the DecFrame and the OutFrame based on the value of nnFilterWeight decoded from the encoded data of the applied SEI, and outputs the FilteredFrame.

[0110] Each pixel in the FilteredFrame is derived as follows.

[0111] FilteredFrame[cIdx][y][x] = (nnFilterWeight[yIdx][xIdx] * OutFrame[cIdx][y][x] + ((1<<shift) - nnFilterWeight[yIdx][xIdx])*DecFrame[cIdx][y][x]) + offset) > > shift Here, cIdx=0..DecNumComp-1, x=0..DecWidth-1, y=0..DecHeight-1. Alternatively, the following settings may be used: xIdx = (weightSizeShift==0) ? 0 : (x>>weightSizeShift) yIdx = (weightSizeShift==0) ? 0 : (y>>weightSizeShift) shift=6 offset=1<<(shift-1) The parameter weightSizeShift indicates the size of the two-dimensional map that indicates the filter weight coefficients.

[0112] Note that nnFilterWeight may take only binary values ​​such as 0 and 1. In other words, the frame image filter unit 308 may switch between applying and not applying a filter, such as not applying a filter when nnFilterWeight is 0 and applying a filter when nnFilterWeight is 1, as shown below.

[0113] FilteredFrame[cIdx][y][x] = nnFilterWeight[yIdx][xIdx] * OutFrame[cIdx][y][x] + (1 - nnFilterWeight[yIdx][xIdx]) * DecFrame[cIdx][y][x] When filtering geometry images, the frame filter unit 308 sets DecFrame = DecGeoFrames[mapIdx][frameIdx], DecNumComp = 0, DecWidth = DecGeoWidth, and DecHeight = DecGeoHeight, and when filtering attribute images, it sets DecFrame = DecAttrFrames[attrIdx][partIdx][mapIdx][frameIdx], DecNumComp = DecAttrNumComp, DecWidth = DecAttrWidth, and DecHeight = DecAttrHeight.

[0114] According to the above configuration, the image quality is improved by individually changing the adaptive SEI for the attribute image i and the degree of application of the characteristic SEI and the NN model according to the characteristics of the frame.

[0115] (Decryption of applicable SEI and application of filter) 10 is a diagram showing a flowchart of the process of the 3D data decoding device. The 3D data decoding device performs the following process including decoding the applied SEI message.

[0116] S6001: The header decoder 301 decodes the cancel_flag from the applicable SEI. The cancel_flag may be nnpfc_cancel_flag. Alternatively, it may be lmpfa_cancel_flag, which will be described later.

[0117] S6002: If cancel_flag is 1, end the processing for the frame that is the target of cancel_flag. If cancel_flag is 0, proceed to S6003.

[0118] S6003: The header decoder 301 decodes the nnpfa_persistence_flag for frame i from the applicable SEI.

[0119] S6004: The header decoder 301 decodes the nnpfa_target_id for frame i from the applicable SEI.

[0120] S6005: Identify a characteristic SEI having the same nnpfc_id as the nnpfa_target_id, and derive parameters of the NN model from the characteristic SEI.

[0121] S6006: The NN filter unit 611 executes filtering using the derived parameters of the NN model. When cancel_flag is lmpfa_cancel_flag, nnpfa_persistence_flag, nnpfa_target_id, and nnpfc_id may be lmpfa_persistence_flag, lmpfa_target_id, and lmpfc_id described below.

[0122] <Syntax configuration example

[0123] (Neural network post-filter characteristics SEI) Figure 11 shows part of the syntax of nn_post_filter_characteristics(payloadSize) (characteristic SEI). The argument payloadSize represents the number of bytes of this SEI message.

[0124] The persistence scope for applying the characteristic SEI is a Coded Atlas Sequence (CAS). That is, the SEI is applied in units of CAS. Note that CAS is an abbreviation of coded atlas sequence, which is a sequence of coded atlas access units in decoding order. More specifically, it is a sequence starting from an IRAP coded atlas access unit with NoOutputBeforeRecoveryFlag==1, followed by coded atlas access units that are not coded atlas access units with NoOutputBeforeRecoveryFlag==1. Note that the IRAP coded atlas access unit may be an instantaneous decoding refresh (IDR) coded access unit, a broken link access (BLA) coded access unit, or a clean random access (CRA) coded access unit.

[0125] In this characteristic SEI, the following syntax elements are coded, encoded, and transmitted:

[0126] The width and height of the decoded image, in units of luma pixels, are denoted here by InpPicWidthInLumaSamples and InpPicHeightInLumaSamples, respectively.

[0127] InpPicWidthInLumaSamples = pps_pic_width_in_luma_samples - SubWidthC * (pps_conf_win_left_offset + pps_conf_win_right_offset) InpPicHeightInLumaSamples = pps_pic_height_in_luma_samples - SubHeightC * (pps_conf_win_top_offset + pps_conf_win_bottom_offset) The input image of the post filter is a two-dimensional array of luminance pixels, DecFrame[0][y][x], with vertical coordinate y and horizontal coordinate x, and two-dimensional arrays of chrominance pixels, DecFrame[1][y][x] and DecFrame[2][y][x]. Here, the y coordinate of the upper left corner of the pixel array is 0, and the x coordinate is 0.

[0128] The pixel bit length of the luminance of the decoded image is BitDepthY. The pixel bit length of the chrominance of the decoded image is BitDepthC. Note that both BitDepthY and BitDepthC are set equal to BitDepth.

[0129] The variable SubWidthC is the chrominance subsampling ratio for luminance in the horizontal direction of the decoded image, and the variable SubHeightC is the chrominance subsampling ratio for luminance in the horizontal direction of the decoded image. Note that SubWidthC is set equal to the variable SubWidthC of the encoded data. SubHeightC is set equal to the variable SubHeightC of the encoded data.

[0130] The variable SliceQPY is set equal to the quantization parameter SliceQpY that is updated at the slice level of the encoded data.

[0131] Use nnpfc_out_sub_c_flag to derive the variables outSubWidthC and outSubHeightC. If nnpfc_out_sub_c_flag is 1, it indicates outSubWidthC=1 and outSubHeightC=1. If nnpfc_out_sub_c_flag is 0, it indicates outSubWidthC=2 and outSubHeightC=1. If nnpfc_out_sub_c_flag does not exist, it is assumed that outSubWidthC=SubWidthC and outSubHeightC=SubHeightC. outSubWidthC and outSubHeightC are the subsampling ratios of the chrominance components to the luma component.

[0132] if (nnpfc_out_sub_c_flag does not exist) { outSubWidthC= SubWidthC outSubHeightC = SubWidthC } else if (nnpfc_out_sub_c_flag == 1) { outSubWidthC= 1 outSubHeightC= 1 } else { outSubWidthC= 2 outSubHeightC= 1 } A value of 0 for the syntax element nnpfc_component_last_flag specifies that the second dimension of the input tensor inputTensor is used for postfiltering, and the output tensor outputTensor resulting from the postfiltering is used for the channels.

[0133] A value of 1 for nnpfc_component_last_flag specifies that the last dimension of the input tensor is used for postfiltering, and the output tensor outputTensor resulting from the postfiltering is used for the channels.

[0134] The syntax element nnpfc_inp_format_idc (input tensor format information) indicates how to convert pixel values ​​of the decoded image into input values ​​for the postfilter process. If the value of nnpfc_inp_format_idc is 0, the input values ​​for the postfilter process (especially the input tensor) are real numbers (floating-point value format) specified by IEEE754, and the functions InpY and InpC are specified as follows: The range of values ​​for the input tensor is 0..1.

[0135] InpY(x) = x÷((1< <BitDepthY)-1) InpC(x) = x ÷ ((1< <BitDepthC)-1) If the value of nnpfc_inp_format_idc is 1, the input values ​​to the postfilter process are unsigned integers, and the functions InpY and InpC are specified as follows: The range of values ​​of the input tensor is 0..1<<( inpTensorBitDepthY)-1.

[0136] shift=BitDepthY-inpTensorBitDepth if(inpTensorBitDepth>=BitDepthY) InpY(x)=x<<(inpTensorBitDepthY-BitDepthY) else InpY(x)=Clip3(0,(1< <inpTensorBitDepthY)-1,(x+(1<<(shift-1)))> >shift) shift=BitDepthC-inpTensorBitDepthC if(inpTensorBitDepth>=BitDepthC) InpC(x)=x<<(inpTensorBitDepthC-BitDepthC) else InpC(x) = Clip3(0,(1< <inpTensorBitDepthC)-1,(x+(1<<(shift-1)))> >shift) The value of the syntax element nnpfc_inp_tensor_bitdepth_luma_minus8 plus 8 indicates the pixel bit length of the luma pixel value of the input integer tensor. The values ​​of the variables inpTensorBitDepthY and inpTensorBitDepthC are derived from the syntax elements nnpfc_inp_tensor_bitdepth_luma_minus8 and nnpfc_inp_tensor_bitdepth_chroma_minus8 as follows:

[0137] inpTensorBitDepthY=nnpfc_inp_tensor_bitdepth_luma_minus8+8 inpTensorBitDepthC=nnpfc_inp_tensor_bitdepth_chroma_minus8+8 The syntax element nnpfc_inp_order_idc indicates how the pixel array of the decoded image is ordered as input to the postfilter process. The semantics of nnpfc_inp_order_idc, which ranges from 0 to 3 inclusive, specifies the process for deriving the input tensor inputTensor for each value of nnpfc_inp_order_idc. It also specifies the top-left pixel position of the patch of pixels contained in the input tensor, given the vertical pixel coordinate cTop and the horizontal pixel coordinate cLeft.

[0138] A patch is a rectangular array of pixels from a component of an image (e.g., luma and chroma components).

[0139] The syntax element nnpfc_constant_patch_size_flag, when set to 0, indicates that the postfilter process accepts as input patch sizes that are positive integer multiples of the patch sizes specified by nnpfc_patch_width_minus1 and nnpfc_patch_height_minus1. When nnpfc_patch_width_minus1 is set to 1, the postfilter process accepts as input patch sizes specified by nnpfc_patch_height_minus1.

[0140] The value of the syntax element nnpfc_patch_width_minus1 plus 1 indicates the number of horizontal pixels of the patch size required for input to postfilter processing when the value of nnpfc_constant_patch_size_flag is 1. When the value of nnpfc_constant_patch_size_flag is 0, any positive integer multiple of (nnpfc_patch_width_minus1+1) can be used as the number of horizontal pixels of the patch size used for input to postfilter processing.

[0141] The value of the syntax element nnpfc_patch_height_minus1 plus 1 indicates the number of vertical pixels of the patch size required for input to postfilter processing when the value of nnpfc_constant_patch_size_flag is 1. When the value of nnpfc_constant_patch_size_flag is 0, any positive integer multiple of (nnpfc_patch_height_minus1+1) can be used as the number of vertical pixels of the patch size used for input to postfilter processing.

[0142] The syntax element nnpfc_overlap specifies the number of horizontal and vertical pixels by which adjacent input tensors overlap. The value of nnpfc_overlap must be between 0 and 16383 inclusive.

[0143] The variables inpPatchWidth, inpPatchHeight, outPatchWidth, outPatchHeight, horCScaling, verCScaling, outPatchCWidth, outPatchCHeight, and overlapSize are derived as follows:

[0144] inpPatchWidth=nnpfc_patch_width_minus1+1 inpPatchHeight=nnpfc_patch_height_minus1+1 outPatchWidth=(nnpfc_pic_width_in_luma_samples*inpPatchWidth) / InpPicWidthInLumaSamples outPatchHeight=(nnpfc_pic_height_in_luma_samples*inpPatchHeight) / InpPicHeightInLumaSamples horCScaling=SubWidthC / outSubWidthC verCScaling=SubHeightC / outSubHeightC outPatchCWidth=outPatchWidth*horCScaling outPatchCHeight=outPatchHeight*verCScaling overlapSize=nnpfc_overlap nnpfc_pic_width_in_luma_samples and nnpfc_pic_height_in_luma_samples are the width and height, respectively, of the luma image to which the postfilter specified by nnpfc_id is applied.

[0145] The syntax element nnpfc_padding_type specifies the padding process when referring to pixel locations outside the boundaries of the decoded image.

[0146] The value of nnpfc_padding_type must be between 0 and 15 inclusive.

[0147] If the value of nnpfc_padding_type is 0, the value of pixel positions outside the boundary of the decoded image is set to 0.

[0148] When the value of nnpfc_padding_type is 1, the value of the pixel position outside the boundary of the decoded image is set as the boundary value.

[0149] When the value of nnpfc_padding_type is 2, the values ​​of pixel positions outside the boundary of the decoded image are mirror-reflected with the boundary value as the boundary. To define the value more precisely, the function InpSampleVal is defined. The function InpSampleVal(y,x,picHeight,picWidth,Pic) takes the vertical pixel position y, horizontal pixel position x, image height picHeight, image width picWidth, and pixel array Pic as input, and returns the value of sampleVal derived as follows: InpSampleVal(y, x, picHeight, picWidth, Pic) { if(nnpfc_padding_type==0) if(y<0 || x<0 || y>=picHeight || x>=picWidth) sampleVal = 0 else sampleVal = Pic[y][x] else if(nnpfc_padding_type==1) sampleVal = Pic[Clip3(0, picHeight-1, y)][Clip3(0, picWidth-1, x)] else / *nnpfc_padding_type==2* / sampleVal = Pic[Reflect(picHeight-1, y)][Reflect(picWidth-1, x)] return sampleVal } The syntax element nnpfc_auxiliary_inp_idc indicates whether there is additional data for inputTensor. If it is 0, there is no additional data, and if it is greater than 0, additional data is input to inputTensor.

[0150] The syntax element nnpfc_out_format_idc, when set to 0, indicates that the pixel values ​​output by post-filter processing are floating-point values ​​specified in IEEE754-2019. The functions OutY and OutC, which convert the luminance pixel values ​​and chrominance pixel values ​​output by post-processing into integer values ​​of the pixel bit length, are specified as follows using the pixel bit lengths BitDepthY and BitDepthC, respectively:

[0151] OutY(x)=Clip3(0,(1< <BitDepthY)-1,Round(x*((1<<BitDepthY)-1))) OutC(x) = Clip3(0,(1< <BitDepthC)-1,Round(x*((1<<BitDepthC)-1))) A value of 1 in nnpfc_out_format_idc indicates that the pixel values ​​output by postfilter processing are unsigned integers. The functions OutY and OutC are specified as follows:

[0152] shift=outTensorBitDepth-BitDepthY if(outTensorBitDepth>=BitDepthY) OutY(x)=Clip3(0,(1< <BitDepthY)-1,(x+(1<<(shift-1)))> >shift) else OutY(x)=x<<(BitDepthY-outTensorBitDepthY) shift=outTensorBitDepthC - BitDepthC if(outTensorBitDepthC>=BitDepthC) OutC(x) = Clip3(0,(1< <BitDepthC)-1,(x+(1<<(shift-1)))> >shift) else OutC(x)=x<<(BitDepthC-outTensorBitDepthC) nnpfc_out_tensor_bitdepth_luma_minus8+8 specifies the pixel bit length of pixel values ​​of the output integer tensor. The values ​​of outTensorBitDepthY and outTensorBitDepthC are derived from the syntax elements nnpfc_out_tensor_luma_bitdepth_minus8 and nnpfc_out_tensor_bitdepth_chroma_minus8 as follows:

[0153] outTensorBitDepthY=nnpfc_out_tensor_bitdepth_luma_minus8+8 outTensorBitDepthC=nnpfc_out_tensor_bitdepth_chroma_minus8+8 nnpfc_out_order_idc specifies the output order of pixels resulting from postfilter processing. The semantics of nnpfc_out_order_idc for values ​​0 to 3 are specified.

[0154] The syntax element nnpfc_reserved_zero_bit_b shall be equal to 0.

[0155] The syntax element nnpfc_payload_byte[i] contains the i-th byte of an ISO / IEC 15938-17 compliant bitstream. All nnpfc_payload_byte[i] must be fully ISO / IEC 15938-17 compliant bitstreams.

[0156] (Neural network post-filter applied SEI

[0157] (Map units example) FIG. 12 shows an example syntax configuration of an application SEI.

[0158] For this applied SEI, the header decoding unit 301 and the atlas information encoding unit 102 decode and encode the following syntax elements on a map-by-map basis.

[0159] nnpfa_atlas_id is the identification number (ID) of the atlas to which the applicable SEI applies. It is set to atlasID. Note that when vuh_atlas_id of the V3C unit is set to atlasID, nnpfa_atlas_id is redundant. Therefore, in the syntax of the applicable SEI in FIG. 12 and subsequent figures, nnpfa_atlas_id may not be signaled.

[0160] nnpfa_map_count_minus1 indicates the number of maps transmitting filter information minus 1. Multiple geometry images included in the encoded data are identified by the number mapIdx (j in the figure). Here, the map is used to identify data at the same time. For example, the map may be used when there are two projected patches, Near and Far. Note that as a syntax element representing the number of maps, the syntax element nnpfa_map_count may be encoded and decoded instead of nnpfa_map_count_minus1, which is the value obtained by subtracting 1 from the number of maps.

[0161] nnpfa_enabled_flag[j] indicates whether or not to perform post-filtering indicated by the characteristic SEI on the image of mapIdx=j. If the value is 1, post-filtering is performed, and if the value is 0, it is not performed.

[0162] nnpfa_target_id[j] indicates the ID of the characteristic SEI (the identifier of the characteristic information of the post-filter, the identification information) to be applied to the image of map j (mapIdx=j). The post-filter processing specified by the characteristic SEI having nnpfa_id whose nnpfc_id is equal to nnpfa_target_id is applied to the image of mapIdx=j.

[0163] nnpfa_cancel_flag[i] is a cancellation flag. nnpfa_cancel_flag[j] is 1, which indicates that the persistence of the postfilter set for the image with mapIdx=j in the already decoded NNPFA SEI is to be cancelled. nnpfa_cancel_flag[j] is 0, which indicates that the following syntax element (nnpfa_persistence_flag[j]) is to be transmitted, encoded and decoded.

[0164] nnpfa_persistence_flag[j] indicates the persistence information of the target post-filter. When nnpfa_persistence_flag[j] is 0, it indicates that the target post-filter is applied only to the picture indicated by atlasID. When nnpfa_persistence_flag[j] is 1, it indicates that the target post-filter indicated by nnpfa_target_id[j] is applied to the current picture and all subsequent pictures until one of the following conditions is met: A new CAS is started. The bitstream ends. · NNPFA SEI messages with nnpfa_cancel_flag[j]=1 are output following the current picture in output order, i.e., are cancelled by nnpfa_cancel_flag[j]=1. The following is also acceptable. · NNPFA SEI messages with the same nnpfa_target_id as the current SEI message and with nnpfa_cancel_flag[j]=1 are output following the current picture in output order, i.e., are cancelled by nnpfa_cancel_flag[j]=1.

[0165] The header decoding unit 301 decodes nnpfa_atlas_id, decodes nnpfa_map_count_minus1 for nnpfa_atlas_id, and decodes nnpfa_enabled_flag[j] for the image with mapIdx=j using a loop from j=0 to nnpfa_map_count_minus1. When nnpfa_enabled_flag[j] is 1, the header decoding unit 301 decodes nnpfa_target_id[j] and nnpfa_cancel_flag[j] for the image with mapIdx=j. Furthermore, when nnpfa_cancel_flag[j] is 0, the header decoding unit 301 decodes nnpfa_persistence_flag[j] for the image with mapIdx=j.

[0166] In this configuration, a 3D data decoding device that decodes encoded data and decodes 3D data including attribute information includes a header decoding unit that decodes post-filter characteristic information and post-filter activation information from the encoded data, a geometry image decoding unit that decodes a geometry image from the encoded data, an attribute image decoding unit that decodes an attribute image from the encoded data, and a frame image filter unit that performs post-filter processing of the attribute image or the geometry image, and the header decoding unit decodes the atlas identifier, map number, characteristic information identifier, cancellation flag, and persistence information from the activation information.

[0167] Furthermore, the header decoding unit 31 may decode the map number, and decode the characteristic information of the post filter for each image j ranging from 0 to the map number-1.

[0168] Further, the header decoding unit 31 may decode the map number, and decode a cancel flag for each image j ranging from 0 to the map number minus 1, an identifier of the characteristic information of the post filter for each image j, and the persistence information for each image j. As a result of the above, it is possible to set the cancellation and duration of postfilter processing in units of maps, while at the same time specifying characteristic information of postfilter processing in the same units.

[0169] Furthermore, for all maps (all mapIdx) provided in the coded access unit, syntax elements such as characteristic information and persistence information may not be decoded, and some syntax elements may be derived by the 3D data decoding device 31 (the header decoding unit 301 of the decoding device). Specifically, when the syntax element nnpfa_target_id[j] does not appear in the bitstream, it may be derived as follows.

[0170] nnpfa_target_id[j] = nnpfa_target_id[0] Furthermore, if nnpfa_cancel_flag[j] and nnpfa_persistence_flag[j] do not appear, the following may be applied.

[0171] nnpfa_cancel_flag[j] = nnpfa_cancel_flag[0] nnpfa_persistence_flag[j] = nnpfa_persistence_flag[0] That is, when the identifier of the characteristic information for map j does not appear, the header decoding unit 301 may use the identifier of the characteristic information for map 0 as the identifier of the characteristic information. According to this configuration, the same processing as for an image with mapIdx=0 can be applied without transmitting characteristic information and persistence information of the post filter to be applied to an image with mapIdx>0, thereby achieving the effect of reducing the amount of code.

[0172] In addition, if nnpfa_target_id[j] does not appear, the following may be applied.

[0173] nnpfa_target_id[j] = nnpfa_target_id[j-1] Furthermore, if nnpfa_cancel_flag[j] and nnpfa_persistence_flag[j] do not appear, the following may be applied.

[0174] nnpfa_cancel_flag[j] = nnpfa_cancel_flag[j-1] nnpfa_persistence_flag[j] = nnpfa_persistence_flag[j-1] That is, when the identifier of the characteristic information for map j does not appear, the header decoding unit 301 may use the identifier of the characteristic information for map j-1 as the identifier of the characteristic information. According to this configuration, the same processing as that for the immediately preceding image mapIdx=j-1 can be applied without transmitting characteristic information and persistence information of the post filter to be applied to the image mapIdx=j, which has the effect of reducing the amount of code. The same processing as that for the image mapIdx=nnpfa_map_count_minus1 can be applied without transmitting characteristic information and persistence information of mapIdx>nnpfa_map_count_minus1.

[0175] (Another example of SEI application 1) 14 and 15 show another syntax configuration example of the application SEI. Here, the cancellation flag and persistence information are not transmitted for each attribute image i or map j, but are common.

[0176] More specifically, the header decoding unit 301 and the atlas information encoding unit 102 decode and encode the syntax elements having the following meanings. The other syntax elements are as already explained.

[0177] nnpfa_cancel_flag equal to 1 indicates that the maintenance of the postfilter by the neural network previously set for the image whose atlasID is equal to nnpfa_atlas_id (the image indicated by nnpfa_atlas_id) is canceled. nnpfa_cancel_flag equal to 0 indicates that the following syntax elements are transmitted, encoded, and decoded.

[0178] nnpfa_persistence_flag indicates the persistence information of the post filter for the picture indicated by nnpfa_atlas_id. When nnpfa_persistence_flag is 0, it indicates that the target post filter is applied only to the picture indicated by nnpfa_atlas_id. When nnpfa_persistence_flag is 1, it indicates that the target post filter is applied to the current picture and all subsequent pictures until one of the following conditions is met: A new CAS is started. The bitstream ends. · NNPFA SEI messages with the same atlasID and nnpfa_cancel_flag=1 are output following the current picture in output order. That is, they are canceled by nnpfa_cancel_flag=1.

[0179] The configuration of FIG. 14 has the advantage that it is possible to specify on / off of postfiltering for each map while setting the cancellation and duration of postfiltering regardless of the map.

[0180] The configuration in FIG. 15 has the advantage that it is possible to specify the on / off of post-filtering for each attribute image and map while setting the cancellation and duration of post-filtering regardless of the attribute image and map.

[0181] (Applying Geometry Images SEI) When filtering a geometry image, an adaptive SEI dedicated to the geometry image may be transmitted. More specifically, as shown in Figures 12 and 14, nnpfa_target_id indicating the ID of the characteristic SEI is decoded, encoded, or transmitted on a map basis.

[0182] Figure 16 is an example of the syntax of an adaptation SEI when applying a filter to a geometry image. The name of nn_post_filter_activation is changed to nn_geometory_post_filter_activation, and the applicable SEI for geometry images is distinguished by the value of payloadType of the SEI. PayloadType is encoded, decoded, or transmitted by sm_payload_type_byte, which is the syntax of the header part of the SEI.

[0183] In this configuration, the header decoding unit decodes characteristic information of the post-filter for the geometry image for each image j from 0 to the number of maps minus 1, and further the header decoding unit decodes the number of attributes and decodes an identifier of the characteristic information of the post-filter for each attribute image i from 0 to the number of attributes minus 1.

[0184] In the application SEI of the above embodiment, examples of the characteristic SEI and the application SEI when a linear filter is used will be described in detail below.

[0185] In the following embodiments, the header decoding unit 301 or the atlas information encoding unit 102 decodes syntax elements from encoded data or encodes syntax elements into encoded data.

[0186] (Linear post-filter characteristic SEI) 17A shows a part of the syntax of filter information, nn_post_filter_characteristics(payloadSize) (characteristics SEI), when the post-filter is a linear filter. The argument payloadSize indicates the number of bytes of this SEI message.

[0187] The persistence scope for applying the feature SEI is CAS, similar to the neural network feature SEI, and the explanation is omitted here.

[0188] nnpfc_id is an identifier used to identify the postfilter.

[0189] filter_param() is coefficient information of a linear filter, the 3D data encoding device 11 encodes and signals the filter coefficient information, and the 3D data decoding device 31 decodes it (hereinafter, the terms encoding and decoding will be omitted and it will simply be described as signaling). Specific examples of syntax elements are shown in Fig. 17(b).

[0190] num_filters_signalled_minus1 plus 1 is the number of linear filters specified by nnpfc_id. The number of linear filters may be the number of classes of the area obtained by dividing the screen. Each class corresponds to classId described in (ALF unit 610). For the linear filter indicated by classId, Ncoeff flt_coeff_abs and flt_coeff_sign are signaled by binarization of ue(v) and u(1), respectively. Ncoeff is the number of filter coefficients to be coded and may be the length of the filter (number of taps). If the filter coefficients are symmetrical with respect to the target pixel, Ncoeff may be a value obtained by adding 1 to 1 / 2 of the filter length. In other words, the number of taps of a 3x3 filter is usually 9, but if there is symmetry, the four filter coefficients except the one at the center and the other four filter coefficients have the same value, so only five signals (9 / 2+1) are sufficient. Note that coeff may be directly signaled using se(v) binarization without dividing into absolute value and sign.

[0191] flt_coeff_abs[i][j] represents the absolute value of the j-th coefficient of the filter indicated by classId=i. If flt_coeff_abs[i][j] is not signaled, it is set to 0.

[0192] flt_coeff_sign[i][j] indicates the sign (positive or negative) of the j-th coefficient of the filter indicated by classId=i. When lmpf_coeff_sign[i][j]=0, the coefficient is positive, and when lmpf_coeff_sign[i][j]=1, the coefficient is negative. When lmpf_coeff_sign[i][j] is not signaled, it is set to 0. The filter coefficient coeff[classId][] described in (ALF unit 610) is derived by the following formula.

[0193] coeff[classId][j] = flt_coeff_abs[i][j] * (1 - 2 * flt_coeff_sign[i][j]) classId=0..num_filters_signalled_minus1, j=0..Ncoeff-1.

[0194] In the above configuration, if a characteristic SEI having an nnpfa_id equal to the nnpfa_target_id specified in the application SEI has an NN model, the NN filter is applied. Otherwise, if the characteristic SEI has a filter coefficient filter_param, a linear filter is applied.

[0195] (Example of SEI common to NN filters and linear filters) FIG. 18 shows an example of the syntax of nn_post_filter_characteristics(payloadSize), which is a characteristic SEI common to NN filters and linear filters.

[0196] The nnpfc_id is an identifier used to identify the postfilter. The SEI message has a specific nnpfc_id for the current CLVS.

[0197] nnpfc_mode_idc=100 indicates that the post-filter applied by this characteristic SEI is a linear filter. When nnpfc_mode_idc=100, filter_param() is signaled. An example of filter_param() is shown in Figure 17(b). nnpfc_mode_idc!=100 indicates that the underlying post-filter associated with nnpfc_id is an NN filter. When nnpfc_mode_idc!=100, the same syntax elements as the NN filter shown in Figure 11 are signaled.

[0198] In the above configuration, instead of signaling individual characteristic SEIs for each NN filter and linear filter, a common SEI is signaled, and by switching between signaling the syntax of the NN model and signaling the syntax of the filter coefficient filter_param, it has the effect of making it easier to select on the encoding side.

[0199] (SEI with linear post-filter applied) <Structure that distinguishes between linear and NN filters> FIG. 19(a) shows the syntax of a characteristic SEI message when the post-filter is a linear filter, and (b) and (c) show the syntax of an application SEI message when the post-filter is a linear filter. The argument payloadSize indicates the number of bytes of this SEI message. The syntax elements are as described in (a) in FIG. 17, (b) in FIG. 14, and (c) in FIG. 15. The syntax configuration is the same except that "nnpfa_" is replaced with "lmpfa_". In this configuration, the filter parameters are transmitted in the characteristic SEI, and the persistence information and the ID of the characteristic SEI are transmitted in the application SEI. In this configuration, the application SEI of the linear filter specifies the characteristic SEI of the linear filter, and the application SEI of the NN filter specifies the characteristic SEI of the NN filter.

[0200] For this applied SEI, the header decoding unit 301 and the atlas information encoding unit 102 decode and encode the following syntax elements on a map basis and an attribute image basis.

[0201] lmpfa_atlas_id is the identification number (ID) of the atlas to which the applicable SEI applies.

[0202] lmpfa_map_count_minus1 indicates the number of maps transmitting filter information minus 1. lmpfa_attribute_count_minus1 indicates the number of attribute images transmitting filter information minus 1. The maps and attribute images are the same as those described in (Neural Network Post Filter Application SEI), and their explanations are omitted. In addition, unless otherwise specified, lmpfa_enabled_flag, lmpfa_target_id, lmpfa_cancel_flag, and lmpfa_persistence_flag have the same definitions as nnpfa_enabled_flag, nnpfa_target_id, nnpfa_cancel_flag, and nnpfa_persistence_flag described in (Neural Network Post Filter Application SEI), and their explanations are omitted.

[0203] Figure 19(b) is an example of the syntax of an applied SEI. In (b), lmpfa_atlas_id and lmpfa_cancel_flag, which are common to this applied SEI, are signaled. If lmpfa_cancel_flag is 0, lmpfa_persistence_flag and lmpfa_map_count_minus1 are signaled. In addition, lmpfa_enabled_flag[j] is signaled on a map basis. If lmpfa_enabled_flag[j] is 1, lmpfa_target_id[j] is signaled. Syntax elements that are not signaled are set to 0.

[0204] Figure 19(c) is another example of the syntax of the applied SEI. In (b), lmpfa_enabled_flag and lmpfa_target_id[j] are signaled on a map-by-map basis, but in (c), they are signaled on a map-by-map and attribute image-by-attribute basis. Specifically, when lmpfa_cancel_flag is 0, lmpfa_persistence_flag, lmpfa_attribute_count_minus1, and lmpfa_map_count_minus1 are signaled. Furthermore, lmpfa_enabled_flag[i][j] is signaled on a map-by-map and attribute image-by-attribute basis. When lmpfa_enabled_flag[i][j] is 1, lmpfa_target_id[i][j] and [i][j] are signaled. Syntax elements that are not signaled are set to 0.

[0205] Apply the post-filtering specified by the linear filter characteristics SEI(lm_) with lmpfc_id having the same value as the lmpfa_target_id[i][j] signaled in the apply SEI to the image with mapIdx=j, attrIdx=i.

[0206] As a result, filter coefficients suitable for various maps and attribute images are added to the bit stream and signaled, improving image quality.

[0207] (Example of Linear Post Filter SEI) In this embodiment, a method is described in which persistence information and filter coefficients are directly signaled in the post-filter SEI without distinguishing between the characteristic SEI and the application SEI.

[0208] Figure 20 shows an example of transmitting filter parameters directly in the post-filter SEI. Here, instead of signaling an identifier lmpfa_target_id that identifies the filter characteristics, the filter coefficient filter_param() that is the filter characteristic is signaled. The syntax of filter_param(j) is shown in Figure 21(a).

[0209] Fig. 21(a) is an example of signaling filter coefficients in units of maps. lmpf_coeff_abs[mapIdx][i][j] and lmpf_coeff_sign[mapIdx][i][j] are signaled in units of mapIdx. The rest is the same as the description of filter_param(), so the description will be omitted.

[0210] In this embodiment as well, image quality is improved because filter coefficients suitable for various maps are added to the bit stream and signaled.

[0211] (Linear postfilter SEI example 2) In this embodiment, an alternative method is described in which the characteristic SEI and the application SEI are not differentiated, and the persistence information and the filter coefficients are directly signaled in the postfilter SEI.

[0212] Fig. 20(b) is an example of signaling filter coefficients on a map or attribute basis. The syntax of filter_param(i,j) is shown in Fig. 21(b).

[0213] In Figure 11(b), the filter coefficient flter_coeff(attrIdx, mapIdx) is signaled on a map and attribute basis. lmpf_coeff_abs[attrIdx][mapIdx][i][j] and lmpf_coeff_sign[attrIdx][mapIdx][i][j] are signaled on a mapIdx and attrIdx basis. The rest is the same as the description of filter_param(), so the description will be omitted.

[0214] In this embodiment as well, the image quality is improved by adding filter coefficients suitable for various maps and attribute images to the bit stream and signaling them.

[0215] Furthermore, in this embodiment, the filter coefficients are directly signaled in the application SEI, so there is no need to separately signal the filter coefficients in the characteristics SEI, and the filter to be used for the post filter can be signaled only in the application SEI.

[0216] (Example 1 of SEI specifying the immediately preceding post-filter) FIG. 22 is an example of signaling by adding nnpfa_prev_flt_flag to FIG. 19. nnpfa_prev_flt_flag is a flag indicating whether the filter is the same as the filter used immediately before. When nnpfa_prev_flt_flag=1, it indicates that the filter used immediately before is used. Therefore, nnpfa_target_id is not signaled and the immediately previous value is copied. When nnpfa_prev_flt_flag=0, it indicates that a filter other than the filter used immediately before is used. Therefore, when nnpfa_prev_flt_flag=0, nnpfa_target_id is signaled. FIG. 22(a) is an example of signaling nnpfa_prev_flt_flag[j] on a map basis, and the signaled filter may be applied to the geometry image. On the other hand, FIG. 22(b) is an example of signaling nnpfa_prev_flt_flag[i][j] on a map basis or attribute image basis, and the signaled filter may be applied to the attribute image.

[0217] (Example 2 of SEI specifying the immediately preceding post-filter) FIG. 23 is an example of signaling by adding nnpfa_prev_flt_flag to FIG. 20. nnpfa_prev_flt_flag is a flag indicating whether the filter is the same as the filter used immediately before. When nnpfa_prev_flt_flag=1, it indicates that the filter used immediately before is used. Therefore, nnpfa_target_id is not signaled and the immediately previous value is copied. When nnpfa_prev_flt_flag=0, it indicates that a filter other than the filter used immediately before is used. Therefore, when nnpfa_prev_flt_flag=0, filter_param() is signaled. FIG. 23(a) is an example of signaling nnpfa_prev_flt_flag[j] on a map basis, and the signaled filter may be applied to the geometry image. On the other hand, FIG. 23(b) is an example of signaling nnpfa_prev_flt_flag[i][j] on a map basis or attribute image basis, and the signaled filter may be applied to the attribute image.

[0218] (Example 1 of SEI common to NN filters and linear filters) Instead of setting an individual application SEI for each of the NN filter and the linear filter, a common SEI may be set.

[0219] In this embodiment, a combination of an application SEI common to an NN filter and a linear filter, and a characteristic SEI for an NN filter and a characteristic SEI for a linear filter will be described.

[0220] Figure 24 is an example of the syntax of the application SEI common to the NN filter and the linear filter. In Figure 24(a), nnpfa_filter_type is additionally signaled to the syntax in Figure 16. nnpfa_filter_type is a parameter that indicates whether the post filter is a linear filter or an NN filter.

[0221] If nnpfa_filter_type=0, the post filter is a linear filter and signals lmpfa_target_id[j]. Then, for each map, the filter_param() indicated by the lmpfc_id same as lmpfa_target_id[j] is decoded from the lmpfc_id of lmpfc_post_filter_characteristics, and the filter coefficient of the linear filter is derived. An example of filter_param() is shown in Figure 17(b).

[0222] If nnpfa_filter_type=1, the post filter is an NN filter and signals nnpfa_target_id[j]. Then, from the nnpfc_id of nnpfc_post_filter_characteristics, the filter parameters indicated by the same nnpfc_id as nnpfa_target_id[j] are decoded to derive the filter coefficients of the NN filter. An example of the filter parameters indicated by nnpfc_id is shown in Figure 11. In Figure 24(b), nnpfa_filter_type is additionally signaled to the syntax in Figure 15. In (a), lmpfa_target_id[j] is signaled on a map basis, but in (b), lmpfa_target_id[i][j] is signaled on a map basis and on an attribute image basis. The rest of the configuration is the same.

[0223] As described above, in this embodiment, the syntax can be simplified by combining the application SEI common to the NN filter and the linear filter, and the characteristic SEI for the NN filter and the characteristic SEI for the linear filter.

[0224] (Configuration of 3D data encoding device) FIG. 9 is a functional block diagram showing a schematic configuration of a 3D data encoding device 11 according to an embodiment of the present invention.

[0225] The 3D data encoding device 11 is composed of a patch generation unit 101, an atlas information encoding unit 102, an occupancy map generation unit 103, an occupancy map encoding unit 104, a geometry image generation unit 105, a geometry image encoding unit 106, an attribute image generation unit 108, an attribute image encoding unit 109, a frame image filter parameter derivation unit 110, and a multiplexing unit 111. The 3D data encoding device 11 inputs a point cloud or a mesh as 3D data, and outputs encoded data.

[0226] The patch generation unit 101 inputs 3D data, generates a set of patches (rectangular images in this case), and outputs it. Specifically, the 3D data is divided into multiple regions, and each region is projected onto one of the planes of a 3D bounding box set in the 3D space to generate multiple patches. The patch generation unit 101 outputs information about the 3D bounding box (coordinates, size, etc.) and information about mapping onto the projection surface (projection surface, coordinates, size, whether or not each patch is rotated, etc.) as atlas information.

[0227] The atlas information encoding unit 102 encodes the atlas information output from the patch generating unit 101 and outputs an atlas information encoded stream. The atlas information encoding unit 102 sets the value of the atlasID to which the above SEI is applied to nnpfa_atlas_id.

[0228] The occupancy map generator 103 inputs a set of patches output from the patch generator 101, and generates an occupancy map that shows the valid area (area where 3D data exists) of each patch as a 2D binary image (e.g., valid area is 1, invalid area is 0). Note that other values ​​such as 255 and 0 may be used as the values ​​of the valid area and invalid area.

[0229] The occupancy map encoding unit 104 inputs the occupancy map output from the occupancy map generation unit 103, and outputs an occupancy map encoded stream and an encoded occupancy map. As an encoding method, VVC, HEVC, or the like is used.

[0230] The geometry image generating unit 105 generates a geometry image storing a depth value for the projection surface of each patch based on the 3D data, the occupancy map, the coded occupancy map, and the atlas information. The geometry image generating unit 105 derives a point having a minimum depth for the projection surface among the points projected onto the pixel g(x,y) as p_min(x,y,z). In addition, the geometry image generating unit 105 derives a point having a maximum depth among the points projected onto the pixel g(x,y) and at a predetermined distance d from p_min(x,y,z) as p_max(x,y,z). The geometry image projected onto all pixels of the projection surface with p_min(x,y,z) is set as the geometry image of the Near layer. The geometry image projected onto all pixels of the projection surface with p_max(x,y,z) is set as the geometry image of the Far layer.

[0231] The geometry image encoding unit 106 receives a geometry image, and outputs an encoded geometry image stream and an encoded geometry image. The encoding method used is VVC, HEVC, or the like.

[0232] The attribute image generating unit 108 generates an attribute image storing color information (e.g., YUV values, RGB values, etc.) for the projection surface of each patch based on the 3D data, the coded occupancy map, the coded geometry image, and the atlas information. The attribute image generating unit 108 obtains the value of the attribute corresponding to the point p_min(x,y,z) with the minimum depth calculated by the geometry image generating unit 106, and sets the attribute image projected with the value as the attribute image of the Near layer. The attribute image obtained in a similar manner for p_max(x,y,z) is set as the attribute image of the Far layer.

[0233] The attribute image encoding unit 109 receives an attribute image, and outputs an attribute image encoded stream and an encoded attribute image. The encoding method used is VVC, HEVC, or the like.

[0234] The frame image filter parameter derivation unit 110 inputs the encoded attribute image and the original attribute image, or the encoded geometry image and the original geometry image, and selects or derives and outputs optimal filter parameters in the NN filter processing.

[0235] The frame image filter parameter derivation unit 110 sets the values ​​of nnpfa_enabled_flag, nnpfa_target_id, nnpfa_cancel_flag, nnpfa_persistance_flag, nnpfa_weight_block_size_idx, nnpfa_weight_map_width_minus1, nnpfa_weight_map_height_minus1, and nnpfa_weight_map to the SEI. It may also set the values ​​of nnpfa_strength_block_size_idx, nnpfa_strength_map_width_minus1, nnpfa_strength_map_height_minus1, and nnpfa_strength_map to the SEI.

[0236] The multiplexing unit 111 inputs the filter parameters output from the frame image filter parameter derivation unit 110, and outputs them in a predetermined format. The predetermined format is, for example, SEI, which is auxiliary extension information of video data, ASPS and AFPS, which are information specifying a data structure in the V3C standard, and ISOBMFF, which is a media file format standard. The multiplexing unit 111 also multiplexes the atlas information coded stream, the occupancy map coded stream, the geometry image coded stream, the attribute image coded stream, and the above filter parameters, and outputs the result as coded data. As a multiplexing method, a byte stream format, ISOBMFF, etc. are used.

[0237] FIG. 13 is a block diagram showing the relationship between the 3D data encoding device 11, the 3D data decoding device 31, and the SEI message according to the embodiment of the present invention.

[0238] The 3D data encoding device 11 is composed of a VVC encoding unit and an SEI encoding unit. In the configuration of Fig. 9, the VVC encoding unit corresponds to the occupancy map encoding unit 104, the geometry image encoding unit 106, and the attribute image encoding unit 109. In the configuration of Fig. 9, the SEI encoding unit corresponds to the atlas information encoding unit 102.

[0239] The 3D data decoding device 31 is composed of a VVC decoding unit, an SEI decoding unit, a switch, and a postfilter unit. In the configuration of Fig. 6, the VVC decoding unit corresponds to the occupancy map decoding unit 303, the geometry image decoding unit 304, and the attribute image decoding unit 307. In the configuration of Fig. 6, the SEI decoding unit corresponds to the header decoding unit 301. In the configuration of Fig. 6, the postfilter unit corresponds to the frame image filter unit 308. The frame image filter unit 308 may include either an NN filter unit 611 or an ALF unit 610, or both.

[0240] The occupancy map, geometry image, and attribute image generated from the 3D data are coded by the VVC coding unit. The coded data is decoded by the VVC decoding unit, and a decoded image is reconstructed. The SEI coding unit generates a characteristic SEI and an applied SEI from the 3D data. The SEI decoding unit decodes these SEI messages. The applied SEI is input to a switch to specify the image to be postfiltered, and only the image to be postfiltered is input to the postfilter unit. The characteristic SEI is input to the postfilter unit to specify the postfilter to be applied to the decoded image. The postfiltered image or the decoded image is displayed on the 3D data display unit 41 (Figure 1).

[0241] Although one embodiment of the present invention has been described in detail above with reference to the drawings, the specific configuration is not limited to the above, and various design changes, etc. are possible within the scope that does not deviate from the gist of the present invention.

[0242] The present invention is not limited to the above-described embodiment, and various modifications are possible within the scope of the claims. In other words, the technical scope of the present invention also includes embodiments obtained by combining technical means that are appropriately modified within the scope of the claims. [Industrial Applicability]

[0243] The embodiments of the present invention can be suitably applied to a 3D data decoding device that decodes encoded data in which 3D data is encoded, and a 3D data encoding device that generates encoded data in which 3D data is encoded. Also, the embodiments of the present invention can be suitably applied to the data structure of encoded data generated by the 3D data encoding device and referenced by the 3D data decoding device. [Explanation of symbols]

[0244] 11 3D data encoding device 101 Patch Generation Unit 102 Atlas Information Encoding Unit 103 Occupancy map generator 104 Occupancy map coding unit 105 Geometry Image Generation Unit 106 Geometry Image Encoding Unit 108 Attribute Image Generator 109 Attribute Image Encoding Unit 110 Frame image filter parameter derivation unit 111 Multiplexer 21 Network 31 3D data decoding device 301 Header Decoding Unit 302 Atlas Information Decoding Unit 303 Occupancy Map Decoding Unit 304 Geometry Image Decoding Unit 306 Geometry Reconstruction Unit 307 Attribute Image Decoding Unit 308 Frame Image Filter Section 309 Attribute Reconstruction Unit 310 3D Data Reconstruction Department 41 3D data display device

Claims

1. In a 3D data decoding device that generates 3D data from encoded data, A header decoding unit decodes the post-filter activation information from the above encoded data, The system includes a frame image filter unit that performs post-filtering on the geometry image or attribute image output from the above encoded data, The header decoding section described above is: The first syntax element indicating the identifier of the V3C unit atlas, the second syntax element for deriving the number related to the post-filter included in the activation information, and the third syntax element indicating whether the post-filter is applied to a particular map are decoded, A 3D data decoding device characterized by decoding a fourth syntax element for identifying the map to which the post-filter is applied, and a fifth syntax element indicating whether or not to cancel the maintenance of the post-filter, depending on the value of the third syntax element described above.

2. The 3D data decoding device according to claim 1, characterized in that the header decoding unit decodes a sixth syntax element, which is included in the activation information, indicating the strength of the post-filter for the input tensor.

3. A 3D data encoding device for encoding 3D data, A header encoding unit that encodes the post-filter activation information, It comprises a frame image filter unit that performs post-filtering on a geometry image or attribute image, The header encoding part described above is: Encode a first syntax element indicating the identifier of the V3C unit atlas, a second syntax element for deriving the number related to the post-filter included in the activation information, and a third syntax element indicating whether the post-filter is applied to a particular map, A 3D data encoding device characterized by encoding a fourth syntax element for identifying the map to which the post-filter is applied, and a fifth syntax element indicating whether or not to cancel the maintenance of the post-filter, depending on the value of the third syntax element described above.