Encoding method, decoding method, encoding device, and decoding device
The method stabilizes the encoding and decoding of three-dimensional meshes by applying offset and scaling processes on attribute information, ensuring accurate transformation and efficient transmission.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
- Filing Date
- 2026-01-09
- Publication Date
- 2026-07-23
AI Technical Summary
Existing methods for encoding and decoding three-dimensional meshes are inadequate, particularly in handling negative and decimal values in attribute information, leading to inefficiencies in transmission and storage.
A method that involves performing an offset process and scaling operations on attribute information, followed by encoding, and including offset and scale values in the bitstream for inverse transformation, ensuring stable mapping and accurate decoding.
Stable encoding and decoding of three-dimensional meshes are achieved, even with negative or decimal values, by using offset and scale values for inverse transformation, maintaining accuracy and efficiency in transmission and storage.
Smart Images

Figure JP2026000600_23072026_PF_FP_ABST
Abstract
Description
Encoding method, decoding method, encoding device, and decoding device
[0001] This disclosure relates to encoding methods, etc.
[0002] Patent Document 1 proposes a method and apparatus for encoding and decoding three-dimensional mesh data.
[0003] Japanese Patent Publication No. 2006-187015
[0004] Further improvements are desired for the encoding or decoding of three-dimensional meshes. This disclosure aims to improve the encoding or decoding of three-dimensional meshes.
[0005] An encoding method according to one aspect of the present disclosure acquires first attribute information, which is attribute information associated with a point in the base mesh of a three-dimensional mesh, performs a first transformation process which involves performing an offset process that performs addition or subtraction on the numerical value indicated by the first attribute information, and a scaling process which performs at least one of multiplication, division, and shift operations on the numerical value indicated by the first attribute information, then encodes the first attribute information, and generates a bitstream which includes encoded first attribute information including the encoded first attribute information, and an offset value and a scale value used in a first inverse transformation process to obtain the first attribute information from the encoded first attribute information.
[0006] A decoding method according to one aspect of the present disclosure obtains a bitstream including encoded first attribute information, which is attribute information associated with points in the base mesh of a three-dimensional mesh, and an offset value and a scale value used for the inverse transformation of the first attribute information obtained from the encoded first attribute information; decodes the encoded first attribute information from the bitstream to obtain the first attribute information, and obtains the offset value and the scale value; and performs a first inverse transformation process on the numerical value indicated by the first attribute information, which includes a scaling process that performs at least one of multiplication / division and shift operations based on the scale value and an offset process that performs addition / subtraction based on the offset value, thereby obtaining the first attribute information.
[0007] These comprehensive or specific embodiments may be implemented as a system, device, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, device, integrated circuit, computer program, and recording medium.
[0008] This disclosure may contribute to improvements in encoding processing and other related processes for three-dimensional meshes.
[0009] This is a conceptual diagram showing a three-dimensional mesh according to Embodiment 1. This is a conceptual diagram showing the basic elements of a three-dimensional mesh according to Embodiment 1. This is a conceptual diagram showing a mapping according to Embodiment 1. This is a block diagram showing an example configuration of an encoding / decoding system according to Embodiment 1. This is a block diagram showing an example configuration of an encoding device according to Embodiment 1. This is a block diagram showing another example configuration of an encoding device according to Embodiment 1. This is a block diagram showing an example configuration of a decoding device according to Embodiment 1. This is a block diagram showing another example configuration of a decoding device according to Embodiment 1. This is a conceptual diagram showing an example configuration of a bitstream according to Embodiment 1. This is a conceptual diagram showing yet another example configuration of a bitstream according to Embodiment 1. This is a block diagram showing a specific example of an encoding / decoding system according to Embodiment 1. This is a conceptual diagram showing an example configuration of point cloud data according to Embodiment 1. This is a conceptual diagram showing an example data file for point cloud data according to Embodiment 1. This is a conceptual diagram showing an example configuration of mesh data according to Embodiment 1. This is a conceptual diagram showing an example data file for mesh data according to Embodiment 1. This is a conceptual diagram showing the types of three-dimensional data according to Embodiment 1. This is a block diagram showing an example configuration of a three-dimensional data encoder according to Embodiment 1. This is a block diagram showing an example configuration of a three-dimensional data decoder according to Embodiment 1. This is a block diagram showing another example configuration of a three-dimensional data encoder according to Embodiment 1. This is a block diagram showing another configuration example of the three-dimensional data decoder according to Embodiment 1. This is a conceptual diagram showing a specific example of the encoding process according to Embodiment 1. This is a conceptual diagram showing a specific example of the decoding process according to Embodiment 1. This is a block diagram showing an implementation example of the encoding device according to Embodiment 1. This is a block diagram showing an implementation example of the decoding device according to Embodiment 1. This is a block diagram showing a different configuration example of the encoding / decoding system according to Embodiment 1. This is a block diagram showing another configuration example of the encoding device according to Embodiment 1. This is a block diagram showing another configuration example of the decoding device according to Embodiment 1. This is a block diagram showing another configuration example of the encoding device according to Embodiment 1. This is a flowchart showing the processing of the encoding device according to Embodiment 1.This is an explanatory diagram conceptually showing the encoding of a mesh frame according to Embodiment 1. This is a flowchart showing the processing of a decoding device according to Embodiment 1. This is an explanatory diagram conceptually showing the decoding of a mesh frame according to Embodiment 1. This is a block diagram showing an example of the configuration of a decoding device according to Embodiment 1. This is a block diagram showing an example of the configuration of a decoding device according to Embodiment 1. This is an explanatory diagram showing an example of subdivision according to Embodiment 1. This is an explanatory diagram showing an example of vertex displacement after subdivision according to Embodiment 1. This is an explanatory diagram showing an example of vertices of the original mesh according to Embodiment 1. This is an explanatory diagram showing an example of a mesh according to Embodiment 1. This is an explanatory diagram showing an example of subdivision of a mesh according to Embodiment 1. This is a first explanatory diagram showing an example of packing displacement information into an image frame according to Embodiment 1. This is a second explanatory diagram showing an example of packing displacement information into an image frame according to Embodiment 1. This is a third explanatory diagram showing an example of packing displacement information into an image frame according to Embodiment 1. This is a diagram for explaining an example of attribute information in mesh data according to Embodiment 1. This is a diagram for explaining an example of attribute information in mesh data according to Embodiment 1. This is a diagram for explaining an example of attribute information in mesh data according to Embodiment 1. This is a diagram for explaining a specific example of an attribute information encoding method. This is a diagram for explaining a specific example of an attribute information encoding method. This is a diagram illustrating a specific example of an attribute information encoding method. This is a diagram illustrating a specific example of an attribute information encoding method. This is a diagram illustrating a specific example of an attribute information encoding method. This is a diagram illustrating a specific example of an attribute information encoding method. This is a diagram showing an example of a flowchart for encoding attribute information of a three-dimensional point according to Embodiment 1. This is a flowchart showing an example of the calculation process for the predicted value of the third prediction mode according to Embodiment 1. This is a diagram showing an example of a flowchart for decoding attribute information of a three-dimensional point according to Embodiment 1. This is a diagram showing an example of the syntax of mesh_attribute_coding according to Embodiment 1. This is a diagram showing an example of the syntax of mesh_attribute_header according to Embodiment 1. This is a diagram showing an example of the syntax of mesh_attribute_data according to Embodiment 1.This figure shows an example of attribute information relating to a modification of Embodiment 1. This figure shows an example of the syntax of mesh_attribute_header relating to a modification of Embodiment 1. This figure shows an example of a geometry map relating to Embodiment 1. This figure shows an example of a texture map relating to Embodiment 1. This figure shows an example of the syntax of mesh_position_coding relating to Embodiment 1. This figure shows an example of the syntax of mesh_position_header relating to Embodiment 1. This figure shows an example of the syntax of mesh_position_data relating to Embodiment 1. This figure shows an example of the configuration of an encoding device relating to Embodiment 1. This flowchart shows an example of an encoding method by an encoding device relating to Embodiment 1. This figure shows an example of the configuration of a decoding device relating to Embodiment 1. This flowchart shows an example of a decoding method by a decoding device relating to Embodiment 1. This flowchart shows another example of the calculation process of the predicted value for the third prediction mode relating to Embodiment 1. This figure shows an example of a vector in three-dimensional space when making fine predictions relating to Embodiment 1. This figure shows an example of a vector in two-dimensional plane when making fine predictions relating to Embodiment 1. This figure shows an example of calculating a predicted value when the magnitude of the absolute value of dot(gNgP, gNgC) exceeds the threshold TH relating to Embodiment 1. This figure shows an example of calculating a predicted value when the magnitude of the absolute value of dot(gNgP, gNgC) according to Embodiment 1 is less than or equal to the threshold TH. This figure shows another example of calculating a predicted value according to Embodiment 1. This flowchart shows an example of the process for calculating a predicted value according to Embodiment 1. This figure shows an example of a vector in three-dimensional space when making a fine prediction according to a modified example of Embodiment 1. This figure shows an example of a vector in two-dimensional plane when making a fine prediction according to a modified example of Embodiment 1. This figure shows another example of calculating a predicted value according to Embodiment 1. This flowchart shows another example of the process for calculating a predicted value according to Embodiment 1. This figure illustrates another specific example of the square root calculation process (sqrt) according to Embodiment 1. This figure illustrates another specific example of the square root calculation process (sqrt) according to Embodiment 1.This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the square root calculation process (sqrt) according to Embodiment 1. This is a diagram illustrating another specific example of the configuration of the encoding device according to Embodiment 1. This is a flowchart showing another example of the encoding method by the encoding device according to Embodiment 1. This is a diagram showing another example of the configuration of the decoding device according to Embodiment 1. This is a flowchart showing another example of the decoding method by the decoding device according to Embodiment 1. This is a block diagram showing the configuration of the three-dimensional data encoding device according to Embodiment 2. This is a block diagram showing the configuration of the three-dimensional data decoding device according to Embodiment 2. This is a flowchart showing the processing procedure of the three-dimensional data encoding device according to Embodiment 2. This is a flowchart showing the processing procedure of the three-dimensional data decoding device according to Embodiment 2. This is a diagram showing an example of the syntax of the VPS (Volumetric Parameter Set) according to Embodiment 2. This is a diagram showing an example of the syntax of the Sequence Parameter Set for VDDMC (SPS_VDDMC) according to Embodiment 2. This is a block diagram for explaining another example of the processing of the three-dimensional data encoding device according to Embodiment 2. This is a block diagram for explaining another example of the processing of the three-dimensional data decoding device according to Embodiment 2.This is a flowchart showing the processing procedure of a three-dimensional data encoding device according to Embodiment 2. This is a flowchart showing the processing procedure of a three-dimensional data decoding device according to Embodiment 2. This is a diagram showing an example of the syntax of an extended information SEI according to Embodiment 2. This is a diagram showing an example of the syntax of the payload of an attribute conversion SEI according to Embodiment 2. This is a diagram showing an example of the syntax of a VDDMC SEI according to Embodiment 2. This is a diagram showing an example of the syntax of a VDDMC SEI according to Embodiment 2. This is a diagram showing an example of the syntax of a VDDMC conversion information SEI according to Embodiment 2. This is an example of the syntax of an attribute conversion SEI according to Embodiment 2. This is a diagram showing an example of the syntax of transform_info() according to Embodiment 2. This is a diagram showing another example of transform_info(attribute_index) according to Embodiment 2. This is a diagram showing an example of the configuration of an encoding device in Embodiment 2. This is a flowchart showing an example of an encoding method by an encoding device in Embodiment 2. This is a diagram showing an example of the configuration of a decoding device in Embodiment 2. This is a flowchart showing an example of a decoding method by a decoding device in Embodiment 2.
[0010] <Summary of this disclosure> For example, a three-dimensional (3D) mesh is used in computer graphics images. For example, a computer graphics image may consist of multiple frames that are different in time, and each frame may be represented by a three-dimensional mesh.
[0011] Furthermore, a three-dimensional mesh consists of vertex information indicating the position of each of several vertices in three-dimensional space, connection information indicating the connections between the multiple vertices, and attribute information indicating the attributes of each vertex or face. Each face is constructed according to the connection relationships of the multiple vertices. Various computer graphics images can be represented using such a three-dimensional mesh.
[0012] Furthermore, efficient encoding and decoding of three-dimensional meshes are expected for the transmission and storage of these meshes. Arithmetic coding and decoding may be used for efficient encoding and decoding of three-dimensional meshes.
[0013] Further improvements are desired in the encoding or decoding of three-dimensional data. This disclosure aims to improve the encoding or decoding of three-dimensional data.
[0014] The following describes examples of inventions that can be obtained from the disclosures in this specification, and explains the effects and other benefits that can be obtained from such inventions.
[0015] An encoding method according to a first aspect of the present disclosure acquires first attribute information, which is attribute information associated with a point in the base mesh of a three-dimensional mesh, performs a first transformation process which involves performing an offset process that performs addition or subtraction on the numerical value indicated by the first attribute information, and a scaling process which performs at least one of multiplication, division, and shift operations on the numerical value indicated by the first attribute information, then encodes the first attribute information, and generates a bitstream which includes encoded first attribute information including the encoded first attribute information, and an offset value and a scale value used in a first inverse transformation process to obtain the first attribute information from the encoded first attribute information.
[0016] According to this method, the first attribute information associated with the points of the base mesh can be encoded after offset and scaling processing. As a result, even if the values of the first attribute information include negative values or decimals, they can be stably mapped to a numerical range suitable for encoding. Furthermore, since the offset and scale values used in the first inverse transform are included in the bitstream and transmitted, the first attribute information can be properly decoded.
[0017] An encoding method according to a second aspect of the present disclosure is an encoding method according to a first aspect, wherein the acquisition further acquires second attribute information which is attribute information of a video, performs a second transformation process on the second attribute information and then encodes the second attribute information, and the bitstream includes the encoded first attribute information, encoded second attribute information which includes the encoded second attribute information, first metadata which includes first parameters which include the offset value and the scale value used in the first inverse transformation process, and second metadata which includes second parameters which are used in the second inverse transformation process which corresponds to the second transformation process.
[0018] According to this method, in addition to the first attribute information for the points of the base mesh, the second attribute information, which is the attribute information of the video, can also be encoded after undergoing a second transformation process. As a result, while transmitting the first and second attribute information in the same bitstream, the first parameters based on the offset and scale values used in the first inverse transformation process and the second parameters used in the second inverse transformation process can be distinguished and maintained as first metadata and second metadata, respectively. Therefore, on the decoding side, an appropriate inverse transformation process can be selected for each type of attribute based on the metadata, and the attribute information for the base mesh and the attribute information for the video can be decoded in appropriate formats.
[0019] An encoding method according to a third aspect of the present disclosure is an encoding method according to the first or second aspect, wherein the first conversion process performs the scaling process after the offset process.
[0020] According to this, the procedure is defined so that scaling is performed after offset processing in the first transformation process. As a result, the order of preprocessing is the same on the encoding and decoding sides, and the first attribute information can be transformed using the same procedure and decoded appropriately.
[0021] The encoding method according to the fourth aspect of this disclosure is an encoding method according to any one aspect of the first to third aspects, wherein the offset value is expressed as a fraction with a fixed denominator, and the scale value is expressed as a fraction with a fixed denominator.
[0022] According to this method, both the offset value and the scale value can be expressed as fractions with a fixed denominator. As a result, the representation format of the transformation parameters can be unified on both the encoding and decoding sides, and the computation process can be implemented with a simple configuration.
[0023] The encoding method according to the fifth aspect of this disclosure is the encoding method according to the fourth aspect, wherein the denominator of the fraction representing the offset value is 2 to the power of 16, and the denominator of the fraction representing the scale value is 2 to the power of 16.
[0024] According to this, the 16th power of 2 can be used as the denominator of the fraction representing the offset value and the scale value. As a result, the operation corresponding to the denominator can be easily realized by a shift operation, and the conversion process of the attribute information can be efficiently executed while maintaining the accuracy of 16 bits.
[0025] The encoding method according to the sixth aspect of the present disclosure is an encoding method according to any one of the first aspect to the fifth aspect, wherein the offset processing is executed with decimal precision.
[0026] According to this, the offset processing can be executed with decimal precision. As a result, a fine offset amount can be applied to the first attribute information, and it can be converted into a value suitable for the encoding process while suppressing the rounding error.
[0027] The decoding method according to the seventh aspect of the present disclosure acquires a bit stream including the encoded first attribute information, which is the attribute information associated with a point in the base mesh of the three-dimensional mesh, and the offset value and the scale value used for the inverse transformation of the first attribute information obtained from the encoded first attribute information, decodes the encoded first attribute information from the bit stream to obtain the first attribute information, and acquires the offset value and the scale value, and performs at least one of multiplication, division, and shift operations based on the scale value and offset processing for performing addition and subtraction based on the offset value on the numerical value indicated by the first attribute information, and executes a first inverse transformation process including the above to obtain the first attribute information.
[0028] According to this, after obtaining the first attribute information based on the bit stream including the encoded first attribute information and the offset value and the scale value, the first attribute information can be obtained by the first inverse transformation process. As a result, the inverse transformation process corresponding to the transformation process applied in the encoding method can be faithfully decoded based on the offset value and the scale value in the bit stream. Further, since the conversion parameters are included in the bit stream, even when the statistical characteristics of the attributes are different, the first attribute information can be appropriately decoded using appropriate conversion conditions for each attribute.
[0029] The decoding method according to the eighth aspect of the present disclosure is the decoding method according to the seventh aspect, wherein the bitstream further includes second encoded data in which second attribute information that is video attribute information is encoded, first metadata including a first parameter used for the first inverse conversion process, and second metadata including a second parameter used for a second inverse conversion process for the second attribute information. The first parameter includes the offset value and the scale value. The decoding method decodes the second encoded data from the bitstream to decode the second attribute information, executes the first inverse conversion process based on the first parameter included in the first metadata, and executes a second inverse conversion process for the decoded second attribute information based on the second parameter included in the second metadata.
[0030] According to this, it is possible to include, in the bitstream, second encoded data in which second attribute information is encoded, first metadata including a first parameter used for the first inverse conversion process, and second metadata including a second parameter used for a second inverse conversion process for the second attribute information. Further, by including the offset value and the scale value in the first parameter, it is possible to clearly indicate the conditions for the inverse conversion process for the first attribute information. Therefore, on the decoding side, based on the first metadata and the second metadata, appropriate inverse conversion processes can be selected and applied to the first attribute information and the second attribute information, respectively, and both the attribute information for the base mesh and the attribute information for the video can be efficiently decoded from one bitstream.
[0031] The decoding method according to the ninth aspect of the present disclosure is the decoding method according to the seventh aspect, wherein in the first inverse conversion process, the offset process is performed after the scaling process.
[0032] According to this, the procedure is defined such that the offset process is performed after the scaling process in the first inverse conversion process. As a result, the order of the conversion process and the inverse conversion process corresponds between the encoding side and the decoding side, and the first attribute information can be inversely converted based on the corresponding order and appropriately decoded.
[0033] A decoding method according to a tenth aspect of this disclosure is a decoding method according to a seventh aspect, wherein the offset value is expressed as a fraction with a fixed denominator, and the scale value is expressed as a fraction with a fixed denominator.
[0034] According to this, both the offset value and scale value used in the decoding method can be treated as fractions with a fixed denominator. As a result, arithmetic processing can be performed based on the same representation format as the transformation parameters used in the encoding method, and the first attribute information can be appropriately decoded based on the same transformation parameters.
[0035] The decoding method according to the eleventh aspect of this disclosure is the decoding method according to the seventh aspect, wherein the denominator of the fraction representing the offset value is 2 to the power of 16, and the denominator of the fraction representing the scale value is 2 to the power of 16.
[0036] According to this, 2 to the power of 16 can be used as the denominator of the fractions representing the offset and scale values used in the decoding method. As a result, an efficient inverse transformation process using shift operations can be achieved with a value that has high affinity with binary representation, and the attribute information of the base mesh can be decoded appropriately with high accuracy.
[0037] The decoding method according to the twelfth aspect of this disclosure is the decoding method according to the seventh aspect, wherein the offset processing is performed with decimal precision.
[0038] According to this, the offset processing in the decoding method can be performed with decimal precision. As a result, a fine offset amount can be set for the first decoded attribute information, and a value close to the first attribute information before encoding can be appropriately decoded.
[0039] An encoding device according to a thirteenth aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires first attribute information which is attribute information associated with a point in the base mesh of a three-dimensional mesh, performs a first transformation process which involves performing an offset process that performs addition or subtraction on the numerical value indicated by the first attribute information and a scaling process which performs at least one of multiplication, division and shift operations, then encodes the first attribute information, and generates a bitstream which includes encoded first attribute information which includes the encoded first attribute information and an offset value and a scale value used in a first inverse transformation process to obtain the first attribute information from the encoded first attribute information.
[0040] According to this method, the first attribute information associated with the points of the base mesh can be encoded after offset and scaling processing. As a result, even if the values of the first attribute information include negative values or decimals, they can be stably mapped to a numerical range suitable for encoding. Furthermore, by including the offset and scale values in the bitstream and transmitting them, the decoding side can perform inverse transformation processing based on the same transformation conditions, allowing for proper decoding of the first attribute information.
[0041] A decoding device according to a fourteenth aspect of the present disclosure comprises a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires a bitstream including encoded first attribute information, which is attribute information associated with points in the base mesh of a three-dimensional mesh, encoded in first attribute information, and offset values and scale values used for the inverse transformation of the first attribute information obtained from the encoded first attribute information, decodes the encoded first attribute information from the bitstream to obtain first attribute information, and acquires the offset values and scale values, and performs a first inverse transformation process on the numerical value indicated by the first attribute information, which includes a scaling process that performs at least one of multiplication / division and shift operations based on the scale values, and an offset process that performs addition / subtraction based on the offset values, thereby acquiring the first attribute information.
[0042] According to this method, the first decoded attribute information can be obtained based on a bitstream containing encoded first attribute information in which the first attribute information is encoded, and the offset and scale values used in the encoding. After that, the first attribute information can be obtained by a first inverse transform process. As a result, the transform process applied in the encoding method and the corresponding inverse transform process can be faithfully decoded based on the offset and scale values in the bitstream. Furthermore, since the transform parameters are included in the bitstream, even if the statistical characteristics of the attributes differ, the first attribute information can be appropriately decoded using appropriate transform conditions for each attribute.
[0043] These comprehensive or specific embodiments may be implemented as a system, device, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, device, integrated circuit, computer program, or recording medium.
[0044] The embodiments will be described in detail below with reference to the drawings.
[0045] The embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit the present invention. Furthermore, among the components related to the following embodiments, components not described in the independent claim representing the highest-level concept will be described as optional components.
[0046] (Embodiment 1) In this embodiment, the encoding method and decoding method will be described.
[0047] <Expressions and Terminology> The following expressions and terminology are used here.
[0048] (1) Three-dimensional mesh A three-dimensional mesh is a collection of multiple faces, for example, representing a three-dimensional object. A three-dimensional mesh mainly consists of vertex information, connection information, and attribute information. A three-dimensional mesh may be expressed as a polygon mesh or a mesh. A three-dimensional mesh may also have temporal changes. A three-dimensional mesh may contain metadata related to vertex information, connection information, and attribute information, as well as other additional information.
[0049] (2) Vertex Information Vertex information is information that indicates a vertex. For example, vertex information indicates the position of a vertex in three-dimensional space. Also, a vertex corresponds to the vertices of the faces that make up a three-dimensional mesh. Vertex information is sometimes expressed as "Geometry". Vertex information is also sometimes expressed as position information.
[0050] (3) Connection Information Connection information is information that indicates the connections between vertices. For example, connection information indicates the connections that make up the faces or edges of a three-dimensional mesh. Connection information is sometimes expressed as "Connectivity". Connection information is also sometimes expressed as face information.
[0051] (4) Attribute Information Attribute information is information that indicates the attributes of a vertex or face. For example, attribute information indicates attributes such as color, image, and normal vector associated with a vertex or face. Attribute information is sometimes expressed as "Texture".
[0052] (5) A face is an element that constitutes a three-dimensional mesh. Specifically, a face is a polygon on a plane in three-dimensional space. For example, a face can be defined as a triangle in three-dimensional space.
[0053] (6) A plane is a two-dimensional plane in three-dimensional space. For example, a polygon is formed on a plane, and multiple polygons are formed on multiple planes.
[0054] (7) Bitstream A bitstream corresponds to encoded information. A bitstream can also be expressed as a stream, encoded bitstream, compressed bitstream, or encoded signal.
[0055] (8) The expressions "encode" and "decode" may be replaced with expressions such as "store," "include," "write," "describe," "signalize," "transmit," "notify," "save," or "compress," and these expressions may be interchangeable. For example, encoding information may mean including information in a bitstream. Also, encoding information into a bitstream may mean encoding information and generating a bitstream that contains the encoded information.
[0056] Furthermore, the expression "decode" may be replaced with expressions such as "read out," "decipher," "read," "load," "derive," "obtain," "receive," "extract," "restore," "reconstruct," "decompress," or "expand," and these expressions may be interchangeable. For example, decoding information may mean obtaining information from a bitstream. Also, decoding information from a bitstream may mean decoding the bitstream and obtaining the information contained in the bitstream.
[0057] (9) In the explanation of ordinal numbers, first and second ordinal numbers may be assigned to components, etc. These ordinal numbers may be rearranged as appropriate. Ordinal numbers may also be newly assigned to components, etc., or removed. Furthermore, these ordinal numbers may be assigned to elements in order to identify them, and may not correspond to a meaningful order.
[0058] <Three-Dimensional Mesh> Figure 1 is a conceptual diagram showing a three-dimensional mesh according to this embodiment. The three-dimensional mesh is composed of multiple faces. For example, each face is a triangle. The vertices of these triangles are defined in three-dimensional space. The three-dimensional mesh represents a three-dimensional object. Each face may have a color or image.
[0059] Figure 2 is a conceptual diagram showing the basic elements of a three-dimensional mesh according to this embodiment. The three-dimensional mesh consists of vertex information, connection information, and attribute information. Vertex information indicates the position of the vertices of a face in three-dimensional space. Connection information indicates the connections between vertices. A face can be identified by the vertex information and connection information. In other words, a colorless three-dimensional object is formed in three-dimensional space by the vertex information and connection information.
[0060] Attribute information may be associated with vertices or faces. Attribute information associated with vertices may be expressed as "Attribute Per Point". Attribute information associated with vertices may indicate the attributes of the vertex itself or the attributes of the faces connected to the vertex.
[0061] For example, a color may be associated with a vertex as attribute information. The color associated with a vertex may be the color of the vertex itself, or the color of the face connected to the vertex. The color of a face may be the average of multiple colors associated with multiple vertices of the face. In addition, a normal vector may be associated with a vertex or face as attribute information. Such a normal vector can represent the front and back of a face.
[0062] Furthermore, a two-dimensional image may be associated with a surface as attribute information. The two-dimensional image associated with a surface is also referred to as a texture image or "Attribute Map". Additionally, information indicating the mapping between the surface and the two-dimensional image may be associated with the surface as attribute information. Such information indicating the mapping may be referred to as mapping information, vertex information of the texture image, texture coordinates, or "Attribute UV Coordinate".
[0063] Furthermore, information such as color, images, and moving images used as attribute information may be expressed as "Parametric Space."
[0064] Such attribute information can be used to reflect textures onto three-dimensional objects. In other words, vertex information, connection information, and attribute information allow a three-dimensional object with color to be formed in three-dimensional space.
[0065] In the above, attribute information is associated with vertices or faces, but it may also be associated with edges.
[0066] Figure 3 is a conceptual diagram illustrating the mapping according to this embodiment. For example, a region of a two-dimensional image in a two-dimensional plane can be mapped to a surface of a three-dimensional mesh in three-dimensional space. Specifically, coordinate information of a region in a two-dimensional image is associated with a surface of a three-dimensional mesh. As a result, the image of the region mapped in the two-dimensional image is reflected on the surface of the three-dimensional mesh.
[0067] By using mapping, a two-dimensional image used as attribute information can be separated from a three-dimensional mesh. For example, in encoding a three-dimensional mesh, the two-dimensional image may be encoded using an image encoding scheme or a video encoding scheme.
[0068] <System Configuration> Figure 4 is a block diagram showing an example configuration of the encoding and decoding system according to this embodiment. In Figure 4, the encoding and decoding system comprises an encoding device 100 and a decoding device 200.
[0069] For example, the encoding device 100 acquires a three-dimensional mesh and encodes the three-dimensional mesh into a bitstream. The encoding device 100 then outputs the bitstream to the network 300. For example, the bitstream includes the encoded three-dimensional mesh and control information for decoding the encoded three-dimensional mesh. By encoding the three-dimensional mesh, the information of the three-dimensional mesh is compressed.
[0070] Network 300 transmits the bitstream from the encoding device 100 to the decoding device 200. Network 300 may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination thereof. Network 300 is not necessarily limited to bidirectional communication; it may also be a one-way communication network for terrestrial digital broadcasting or satellite broadcasting, etc.
[0071] Furthermore, the network 300 may be replaced by recording media such as DVD (Digital Versatile Disc) or BD (Blu-Ray Disc®).
[0072] The decoding device 200 acquires the bitstream and decodes the three-dimensional mesh from the bitstream. The decoding of the three-dimensional mesh expands the information of the three-dimensional mesh. For example, the decoding device 200 decodes the three-dimensional mesh according to a decoding method that corresponds to the encoding method used by the encoding device 100 to encode the three-dimensional mesh. That is, the encoding device 100 and the decoding device 200 perform encoding and decoding according to their respective encoding and decoding methods.
[0073] The three-dimensional mesh before encoding can also be referred to as the original three-dimensional mesh. Similarly, the three-dimensional mesh after decoding can be referred to as the reconstructed three-dimensional mesh.
[0074] <Encoding Device> Figure 5 is a block diagram showing an example configuration of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a vertex information encoder 101, a connection information encoder 102, and an attribute information encoder 103.
[0075] The vertex information encoder 101 is an electrical circuit that encodes vertex information. For example, the vertex information encoder 101 encodes vertex information into a bitstream according to a defined format for vertex information.
[0076] The connection information encoder 102 is an electrical circuit that encodes connection information. For example, the connection information encoder 102 encodes connection information into a bitstream according to a specified format for connection information.
[0077] The attribute information encoder 103 is an electrical circuit that encodes attribute information. For example, the attribute information encoder 103 encodes attribute information into a bitstream according to a format defined for the attribute information.
[0078] Variable-length coding or fixed-length coding may be used to encode vertex information, connection information, and attribute information. Variable-length coding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.
[0079] The vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be integrated into a single unit. Alternatively, each of the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 may be further subdivided into multiple components.
[0080] Figure 6 is a block diagram showing another configuration example of the encoding device 100 according to this embodiment. For example, the encoding device 100 includes a preprocessor 104 and a postprocessor 105 in addition to the configuration shown in Figure 5.
[0081] The preprocessor 104 is an electrical circuit that performs processing before encoding vertex information, connection information, and attribute information. For example, the preprocessor 104 may perform transformation processing, separation processing, or multiplexing processing on the three-dimensional mesh before encoding. More specifically, for example, the preprocessor 104 may separate vertex information, connection information, and attribute information from the three-dimensional mesh before encoding.
[0082] The post-processor 105 is an electrical circuit that performs processing after encoding the vertex information, connection information, and attribute information. For example, the post-processor 105 may perform conversion processing, separation processing, or multiplexing processing on the encoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 105 may multiplex the encoded vertex information, connection information, and attribute information into a bitstream. Alternatively, for example, the post-processor 105 may further perform variable-length encoding on the encoded vertex information, connection information, and attribute information.
[0083] <Decoding Device> Figure 7 is a block diagram showing an example configuration of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a vertex information decoder 201, a connection information decoder 202, and an attribute information decoder 203.
[0084] The vertex information decoder 201 is an electrical circuit that decodes vertex information. For example, the vertex information decoder 201 decodes vertex information from a bitstream according to a format defined for vertex information.
[0085] The connection information decoder 202 is an electrical circuit that decodes connection information. For example, the connection information decoder 202 decodes connection information from a bitstream according to a format defined for connection information.
[0086] The attribute information decoder 203 is an electrical circuit that decodes attribute information. For example, the attribute information decoder 203 decodes attribute information from a bitstream according to a format defined for attribute information.
[0087] Variable-length decoding or fixed-length decoding may be used for decoding vertex information, connection information, and attribute information. Variable-length decoding may correspond to Huffman coding or context-adaptive binary arithmetic coding (CABAC), etc.
[0088] The vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be integrated. Alternatively, each of the vertex information decoder 201, the connection information decoder 202, and the attribute information decoder 203 may be subdivided into multiple components.
[0089] Figure 8 is a block diagram showing another configuration example of the decoding device 200 according to this embodiment. For example, the decoding device 200 includes a pre-processor 204 and a post-processor 205 in addition to the configuration shown in Figure 7.
[0090] The preprocessor 204 is an electrical circuit that performs processing before decoding the vertex information, connection information, and attribute information. For example, the preprocessor 204 may perform conversion processing, separation processing, or multiplexing processing on the bitstream before decoding the vertex information, connection information, and attribute information.
[0091] More specifically, for example, the preprocessor 204 may separate the bitstream into sub-bitstreams corresponding to vertex information, connection information, and attribute information. Alternatively, for example, the preprocessor 204 may perform variable-length decoding on the bitstream before decoding the vertex information, connection information, and attribute information.
[0092] The post-processor 205 is an electrical circuit that performs processing after decoding the vertex information, connection information, and attribute information. For example, the post-processor 205 may perform conversion processing, separation processing, or multiplexing processing on the decoded vertex information, connection information, and attribute information. More specifically, for example, the post-processor 205 may multiplex the decoded vertex information, connection information, and attribute information into a three-dimensional mesh.
[0093] <Bitstream> Vertex information, connection information, and attribute information are encoded and stored in the bitstream. The relationship between this information and the bitstream is shown below.
[0094] Figure 9 is a conceptual diagram showing an example of the bitstream configuration according to this embodiment. In this example, connection information, vertex information, and attribute information are integrated in the bitstream. For example, connection information, vertex information, and attribute information may be included in a single file.
[0095] Furthermore, multiple parts of this information may be stored sequentially, such as the first part of connection information, the first part of vertex information, the first part of attribute information, the second part of connection information, the second part of vertex information, the second part of attribute information, and so on. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.
[0096] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.
[0097] Figure 10 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, multiple files are included in the bitstream, and connection information, vertex information, and attribute information are stored in different files. Here, a file containing connection information, a file containing vertex information, and a file containing attribute information are shown, but the storage format is not limited to this example. For example, two types of information from connection information, vertex information, and attribute information may be included in one file, and the remaining type of information may be included in another file.
[0098] Alternatively, this information may be divided and stored in more files. For example, multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files. These multiple parts may correspond to multiple parts that are different in time, multiple parts that are different in space, or multiple faces that are different.
[0099] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the example above, and a different storage order may be used.
[0100] Figure 11 is a conceptual diagram showing another example of the bitstream configuration according to this embodiment. In this example, the bitstream is composed of multiple separable sub-bitstreams, and connection information, vertex information, and attribute information are stored in different sub-bitstreams.
[0101] Here, sub-bitstreams containing connection information, sub-bitstreams containing vertex information, and sub-bitstreams containing attribute information are shown, but the storage format is not limited to these examples.
[0102] For example, two types of information from connection information, vertex information, and attribute information may be included in one sub-bitstream, and the remaining type of information may be included in another sub-bitstream. Specifically, attribute information of a two-dimensional image, etc., may be stored in a sub-bitstream compliant with an image encoding scheme, separate from the sub-bitstreams of connection information and vertex information.
[0103] Furthermore, each sub-bitstream may contain multiple files. Multiple parts of connection information may be stored in multiple files, multiple parts of vertex information may be stored in multiple files, and multiple parts of attribute information may be stored in multiple files.
[0104] Furthermore, the storage order of connection information, vertex information, and attribute information is not limited to the examples in Figures 9, 10, and 11, and a different storage order may be used. For example, they may be stored in the bitstream in the order of vertex information, connection information, and attribute information. Alternatively, they may be stored in the bitstream in any of the following orders: connection information, attribute information, and vertex information; vertex information, attribute information, and connection information; attribute information, connection information, and vertex information; or attribute information, vertex information, and connection information.
[0105] Furthermore, connection information, vertex information, and attribute information may each be divided into multiple data points, and these multiple data points may be stored in a bitstream in a periodic or random order.
[0106] <Specific Example> Figure 12 is a block diagram showing a specific example of the encoding and decoding system according to this embodiment. In Figure 12, the encoding and decoding system comprises a three-dimensional data encoding system 110, a three-dimensional data decoding system 210, and an external connector 310.
[0107] The three-dimensional data encoding system 110 comprises a controller 111, an input / output processor 112, a three-dimensional data encoder 113, a three-dimensional data generator 115, and a system multiplexer 114. The three-dimensional data decoding system 210 comprises a controller 211, an input / output processor 212, a three-dimensional data decoder 213, a system demultiplexer 214, a presenter 215, and a user interface 216.
[0108] In the three-dimensional data encoding system 110, sensor data is input from the sensor terminal to the three-dimensional data generator 115. The three-dimensional data generator 115 generates three-dimensional data, such as point cloud data or mesh data, from the sensor data and inputs it to the three-dimensional data encoder 113.
[0109] For example, the 3D data generator 115 generates vertex information and connection information and attribute information corresponding to the vertex information. The 3D data generator 115 may process the vertex information when generating the connection information and attribute information. For example, the 3D data generator 115 may reduce the amount of data by deleting duplicate vertices or perform transformations on the vertex information (such as position shifting, rotation, or normalization). The 3D data generator 115 may also render the attribute information.
[0110] Furthermore, although the three-dimensional data generator 115 is a component of the three-dimensional data encoding system 110 in Figure 12, it may also be located externally and independently of the three-dimensional data encoding system 110.
[0111] The sensor terminal that provides sensor data for generating three-dimensional data may be, for example, a moving object such as an automobile, a flying object such as an airplane, a mobile terminal, or a camera. Alternatively, distance sensors such as LIDAR, millimeter-wave radar, infrared sensors, or rangefinders, stereo cameras, or combinations of multiple monocular cameras may be used as sensor terminals.
[0112] The sensor data may include the distance (position) of the object, monocular camera images, stereo camera images, color, reflectivity, sensor attitude, orientation, gyroscope, sensing position (GPS information or altitude), speed, acceleration, sensing time, temperature, atmospheric pressure, humidity, or magnetism.
[0113] The three-dimensional data encoder 113 corresponds to the encoding device 100 shown in Figure 5, etc. For example, the three-dimensional data encoder 113 encodes three-dimensional data and generates encoded data. The three-dimensional data encoder 113 also generates control information during the encoding of three-dimensional data. The three-dimensional data encoder 113 then inputs the encoded data along with the control information to the system multiplexer 114.
[0114] The encoding method for three-dimensional data may be a geometry-based encoding method or a video codec-based encoding method. Here, the geometry-based encoding method can also be referred to as a geometry-based encoding method. The video codec-based encoding method can also be referred to as a video-based encoding method.
[0115] The system multiplexer 114 multiplexes the encoded data and control information input from the three-dimensional data encoder 113 and generates multiplexed data using a predetermined multiplexing scheme. The system multiplexer 114 may also multiplex other media such as video, audio, subtitles, application data, or document files, or reference time information, along with the encoded data and control information of the three-dimensional data. Furthermore, the system multiplexer 114 may also multiplex sensor data or attribute information related to the three-dimensional data.
[0116] For example, the multiplexed data may have a file format for storage or a packet format for transmission. ISOBMFF or an ISOBMFF-based format may be used as these formats. Alternatively, MPEG-DASH, MMT, MPEG-2 TS Systems, or RTP may be used.
[0117] The multiplexed data is then output as a transmission signal to the external connector 310 by the input / output processor 112. The multiplexed data may be transmitted as a transmission signal by wire or by wireless. Alternatively, the multiplexed data may be stored in internal memory or storage device. The multiplexed data may be transmitted to a cloud server via the internet or stored in an external storage device.
[0118] For example, the transmission or storage of multiplexed data is carried out in a manner appropriate to the medium for transmission or storage, such as broadcasting or telecommunications. Communication protocols such as HTTP, FTP, TCP, UDP, IP, or combinations thereof may be used. Furthermore, either a pull-type or push-type communication method may be used.
[0119] For wired transmission, Ethernet®, USB, RS-232C, HDMI®, or coaxial cable may be used. For wireless transmission, 3GPP®, IEEE 3G / 4G / 5G, wireless LAN, Wi-Fi, Bluetooth, or millimeter wave may be used. Furthermore, as a broadcasting method, for example, DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.
[0120] The sensor data may also be input to the three-dimensional data generator 115 or the system multiplexer 114. Alternatively, the three-dimensional data or encoded data may be output directly as a transmission signal to the external connector 310 via the input / output processor 112. The transmission signal output from the three-dimensional data encoding system 110 is input to the three-dimensional data decoding system 210 via the external connector 310.
[0121] Furthermore, each operation of the three-dimensional data encoding system 110 may be controlled by a controller 111 that executes an application program.
[0122] In the three-dimensional data decoding system 210, the transmission signal is input to the input / output processor 212. The input / output processor 212 decodes the transmission signal into multiplexed data in file format or packet format and inputs the multiplexed data to the system demultiplexer 214. The system demultiplexer 214 obtains encoded data and control information from the multiplexed data and inputs them to the three-dimensional data decoder 213. The system demultiplexer 214 may also extract other media or reference time information from the multiplexed data.
[0123] The three-dimensional data decoder 213 corresponds to the decoding device 200 shown in Figure 7, etc. For example, the three-dimensional data decoder 213 decodes three-dimensional data from encoded data based on a predetermined encoding scheme. The three-dimensional data is then presented to the user by the presenter 215.
[0124] In addition, additional information such as sensor data may be input to the display device 215. The display device 215 may present three-dimensional data based on the additional information. Furthermore, user instructions may be input from the user terminal to the user interface 216. The display device 215 may then present three-dimensional data based on the input instructions.
[0125] The input / output processor 212 may also acquire three-dimensional data and encoded data from the external connector 310.
[0126] Furthermore, each operation of the three-dimensional data decoding system 210 may be controlled by a controller 211 that executes an application program.
[0127] Figure 13 is a conceptual diagram showing an example of the configuration of point cloud data according to this embodiment. Point cloud data is data representing a three-dimensional object.
[0128] Specifically, a point cloud consists of multiple points and contains positional information indicating the three-dimensional coordinate position of each point, as well as attribute information indicating the attributes of each point. Positional information is also expressed as geometry.
[0129] The type of attribute information may be, for example, color or reflectance. A single point may be associated with attribute information relating to one type, a single point may be associated with attribute information relating to multiple different types, or a single point may be associated with attribute information having multiple values for the same type.
[0130] Figure 14 is a conceptual diagram showing an example of a data file for point cloud data according to this embodiment. In this example, there is a one-to-one correspondence between location information items and attribute information items, and it shows the location information and attribute information of N points that constitute the point cloud data. In this example, the location information is information that indicates the three-dimensional coordinate position on the three axes x, y, and z, and the attribute information is information that indicates the color in RGB. A PLY file or the like can be used as a typical data file for point cloud data.
[0131] Figure 15 is a conceptual diagram showing an example of the configuration of mesh data according to this embodiment. Mesh data is data used in CG (Computer Graphics), etc., and is three-dimensional mesh data that shows the three-dimensional shape of an object with multiple faces. Each face is also represented as a polygon and has the shape of a polygon such as a triangle or quadrilateral.
[0132] Specifically, a three-dimensional mesh consists of multiple points that make up a point cloud, as well as multiple edges and multiple faces. Each point can also be expressed as a vertex or position. Each edge corresponds to a line segment connected by two vertices. Each face corresponds to a region enclosed by three or more edges.
[0133] Furthermore, a three-dimensional mesh contains positional information indicating the three-dimensional coordinate positions of its vertices. This positional information is also referred to as vertex information or geometry. A three-dimensional mesh also contains connection information indicating the relationships between multiple vertices that constitute an edge or face. This connection information is also referred to as connectivity. Finally, a three-dimensional mesh contains attribute information indicating the attributes of vertices, edges, or faces. This attribute information in a three-dimensional mesh is also referred to as texture.
[0134] For example, attribute information may indicate the color, reflectance, or normal vector for a vertex, edge, or face. The direction of the normal vector can represent the front and back of the face.
[0135] Object files or similar formats may be used as the data file format for mesh data.
[0136] Figure 16 is a conceptual diagram showing an example of a data file for mesh data according to this embodiment. In this example, the data file includes position information G(1) to G(N) for N vertices constituting the three-dimensional mesh, and attribute information A1(1) to A1(N) for the N vertices. In addition, this example includes M pieces of attribute information A2(1) to A2(M). The items of the attribute information do not have to correspond one-to-one with vertices, nor do they have to correspond one-to-one with faces. Furthermore, attribute information may not exist at all.
[0137] Connection information is indicated by a combination of vertex indices. n[1, 3, 4] represents a triangular face composed of three vertices n=1, n=3, and n=4. Also, m[2, 4, 6] indicates that attribute information m=2, m=4, and m=6 correspond to the three vertices, respectively.
[0138] Furthermore, the actual content of the attribute information may be described in a separate file. A pointer to that content may be associated with a vertex or face, etc. For example, attribute information indicating an image for a face may be stored in a two-dimensional attribute map file. The file name of the attribute map and the two-dimensional coordinate values in the attribute map may be described in attribute information A2(1) to A2(M). The method of specifying attribute information for a face is not limited to these methods, and any method may be used.
[0139] Figure 17 is a conceptual diagram showing the types of three-dimensional data according to this embodiment. Point cloud data and mesh data may represent static objects or dynamic objects. Static objects are objects that do not change over time, while dynamic objects are objects that change over time. Static objects may correspond to three-dimensional data for any given point in time.
[0140] For example, point cloud data for any given time point may be referred to as a PCC frame. Similarly, mesh data for any given time point may be referred to as a mesh frame. Furthermore, both PCC frames and mesh frames may simply be referred to as frames.
[0141] Furthermore, the object's area may be limited to a certain range, like in regular video data, or it may not be limited, like in map data. Also, the density of points or surfaces can be determined in various ways. Sparse point cloud data or sparse mesh data may be used, or dense point cloud data or dense mesh data may be used.
[0142] Next, the encoding and decoding of point clouds or three-dimensional meshes will be described. The apparatus, processing, or syntax for encoding and decoding vertex information of a three-dimensional mesh in this disclosure may be applied to the encoding and decoding of point clouds. The apparatus, processing, or syntax for encoding and decoding of point clouds in this disclosure may be applied to the encoding and decoding of vertex information of a three-dimensional mesh.
[0143] Furthermore, the apparatus, processing, or syntax for encoding and decoding point cloud attribute information in this disclosure may also be applied to the encoding and decoding of connection information or attribute information of a three-dimensional mesh.
[0144] Furthermore, at least some of the processing may be shared between the encoding and decoding of point cloud data and the encoding and decoding of mesh data. This can help reduce the size of the circuit and software programs.
[0145] Figure 18 is a block diagram showing an example configuration of a three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 comprises a vertex information encoder 121, an attribute information encoder 122, a metadata encoder 123, and a multiplexer 124. The vertex information encoder 121, the attribute information encoder 122, and the multiplexer 124 may correspond to the vertex information encoder 101, the attribute information encoder 103, and the post-processor 105 in Figure 6, etc.
[0146] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a geometry-based encoding scheme. In encoding according to a geometry-based encoding scheme, the three-dimensional structure is taken into consideration. In addition, in encoding according to a geometry-based encoding scheme, attribute information is encoded using the configuration information obtained in encoding the vertex information.
[0147] Specifically, first, vertex information, attribute information, and metadata contained in the three-dimensional data generated from sensor data are input to the vertex information encoder 121, attribute information encoder 122, and metadata encoder 123, respectively. Here, connection information contained in the three-dimensional data may be treated in the same way as attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0148] The vertex information encoder 121 encodes the vertex information into compressed vertex information and outputs the compressed vertex information as encoded data to the multiplexer 124. The vertex information encoder 121 also generates metadata for the compressed vertex information and outputs it to the multiplexer 124. Furthermore, the vertex information encoder 121 generates configuration information and outputs it to the attribute information encoder 122.
[0149] The attribute information encoder 122 uses the configuration information generated by the vertex information encoder 121 to encode the attribute information into compressed attribute information and outputs the compressed attribute information as encoded data to the multiplexer 124. The attribute information encoder 122 also generates metadata for the compressed attribute information and outputs it to the multiplexer 124.
[0150] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The metadata encoded by the metadata encoder 123 may be used for encoding vertex information and attribute information.
[0151] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.
[0152] Figure 19 is a block diagram showing an example configuration of a three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information decoder 221, an attribute information decoder 222, a metadata decoder 223, and a demultiplexer 224. The vertex information decoder 221, attribute information decoder 222, and demultiplexer 224 may correspond to the vertex information decoder 201, attribute information decoder 203, and preprocessor 204 in Figure 8, etc.
[0153] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a geometry-based coding scheme. In decoding according to a geometry-based coding scheme, the three-dimensional structure is taken into consideration. In addition, in decoding according to a geometry-based coding scheme, attribute information is decoded using the configuration information obtained in the decoding of vertex information.
[0154] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information and the metadata of the compressed vertex information are input to the vertex information decoder 221. The compressed attribute information and the metadata of the compressed attribute information are input to the attribute information decoder 222. The metadata is input to the metadata decoder 223.
[0155] The vertex information decoder 221 decodes vertex information from compressed vertex information using the metadata of the compressed vertex information. The vertex information decoder 221 also generates configuration information and outputs it to the attribute information decoder 222. The attribute information decoder 222 decodes attribute information from compressed attribute information using the configuration information generated by the vertex information decoder 221 and the metadata of the compressed attribute information. The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 may be used for decoding vertex information and attribute information.
[0156] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.
[0157] Figure 20 is a block diagram showing another configuration example of the three-dimensional data encoder 113 according to this embodiment. In this example, the three-dimensional data encoder 113 includes a vertex image generator 131, an attribute image generator 132, a metadata generator 133, a video encoder 134, a metadata encoder 123, and a multiplexer 124. The vertex image generator 131, the attribute image generator 132, and the video encoder 134 may correspond to the vertex information encoder 101 and the attribute information encoder 103 in Figure 6, etc.
[0158] In this example, the three-dimensional data encoder 113 encodes the three-dimensional data according to a video-based encoding scheme. In encoding according to a video-based encoding scheme, multiple two-dimensional images are generated from the three-dimensional data, and these multiple two-dimensional images are encoded according to a video encoding scheme. Here, the video encoding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.
[0159] Specifically, first, vertex information and attribute information contained in the three-dimensional data generated from sensor data are input to the metadata generator 133. The vertex information and attribute information are then input to the vertex image generator 131 and the attribute image generator 132, respectively. Furthermore, metadata contained in the three-dimensional data is input to the metadata encoder 123. Here, connection information contained in the three-dimensional data may be treated similarly to attribute information. In the case of point cloud data, position information may be treated as vertex information.
[0160] The metadata generator 133 generates map information for multiple two-dimensional images from vertex information and attribute information. The metadata generator 133 then inputs the map information to the vertex image generator 131, the attribute image generator 132, and the metadata encoder 123.
[0161] The vertex image generator 131 generates a vertex image based on the vertex information and map information and inputs it to the video encoder 134. The attribute image generator 132 generates an attribute image based on the attribute information and map information and inputs it to the video encoder 134.
[0162] The video encoder 134 encodes the vertex image and attribute image into compressed vertex information and compressed attribute information, respectively, according to the video encoding scheme, and outputs the compressed vertex information and compressed attribute information as encoded data to the multiplexer 124. The video encoder 134 also generates metadata for the compressed vertex information and metadata for the compressed attribute information and outputs them to the multiplexer 124.
[0163] The metadata encoder 123 encodes compressible metadata into compressed metadata and outputs the compressed metadata as encoded data to the multiplexer 124. The compressible metadata includes map information. The metadata encoded by the metadata encoder 123 may also be used for encoding vertex information and attribute information.
[0164] The multiplexer 124 multiplexes the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata into a bitstream. The multiplexer 124 then inputs the bitstream to the system layer.
[0165] Figure 21 is a block diagram showing another configuration example of the three-dimensional data decoder 213 according to this embodiment. In this example, the three-dimensional data decoder 213 includes a vertex information generator 231, an attribute information generator 232, a video decoder 234, a metadata decoder 223, and a demultiplexer 224. The vertex information generator 231, the attribute information generator 232, and the video decoder 234 may correspond to the vertex information decoder 201 and the attribute information decoder 203 in Figure 8, etc.
[0166] In this example, the three-dimensional data decoder 213 decodes the three-dimensional data according to a video-based coding scheme. In decoding according to a video-based coding scheme, multiple two-dimensional images are decoded according to the video coding scheme, and three-dimensional data is generated from the multiple two-dimensional images. Here, the video coding scheme may be HEVC (High Efficiency Video Coding) or VVC (Versatile Video Coding), etc.
[0167] Specifically, first, the bitstream is input from the system layer to the demultiplexer 224. The demultiplexer 224 separates the compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, the metadata of the compressed attribute information, and the compressed metadata from the bitstream. The compressed vertex information, the metadata of the compressed vertex information, the compressed attribute information, and the metadata of the compressed attribute information are input to the video decoder 234. The compressed metadata is input to the metadata decoder 223.
[0168] The video decoder 234 decodes the vertex image according to the video encoding scheme. In doing so, the video decoder 234 decodes the vertex image from the compressed vertex information using the metadata of the compressed vertex information. The video decoder 234 then inputs the vertex image to the vertex information generator 231. The video decoder 234 also decodes the attribute image according to the video encoding scheme. In doing so, the video decoder 234 decodes the attribute image from the compressed attribute information using the metadata of the compressed attribute information. The video decoder 234 then inputs the attribute image to the attribute information generator 232.
[0169] The metadata decoder 223 decodes metadata from the compressed metadata. The metadata decoded by the metadata decoder 223 includes map information used for generating vertex information and attribute information. The metadata decoded by the metadata decoder 223 may also be used for decoding vertex images and attribute images.
[0170] The vertex information generator 231 reconstructs vertex information from the vertex image according to the map information contained in the metadata decoded by the metadata decoder 223. The attribute information generator 232 reconstructs attribute information from the attribute image according to the map information contained in the metadata decoded by the metadata decoder 223.
[0171] Subsequently, vertex information, attribute information, and metadata are output as three-dimensional data from the three-dimensional data decoder 213. For example, this metadata is metadata for vertex information and attribute information, and can be used in application programs.
[0172] Figure 22 is a conceptual diagram showing a specific example of the encoding process according to this embodiment. Figure 22 shows a three-dimensional data encoder 113 and a description encoder 148. In this example, the three-dimensional data encoder 113 comprises a two-dimensional data encoder 141 and a mesh data encoder 142. The two-dimensional data encoder 141 comprises a texture encoder 143. The mesh data encoder 142 comprises a vertex information encoder 144 and a connectivity information encoder 145.
[0173] The vertex information encoder 144, the connection information encoder 145, and the texture encoder 143 may correspond to the vertex information encoder 101, the connection information encoder 102, and the attribute information encoder 103 in Figure 6, etc.
[0174] For example, the two-dimensional data encoder 141 operates as a texture encoder 143 and generates a texture file by encoding the texture corresponding to the attribute information as two-dimensional data according to an image encoding scheme or video encoding scheme.
[0175] Furthermore, the mesh data encoder 142 operates as a vertex information encoder 144 and a connection information encoder 145, and generates a mesh file by encoding vertex information and connection information. The mesh data encoder 142 may also encode mapping information for textures. The encoded mapping information may then be included in the mesh file.
[0176] Furthermore, the description encoder 148 generates a description file by encoding the description corresponding to metadata such as text data. The description encoder 148 may encode the description at the system layer. For example, the description encoder 148 may be included in the system multiplexer 114 in Figure 12.
[0177] The above operation generates a bitstream containing texture files, mesh files, and description files. These files may be multiplexed into the bitstream in a file format such as glTF (Graphics Language Transmission Format) or USD (Universal Scene Description).
[0178] The three-dimensional data encoder 113 may also include two mesh data encoders, namely the mesh data encoder 142. For example, one mesh data encoder encodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data encoder encodes the vertex and connection information of a dynamic three-dimensional mesh.
[0179] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.
[0180] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.
[0181] Figure 23 is a conceptual diagram showing a specific example of the decoding process according to this embodiment. Figure 23 shows a three-dimensional data decoder 213, a description decoder 248, and an presenter 247. In this example, the three-dimensional data decoder 213 comprises a two-dimensional data decoder 241, a mesh data decoder 242, and a mesh reconstructor 246. The two-dimensional data decoder 241 comprises a texture decoder 243. The mesh data decoder 242 comprises a vertex information decoder 244 and a connection information decoder 245.
[0182] The vertex information decoder 244, connection information decoder 245, texture decoder 243, and mesh reconstructor 246 may correspond to the vertex information decoder 201, connection information decoder 202, attribute information decoder 203, and post-processor 205, etc., shown in Figure 8. The presenter 247 may correspond to the presenter 215, etc., shown in Figure 12.
[0183] For example, the two-dimensional data decoder 241 operates as a texture decoder 243 and decodes the texture corresponding to the attribute information from the texture file as two-dimensional data according to the image encoding scheme or video encoding scheme.
[0184] Furthermore, the mesh data decoder 242 operates as a vertex information decoder 244 and a connection information decoder 245, decoding vertex information and connection information from the mesh file. The mesh data decoder 242 may also decode mapping information to textures from the mesh file.
[0185] Furthermore, the description decoder 248 decodes the description corresponding to metadata such as text data from the description file. The description decoder 248 may decode the description at the system layer. For example, the description decoder 248 may be included in the system demultiplexer 214 in Figure 12.
[0186] The mesh reconstructor 246 reconstructs a three-dimensional mesh from vertex information, connection information, and textures according to the description. The presenter 247 renders and outputs the three-dimensional mesh according to the description.
[0187] The above process reconstructs and outputs a 3D mesh from a bitstream containing texture files, mesh files, and description files.
[0188] The three-dimensional data decoder 213 may also include two mesh data decoders, which are mesh data decoders 242. For example, one mesh data decoder decodes the vertex and connection information of a static three-dimensional mesh, while the other mesh data decoder decodes the vertex and connection information of a dynamic three-dimensional mesh.
[0189] Correspondingly, two mesh files may be included in the bitstream. For example, one mesh file may correspond to a static 3D mesh, and the other mesh file may correspond to a dynamic 3D mesh.
[0190] Furthermore, a static three-dimensional mesh may be a three-dimensional mesh of an intraframe encoded using intraprediction, and a dynamic three-dimensional mesh may be a three-dimensional mesh of an interframe encoded using interprediction. In addition, as information for the dynamic three-dimensional mesh, the difference information between the vertex information or connection information of the intraframe three-dimensional mesh and the vertex information or connection information of the interframe three-dimensional mesh may be used.
[0191] A coding scheme for dynamic three-dimensional meshes is sometimes called DMC (Dynamic Mesh Coding). Similarly, a video-based coding scheme for dynamic three-dimensional meshes is sometimes called V-DMC (Video-based Dynamic Mesh Coding).
[0192] The encoding method for point clouds is sometimes called PCC (Point Cloud Compression). The video-based encoding method for point clouds is sometimes called V-PCC (Video-based Point Cloud Compression). The geometry-based encoding method for point clouds is sometimes called G-PCC (Geometry-based Point Cloud Compression).
[0193] <Implementation Example> Figure 24 is a block diagram showing an implementation example of the encoding device 100 according to this embodiment. The encoding device 100 includes a circuit 151 and a memory 152. For example, the multiple components of the encoding device 100 shown in Figure 5, etc., are implemented by the circuit 151 and memory 152 shown in Figure 24.
[0194] Circuit 151 is an information processing circuit and is a circuit that can access memory 152. For example, circuit 151 is a dedicated or general-purpose electrical circuit for encoding a three-dimensional mesh. Circuit 151 may also be a processor such as a CPU. Alternatively, circuit 151 may be a collection of multiple electrical circuits.
[0195] Memory 152 is a dedicated or general-purpose memory in which information for the circuit 151 to encode a three-dimensional mesh is stored. Memory 152 may be an electrical circuit, or it may be connected to circuit 151. Memory 152 may also be included in circuit 151. Memory 152 may also be a collection of multiple electrical circuits. Memory 152 may also be a magnetic disk or an optical disk, or it may be described as storage or a recording medium. Memory 152 may also be a non-volatile memory or a volatile memory.
[0196] For example, memory 152 may store a three-dimensional mesh or a bitstream. Alternatively, memory 152 may store a program for circuit 151 to encode the three-dimensional mesh.
[0197] Furthermore, not all of the multiple components shown in Figure 5, etc., are to be implemented in the encoding device 100, nor are all of the multiple processes shown herein to be performed. Some of the multiple components shown in Figure 5, etc., may be included in other devices, and some of the multiple processes shown herein may be executed by other devices. In addition, the multiple components of this disclosure may be implemented in any combination in the encoding device 100, and the multiple processes of this disclosure may be performed in any combination.
[0198] Figure 25 is a block diagram showing an example of the implementation of the decoding device 200 according to this embodiment. The decoding device 200 includes a circuit 251 and a memory 252. For example, the multiple components of the decoding device 200 shown in Figure 7, etc., are implemented by the circuit 251 and memory 252 shown in Figure 25.
[0199] Circuit 251 is an information processing circuit and is a circuit that can access memory 252. For example, circuit 251 is a dedicated or general-purpose electrical circuit for decoding a three-dimensional mesh. Circuit 251 may also be a processor such as a CPU. Alternatively, circuit 251 may be a collection of multiple electrical circuits.
[0200] Memory 252 is a dedicated or general-purpose memory that stores information for circuit 251 to decode the three-dimensional mesh. Memory 252 may be an electrical circuit and may be connected to circuit 251. Memory 252 may also be included in circuit 251. Memory 252 may also be a collection of multiple electrical circuits. Memory 252 may also be a magnetic disk or an optical disk, or may be described as storage or a recording medium. Memory 252 may also be a non-volatile memory or a volatile memory.
[0201] For example, memory 252 may store a three-dimensional mesh or a bitstream. Alternatively, memory 252 may store a program for circuit 251 to decode the three-dimensional mesh.
[0202] Furthermore, the decoding device 200 does not need to implement all of the components shown in Figure 7, etc., nor does it need to perform all of the processes shown herein. Some of the components shown in Figure 7, etc., may be included in other devices, and some of the processes shown herein may be performed by other devices. In addition, the decoding device 200 may implement the components of this disclosure in any combination, and the processes of this disclosure may be performed in any combination.
[0203] The encoding and decoding methods, including the steps performed by each component of the encoding device 100 and decoding device 200 of this disclosure, may be performed by any device or system. For example, part or all of the encoding and decoding methods may be performed by a computer equipped with a processor, memory, and input / output circuits, etc. In this case, the encoding and decoding methods may be performed by the computer executing a program that causes the computer to perform the encoding and decoding methods.
[0204] Furthermore, a non-temporary computer-readable recording medium such as a CD-ROM may contain either a program or a bitstream.
[0205] An example of a program may be a bitstream. For instance, a bitstream containing an encoded three-dimensional mesh may include syntax elements to cause the decoding device 200 to decode the three-dimensional mesh. The bitstream then causes the decoding device 200 to decode the three-dimensional mesh according to the syntax elements contained within the bitstream. Therefore, a bitstream can perform a similar role to a program.
[0206] The above bitstream may be an encoded bitstream containing an encoded three-dimensional mesh, or a multiplexed bitstream containing an encoded three-dimensional mesh and other information.
[0207] Furthermore, each component of the encoding device 100 and the decoding device 200 may be made of dedicated hardware, general-purpose hardware that executes the above-mentioned program, or a combination thereof. The general-purpose hardware may also consist of a memory on which the program is stored, and a general-purpose processor that reads the program from the memory and executes it. Here, the memory may be semiconductor memory or a hard disk, and the general-purpose processor may be a CPU.
[0208] Furthermore, dedicated hardware may consist of memory and a dedicated processor, etc. For example, a dedicated processor may refer to memory for recording data and execute an encoding method and a decoding method.
[0209] Furthermore, each component of the encoding device 100 and the decoding device 200 may be an electrical circuit, as described above. These electrical circuits may form a single electrical circuit as a whole, or they may be separate electrical circuits. These electrical circuits may correspond to dedicated hardware, or they may correspond to general-purpose hardware that executes the above-mentioned program, etc. Furthermore, the encoding device 100 and the decoding device 200 may be implemented as integrated circuits.
[0210] Furthermore, the encoding device 100 may be a transmitting device that transmits a three-dimensional mesh. The decoding device 200 may be a receiving device that receives a three-dimensional mesh.
[0211] <Encoding and Decoding of Displacements> Here, the following terms are used as examples.
[0212] (1) Image An image is a data unit composed of a collection of pixels, and includes a picture or a block smaller than a picture. Images include both moving images and still images.
[0213] (2) A picture is an image processing unit composed of a collection of pixels, and is also called a frame or field.
[0214] (3) A block is a processing unit consisting of a specific number of pixels. The term block is also used in the examples shown below. The shape of a block is not particularly limited. A block may be, for example, a rectangular shape of M x N pixels, or a square shape of M x M pixels. A block may also be triangular, circular, or have other shapes. Examples of blocks are as follows.
[0215] - Slice, tile, or brick; CTU, superblock, or basic partitioning unit; VPDU, hardware processing partitioning unit; CU, processing block unit, prediction block unit (PU), or orthogonal transformation block unit (TU); subblock
[0216] (4) A pixel or sample is the smallest point, or in other words, the smallest unit, of an image. A pixel or sample includes not only pixels at integer positions but also pixels at sub-pixel positions generated from pixels at integer positions.
[0217] (5) Pixel values or sample values Pixel values or sample values are eigenvalues of a pixel. Pixel values or sample values may include luma values, chroma values, or RGB tonal levels, and may also include depth values or binary values of 0 or 1.
[0218] (6) Flags A flag represents one or more bits. A flag is, for example, a parameter or index represented by two or more bits. A flag may also represent a value that is not represented by a binary number, but by a non-binary number.
[0219] (7) A signal is a symbol or encoded information used to transmit information. Signals include discrete digital signals or continuous analog signals.
[0220] (8) Stream or Bitstream A stream or bitstream is a sequence of digital data that represents a flow of digital data. A stream or bitstream may be a single stream or may consist of multiple streams having multiple layers. A stream or bitstream may be transmitted by serial communication using a single transmission path or by packet communication using multiple transmission paths.
[0221] (9) Differences: For scalar quantities, differences can include simple differences (x - y) and difference calculations. Differences can include the absolute value of the difference (|x - y|), the square of the difference (x^2 - y^2), the square root of the difference (√(x - y)), weighted differences (ax - by, where a and b are constants), or offset differences (x - y + a, where a is the offset).
[0222] (10) For scalar quantities, sums can include simple sums (x + y) and addition operations. Sums can include the absolute value of the sum (|x + y|), the sum of squares (x^2 + y^2), the square root of the sum (√(x + y)), a weighted sum (ax + by, where a and b are constants), or an offset sum (x + y + a, where a is the offset).
[0223] (11) "Based on" The expression "based on something" means that other things besides that "something" may also be considered. Also, "based on" can be used when a direct result is obtained, or when a result is obtained after intermediate results.
[0224] (12) "Used" or "Used" The expression "something was used" or "something was used" means that something other than that "something" may also be considered. Also, the expression "used" or "used" can be used when a direct result is obtained, or when a result is obtained after an intermediate result.
[0225] (13) Prohibition "To prohibit" can be rephrased as "not to permit." Also, "not prohibited / prohibited" or "permitted / permitted" does not necessarily imply an obligation.
[0226] (14) "Restriction" or "Limitation" "Restriction" or "Limitation" can be rephrased as "not permitted / not allowed" or "not permitted / permitted." Also, "prohibited / not prohibited" or "not permitted / permitted" does not necessarily mean "obligation." Furthermore, the prohibition, quantitatively or qualitatively, may be partial or entire.
[0227] (15) Chroma The term chroma is an adjective represented by the symbols Cb or Cr, indicating that a sample sequence or a single sample represents one of two color difference signals associated with a primary color. The term chroma is sometimes used instead of the term chrominance.
[0228] (16) Luma The term luma is an adjective represented by the symbol or subscript Y or L, indicating that a sample sequence or single sample represents a monochrome signal relating to a primary color. The term luma is sometimes used as an alternative to the term luminance.
[0229] The coding and decoding system of this embodiment will be described below.
[0230] A typical three-dimensional model (also called a 3D model) digitally represents an object so that the user can explore the model using zoom, pan, and rotate in all three dimensions while rendering it over time. One way to construct such a representation is to build a 3D mesh using triangles. In the above model, the positions of the triangle vertices, the connectivity of the triangle vertices to each other, and their associated attributes (such as normals or UV patches) are stored.
[0231] Storing all this information in an uncompressed format requires a very large amount of memory, and therefore a very large bandwidth for transmission. The triangles that form a mesh often have attributes similar to repeating patterns, especially in temporal and spatial neighborhoods. These repetitions can be used to develop efficient encoding and decoding methods for storage and transmission. One such encoding and decoding method is Video-based Dynamic Mesh Coding (V-DMC).
[0232] Figure 26 is a block diagram showing different configuration examples of the coding and decoding system according to this embodiment. As shown in Figure 26, the coding and decoding system includes an coding device 100 and a decoding device 200.
[0233] The encoding / decoding system accepts a three-dimensional mesh (also called a 3D mesh) as input in the form of three-dimensional coordinates of vertices (vertex information), connectivity (connection information), and associated attributes (attribute information). Note that the 3D mesh can include not only geometry but also texture maps.
[0234] The encoding device 100 captures the input 3D mesh (also called the input 3D mesh or input mesh) in the form of the three-dimensional coordinates of the vertices, connectivity, and associated attributes. The encoding device 100 encodes all the associated information into a stream. The stream may consist of a single bitstream or multiple bitstreams.
[0235] Network 300 transmits the stream generated by the encoding device to the decoding device 200. Network 300 may be the Internet, a WAN (Wide Area Network), a LAN (Local Area Network), or any combination thereof. Furthermore, network 300 is not necessarily limited to a bidirectional communication network, but may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. Alternatively, a recording medium such as a DVD (Digital Versatile Disc) or a BD (Blue-Ray Disc) on which the stream is recorded may be used instead of network 300.
[0236] The stream is transmitted to the decoder 200 via the network 300. The decoder 200 decodes the bitstream and generates a three-dimensional mesh using the three-dimensional coordinates, connectivity, and associated attributes of the decoded vertices. The decoder 200 outputs the generated three-dimensional mesh (also called the output 3D mesh or output mesh).
[0237] Figure 27 shows another example of the encoding device 100 configuration.
[0238] As shown in Figure 27, the encoding device 100 includes a preprocessor 1103 and a compressor 1106.
[0239] The encoding device 100 reads the input mesh 1101 and attribute map 1102 and passes them to the preprocessor 1103. The preprocessor 1103 processes the input mesh and extracts the base mesh 1104 and displacement data 1105. The attribute map 1102, along with the extracted base mesh 1104 and displacement data 1105, is passed to the compressor 1106.
[0240] Furthermore, the compressor 1106 compresses the base mesh 1104, displacement data 1105, and attribute map 1102 to generate a bitstream 1107. The compressor 1106 can transmit additional information to the decoder 200 by further including metadata 1108 in the bitstream 1107.
[0241] Figure 28 shows another example of the configuration of the decoding device 200.
[0242] As shown in Figure 28, the decoding device 200 includes an expander 2102 and a post-processing unit 2106.
[0243] The decoding device 200 reads the bitstream 2101 and passes it to the decompressor 2102. The decompressor 2102 decompresses the base mesh 2103, displacement data 2104, and attribute map 2108 from the bitstream 2101 and passes them to the post-processor 2106. An example of displacement data 2104 is a displacement vector.
[0244] Furthermore, the post-processor 2106 generates the output mesh 2107 by processing the base mesh 2103 according to the displacement data 2104 and attribute map 2108. The post-processor 2106 may also use information from metadata 2105 to generate the output mesh 2107.
[0245] Figure 29 is a block diagram showing yet another configuration example of the encoding device 100 according to this embodiment.
[0246] In this example, the encoding device 100 includes a volumetric capturer 511, a projector 512, a base mesh encoder 513, a displacement encoder 514, an attribute encoder 515, and optionally one or more other type encoders 516.
[0247] The volumetric capture device 511 captures the content and outputs the captured content to the projector 512.
[0248] The projector 512 projects content onto a three-dimensional mesh frame containing vertex geometry coordinates (vertex coordinates indicating the position of vertices), texture coordinates, and connectivity data (connection information). This data is output to the base mesh encoder 513, the displacement encoder 514, the attribute encoder 515, and optionally one or more other type encoders 516. Each encoder compresses the data into a bitstream.
[0249] Figure 30 is a block diagram showing yet another configuration example of the decoding device 200 according to this embodiment.
[0250] In this example, the decoding device 200 comprises a base mesh decoder 613, a displacement decoder 614, an attribute decoder 615, one or more other type decoders 616, and a three-dimensional reconstructor 617.
[0251] The bitstream is sent to the base mesh decoder 613, the displacement decoder 614, the attribute decoder 615, and optionally one or more other type decoders 616. These decoders decode the bitstream to generate decoded data containing vertex geometry coordinates, texture coordinates, and connectivity data. The decoded data is then sent to the 3D reconstructor 617, where the 3D mesh frame is reconstructed.
[0252] The encoding process performed by the encoding device 100 will be described in detail below.
[0253] Figure 31 is a flowchart showing the processing of the encoding device 100. Figure 32 is an explanatory diagram conceptually illustrating the encoding of a mesh frame. The processing of the encoding device 100 will be explained with reference to Figures 31 and 32.
[0254] In step S101, the encoding device 100 reads the input mesh frame, which is a 3D mesh frame, and its attributes. The input mesh frame is the mesh frame input to the encoding device 100. An example of the input mesh frame, a 3D mesh frame, is shown as mesh frame 1301 (see Figure 32).
[0255] In step S102, the encoding device 100 generates a base mesh frame with fewer vertices than the input mesh frame by performing a decimation process on the input mesh frame read in step S101. The base mesh frame generated by decimating the mesh frame 1301 is shown as the base mesh frame 1302 (see Figure 32).
[0256] In step S103, the encoding device 100 calculates displacement information that the decoding device 200 uses to reconstruct the mesh frame. The displacement information corresponds to a displacement vector from the vertices of the base mesh frame generated in step S102 to the vertices of the input mesh frame. One method for calculating the displacement information is to subtract the coordinates of the vertices of the base mesh frame from the coordinates of the vertices of the input mesh frame. The displacement information calculated from the mesh frame 1301 and the base mesh frame 1302 is shown as displacement information 1303 (see Figure 32). The displacement information 1303 is in vector form, or in other words, it is expressed as a displacement vector.
[0257] In step S104, the encoding device 100 encodes the base mesh frame generated in step S102, the displacement information generated in step S103, and the attributes of the input mesh frame into a bitstream (corresponding to a compressed bitstream). An example of a bitstream is shown as bitstream 1304 (see Figure 32).
[0258] Specifically, bitstream 1304 includes a video bitstream containing vertex coordinates and connection information for vertices A, C, E, and F, displacement information, and texture data, as well as a compressed attribute map (see Figure 32). The displacement information includes displacement information for displacing vertices based on vertex coordinates obtained from the subdivided base mesh frame. The compressed attribute map is texture coordinates for applying texture data to the mesh frame reconstructed using the base mesh frame and displacement information.
[0259] The decoding process performed by the decoding device 200 will be described in detail below.
[0260] Figure 33 is a flowchart showing the processing of the decoding device 200. Figure 34 is an explanatory diagram conceptually illustrating the decoding of a 3D mesh. The processing of the decoding device 200 will be explained with reference to Figures 33 and 34.
[0261] In step S201, the decoding device 200 decodes the base mesh frame and attributes from the bitstream (corresponding to the compressed bitstream). An example of the decoded base mesh frame (corresponding to the decoded base mesh frame) is shown as the decoded base mesh frame 2301 (see Figure 34).
[0262] In step S202, the decoding device 200 generates subdivided vertices by performing a subdivision process on the base mesh frame decoded in step S201. An example of a base mesh frame containing subdivided vertices is shown as base mesh frame 2302 (see Figure 34).
[0263] In step S203, the decoding device 200 decodes displacement information from the bitstream (corresponding to a compressed bitstream). An example of the decoded displacement information is shown as displacement information 2303 (see Figure 34). Displacement information 2303 is in vector form, or in other words, it is represented as a displacement vector.
[0264] In step S204, the decoding device 200 reconstructs the shape of the mesh frame by moving the vertices of the base mesh frame, including the sub-divided vertices, to new positions using displacement information, and then restores the mesh frame by applying attribute information. An example of an attribute is a texture. An example of the reconstructed mesh frame is shown as mesh frame 2304 (see Figure 34).
[0265] Figure 35 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0266] Figure 35 shows an example of a typical intra-decryption block diagram.
[0267] The decoding device shown in Figure 35 comprises an inverse multiplexer 1231, a switch 1232, a static mesh decoder 1233, a mesh buffer 1234, a motion decoder 1235, a base mesh reconstructor 1236, an inverse quantizer 1237, a video decoder 1238, an image umpacker 1239, an inverse quantizer 1240, an inverse wavelet converter 1241, a reconstructor 1242, a video decoder 1243, and a color converter 1244.
[0268] The demultiplexer 1231 acquires the compressed bitstream and separates the compressed data for the base mesh, the video containing displacement data (also called the displacement bitstream), and the video containing attribute data (also called the attribute bitstream). The compressed data for the base mesh is passed to the switch 1232. The switch 1232 decides whether to perform intra-decoding or inter-decoding based on the parameters in the bitstream.
[0269] If intra-decoding is selected, the bitstream is passed to a static mesh decoder 1233 that generates a quantized base mesh. The static mesh decoder 1233 is, for example, a decoder that uses an edge breaker algorithm to decode 3D mesh data. The static mesh decoder 1233 generates a quantized base mesh from the bitstream. The quantized base mesh generated by the static mesh decoder 1233 is stored in the mesh buffer 1234 for reference when inter-decoding is selected.
[0270] Switch 1232 passes compressed data about the base mesh to the motion decoder 1235 if inter-decoding is selected. The motion decoder 1235 receives the previously decoded, quantized base mesh and decodes motion data representing the difference in vertex coordinates between the quantized base mesh stored in the mesh buffer 1234 and the current quantized base mesh. The motion data and the quantized base mesh stored in the mesh buffer 1234 are used by the base mesh reconstructor 1236 to reconstruct the current quantized base mesh. The quantized base mesh obtained from inter-decoding or intra-decoding is passed to the inverse quantizer 1237 to obtain the decoded base mesh.
[0271] The video containing displacement data is passed to the video decoder 1238 because the bitstream contains displacement data in an image format having two chroma information and one luma information. The video decoder 1238 decodes the data using the video frame decompression method. In another example, the displacement data is decoded using an arithmetic decoder. This decompressed data is passed to the image umpacker 1239, which extracts wavelet coefficients associated with each vertex from the decompressed image data. The inverse quantizer 1240 inverse quantizes the quantized wavelet coefficients with three components associated with each vertex. The inverse wavelet transformer 1241 inverse transforms the result to obtain the finally decoded displacement data. The decoded displacement data and the decoded base mesh are passed to the reconstructor 1242. The reconstructor 1242 subdivides the edges of the decoded base mesh, displaces the vertices using the decoded displacement data, and obtains the decoded mesh.
[0272] The video containing attribute data is passed to another video decoder 1243 to obtain a decoded attribute bitstream. The decoded attribute bitstream is further processed by a color converter 1244 for color space and color format conversion to obtain a decoded attribute map.
[0273] Figure 36 is a block diagram showing an example of the configuration of a decoding device according to this embodiment.
[0274] Figure 36 shows an example of a reconstructor that obtains a decoded 3D mesh 1256 from a decoded base mesh 1251 and decoded displacement data 1254.
[0275] The decoded base mesh 1251 is passed to the sub-divider 1252.
[0276] The sub-decomposer 1252 subdivides any two connected vertices in the entire 3D mesh by adding a new vertex between them. This process can be repeated several times, including the vertices created in the previous sub-decomposition step, to generate a predetermined number of vertices. Each iteration of sub-decomposition across the entire 3D mesh generates a new level of detail (LoD). The subdivided mesh 1253 and the decoded displacement data 1254 are passed to the displacementr 1255. The displacementr 1255 generates the decoded 3D mesh 1256 by moving each vertex to a new position according to the corresponding displacement data.
[0277] Subpartitioning will be explained below. Subpartitioning is performed by a subpartitioner (specifically, subpartitioner 1206 or subpartitioner 2204).
[0278] Figure 37 is an explanatory diagram showing an example of subdivision.
[0279] The base mesh shown in Figure 37(a) includes vertices A, B, and C, and connectivity information indicating their connectivity.
[0280] Figure 37(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision. In the first subdivision, the subdivision generator generates vertices D, E, and F, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD1 or the first LoD.
[0281] Vertex D of the mesh after the first subdivision is a vertex generated by the subdivision based on vertices A and B. Similarly, vertex E is a vertex generated by the subdivision based on vertices B and C. Vertex F is a vertex generated by the subdivision based on vertices A and C.
[0282] For example, vertex D could be the midpoint of the line segment AB (or side AB) connecting vertices A and B, which were the source of its generation. Similarly, vertex E could be the midpoint of line segment AC, and vertex F could be the midpoint of line segment BC.
[0283] Figure 37(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision. In the second subdivision, the subdivision generator generates vertices G, H, I, J, K, L, M, N, and O, and connection information indicating their connectivity. The mesh generated by the subdivision generator is also called LoD2 or the second LoD.
[0284] Vertex G of the mesh after the second subdivision is a vertex generated by the subdivision based on vertices A and D. Similarly, vertex H is a vertex generated by the subdivision based on vertices A and E. Vertex I is a vertex generated by the subdivision based on vertices B and D. Vertex J is a vertex generated by the subdivision based on vertices D and F. Vertex K is a vertex generated by the subdivision based on vertices E and F. Vertex L is a vertex generated by the subdivision based on vertices C and E. Vertex M is a vertex generated by the subdivision based on vertices B and F. Vertex N is a vertex generated by the subdivision based on vertices C and F. Vertex O is a vertex generated by the subdivision based on vertices D and E.
[0285] For example, vertex G could be the midpoint of the line segment AD (or edge AD) connecting vertices A and D, which were the source of its generation. Similarly, vertex H could be the midpoint of line segment AE. Vertex I could be the midpoint of line segment BD. Vertex J could be the midpoint of line segment DF. Vertex K could be the midpoint of line segment EF. Vertex L could be the midpoint of line segment CE. Vertex M could be the midpoint of line segment BF. Vertex N could be the midpoint of line segment CF. Vertex O could be the midpoint of line segment DE.
[0286] In the following sections, the displacement of the vertices will be explained with reference to Figures 38 and 39. The displacement of the vertices is performed by the reconstructor 2209.
[0287] Figure 38 is an explanatory diagram showing an example of vertex displacement after subdivision. Figure 39 is an explanatory diagram showing an example of vertices in the original mesh.
[0288] The base mesh shown in Figure 38(a) includes vertices A, B, C, and Z, and connectivity information indicating their connectivity.
[0289] Figure 38(b) shows the mesh generated by the first subdivision, in other words, the mesh after the first subdivision (i.e., the first LoD). In the first subdivision, the subdivision generator generates vertices S, T, U, X, or Y and connection information indicating their connectivity. Vertices S, T, U, X, or Y are the same as vertices D, E, and F shown in Figure 37(b).
[0290] Figure 38(c) shows the mesh generated by the second subdivision, in other words, the mesh after the second subdivision (i.e., the second LoD). In the second subdivision, the subdivision tool generates vertices D, E, F, G, and H, and connection information indicating their connectivity. Vertices D, E, F, G, and H are the same as vertices G, H, I, J, K, L, M, N, or O shown in Figure 37(c).
[0291] Figure 38(d) shows the mesh including the vertices after displacement following subdivision. The vertices A, B, C, D, E, F, G, H, S, T, U, X, Y, and Z shown in Figure 38(d) are located at positions that have been displaced using displacement information from the positions of those vertices shown in Figure 38(c).
[0292] The original mesh shown in Figure 39 is an example of the mesh input to the encoding device 100, that is, the mesh before encoding.
[0293] The mesh shown in Figure 38 has a shape similar to the original mesh shown in Figure 39. The displacement information is generated by the displacement vector calculator 1207 of the encoding device 100 as information indicating the displacement from the vertices of the base mesh to the vertices of the original mesh. By reconstructing the mesh using the displacement information generated in this way, a mesh with a shape similar to the original mesh is generated.
[0294] The decoding device 200 can output the mesh shown in Figure 38(d).
[0295] Next, we will explain the division of the mesh into submeshes with reference to Figures 40 and 41.
[0296] A mesh can be divided into multiple smaller parts, and each divided part can be encoded. When dividing a mesh, the vertices of the mesh are divided in such a way that the coordinates and connectivity of the vertices included in each part can be encoded independently.
[0297] Figure 40 is an explanatory diagram showing an example of a mesh. Figure 41 is an explanatory diagram showing an example of dividing a mesh into submeshes.
[0298] The mesh shown in Figure 40 is the original mesh, and is sometimes called the full mesh in contrast to the submesh.
[0299] Figure 41 shows how the full mesh shown in Figure 40 is divided into two submeshes. For vertices A, B, and C of the full mesh (see Figure 40), vertex A is duplicated into vertices A1 and A2, vertex B is duplicated into vertices B1 and B2, and vertex C is duplicated into vertices C1 and C2, thereby creating two submeshes (i.e., the first submesh and the second submesh) from the full mesh. The first submesh and the second submesh are meshes that can be decoded independently.
[0300] In the following sections, the packing of displacement information into image frames will be explained with reference to Figures 42, 43, and 44.
[0301] Figures 42, 43, and 44 are explanatory diagrams illustrating examples of packing displacement information into image frames. Note that image frames can also be referred to as video frames.
[0302] Vertex displacement data is encoded as image frame data by mapping it to each component of an image frame in YUV format (i.e., the Y component (Y Plane), U component (U Plane), and V component (V Plane) respectively). This case is explained below as an example. Alternatively, vertex displacement data may be encoded as image frame data by mapping it to each component of an image frame in RGB format (the R component, G component, and B component, respectively).
[0303] The decoding device 200 can use an image coding module to extract displacement data. The displacement data may be in the form of X, Y, or Z components in a global coordinate system (e.g., a Cartesian coordinate system), or normal, tangent, or both tangent components in a local coordinate system. Methods for mapping displacement data to an image frame include the following:
[0304] For example, in the first method, displacement data is arranged in scan order within the image frame. An example of packing displacement data in this case is shown in Figure 42. The displacement data is directly mapped to the image frame according to a predefined scan order.
[0305] Note that since the height and width of the image frame are fixed, the displacement data may not fit perfectly within the frame. In such cases, the remaining portion of the image frame is padded with padding data (also called padded data) (see Figure 42).
[0306] For example, in the second method, displacement data is separated into multiple Lines of Data (LDs) and mapped to the Y, U, and V components of the image frame. An example of packing the displacement data in this case is shown in Figure 43. Here, the displacement data for the image frame of the next LD starts immediately after the displacement data for the previous LD ends. Similar to the first method, if the displacement data does not fit perfectly into the image frame, padding is applied to the end of the image frame (see Figure 43).
[0307] For example, in the third method, the displacement data corresponding to the LoD is mapped to the Y, U, and V components of the image frame in a different manner than in the second method. An example of the packing of the displacement data in this case is shown in Figure 44. In this way, each LoD can be decoded independently. In the third method, intermediate padding is performed on the displacement data of each LoD, and CTU alignment is performed together with padding at the end of the video frame (see Figure 44).
[0308] <Other Examples> Although the embodiments of the encoding and decoding devices have been described above according to the embodiments, the embodiments of the encoding and decoding devices are not limited to the embodiments. Modifications that a person skilled in the art can conceive of may be made to the embodiments, and multiple components according to the embodiments may be arbitrarily combined.
[0309] For example, in the embodiment, a process performed by a specific component may be performed by another component instead of that specific component. Also, the order of multiple processes may be changed, or multiple processes may be executed in parallel.
[0310] Furthermore, as described above, at least some of the configurations of this disclosure can be implemented as an integrated circuit. At least some of the processes of this disclosure may be used as an encoding method or a decoding method. A program for causing a computer to execute such encoding method or decoding method may be used. A non-temporary computer-readable recording medium on which such program is recorded may also be used. A bitstream for causing a decoding device to perform the decoding process may also be used.
[0311] Furthermore, at least some of the configurations and processes of this disclosure may be used as transmitting devices, receiving devices, transmitting methods, and receiving methods. A program for causing a computer to execute such transmitting method or receiving method may be used. A non-temporary computer-readable recording medium on which such program is recorded may also be used.
[0312] In the above embodiment, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Here, the software that implements the encoding device, etc. in the above embodiment is the following program.
[0313] In other words, the above software is a program that causes a computer to acquire a first frame to be encoded, encode first data having one or more layers contained in the first frame by referring to second data having one or more layers contained in a second frame, and when encoding the first data, for each of the one or more layers contained in the first data, determine one of several processes using a value indicating that layer and a value relating to the one or more layers contained in the second data, and execute the determined one process on the first data of that layer.
[0314] Furthermore, the software that implements the decoding device and the like in the above embodiment is the following program.
[0315] In other words, the above software is a program that causes a computer to execute a decoding method in which it obtains a first frame to be decoded, decodes first data having one or more layers contained in the first frame by referring to second data having one or more layers contained in the second frame, and when decoding the first data, determines one of several processes for each of the one or more layers contained in the first data using a value indicating that layer and a value relating to the one or more layers contained in the second data, and executes the determined one process on the first data of that layer.
[0316] Next, an example of attribute information in mesh data will be described. Here, "device" refers to an encoding device or a decoding device. Figures 45 to 47 are diagrams illustrating an example of attribute information in mesh data according to Embodiment 1.
[0317] As shown in Figure 45, multiple three-dimensional points G (vertex) in mesh data are projected onto a plane X, and the resulting two-dimensional coordinates on plane X are sometimes set as attribute information for these three-dimensional points. For example, if the two-dimensional coordinates after projecting a three-dimensional point (x, y, z) onto plane X are (u, v), then the two-dimensional coordinates (u, v) are set as attribute information for the three-dimensional point (x, y, z). Figure 45 shows an example in which a set of vertex coordinates 1501 of a three-dimensional mesh represented by mesh data is projected onto plane X, transforming it into a set of two-dimensional coordinates 1502 on plane X.
[0318] Furthermore, as shown in Figure 46, multiple three-dimensional points within the mesh data are projected onto plane X, and the plane information to which the two-dimensional coordinates on plane X after the projection belong may be set as attribute information of the three-dimensional point. For example, if the two-dimensional coordinates (u, v) after projecting a three-dimensional point (x, y, z) onto plane X belong to the plane information face0, then the plane information face0 is set as attribute information of the three-dimensional point (x, y, z). In this case, by setting (u, v) = (face0, 0), the three-dimensional point (x, y, z) and the plane information face0 are associated as uv coordinates. This method allows the relationship between the three-dimensional point (x, y, z) and the plane information face0 to be defined using the same mechanism as uv coordinates.
[0319] Furthermore, when projecting connected three-dimensional points onto a two-dimensional image, attribute information for those three-dimensional points can be added by setting the same plane information. For example, this can be done as follows:
[0320] In the case of Figure 46, `face0` is set as attribute information for 3D point G and 3D points connected to 3D point G. Similarly, `face1` is set as attribute information for 3D points G' and 3D points connected to 3D point G'. By this method, the decoding device can obtain the set planar information by decoding the attribute information of the 3D points. Then, using that planar information, it can obtain, for example, related projection information.
[0321] Furthermore, as shown in Figure 47, the normal vector corresponding to a three-dimensional point in the mesh data, or to the plane to which that three-dimensional point belongs, may be set as attribute information. For example, the device can set the normal vector (nx, ny, nz) of a three-dimensional point (x, y, z) as attribute information. It is also possible to set the normal vector (nx, ny, nz) corresponding to a three-dimensional point G or to the plane to which the three-dimensional point G belongs as attribute information for G.
[0322] The attribute information in mesh data is not limited to this. Attribute information can include any information, such as the color information, reflectance, or group ID of a three-dimensional point. Furthermore, attribute information may include information with various dimensions, such as two-dimensional information like UV coordinates, one-dimensional information like plane information face0, and three-dimensional information like normal vectors.
[0323] Furthermore, it is possible to assign multiple attribute information to a single three-dimensional point. For example, the device can assign multiple attribute information, such as UV coordinates and normal vectors, to a three-dimensional point (x, y, z). It is also possible for three-dimensional points to exist without any attribute information assigned. By assigning various types of attribute information to three-dimensional points in this way, it is possible to generate mesh data with greater expressive power.
[0324] This embodiment provides an example of a method for encoding or decoding attribute information using a device. The example described here describes the case of encoding or decoding one attribute piece of information assigned to a three-dimensional point, but is not limited to this. For example, if two or more attribute pieces of information are assigned to a three-dimensional point, it is also possible to encode or decode each attribute piece of information individually using this embodiment. This makes it possible to efficiently encode or decode attribute pieces of information even for three-dimensional points that have multiple attribute pieces of information assigned to them.
[0325] In this embodiment, an example using UV coordinates as attribute information is shown, but it can also be applied to other types of attribute information. Figures 48 to 53 are diagrams illustrating specific examples of attribute information encoding methods.
[0326] First, as shown in Figure 48, the device selects a first initial three-dimensional point v0 and encodes its uv coordinates without prediction. For example, in the following case, the device may encode the values of (u0, v0) directly. Note that the order in which the three-dimensional points for encoding attribute information are selected is arbitrary; for example, the device may select them in the same order as when encoding the position information of the three-dimensional points. This allows the device to apply predictive encoding while reducing the amount of processing required to select three-dimensional points.
[0327] Next, as shown in Figure 49, the device selects a three-dimensional point v1. If the three-dimensional point v1 has connectivity to an encoded three-dimensional point v, i.e., an edge, the device can use the uv coordinates of three-dimensional point v0 as a predicted value and differentially encode the uv coordinates of three-dimensional point v1. For example, in the following case, the device can encode (u1 - u0, v1 - v0). This reduces the amount of coding required when the uv coordinate values of three-dimensional points v0 and v1 are close. This method of predictive coding using the uv coordinates of a single connected three-dimensional point is called "coarse prediction".
[0328] Next, as shown in Figure 50, the device selects a three-dimensional point v2. If the three-dimensional point v2 is connected to the encoded three-dimensional points v0 and v1, that is, if there are edges between three-dimensional point v2 and three-dimensional point v1, and between three-dimensional point v2 and three-dimensional point v0, the device can generate a predicted value using the uv coordinates of three-dimensional points v0 and v1, and differentially encode the uv coordinate of three-dimensional point v2. For example, in the following case, the device can encode (u2 - predx, v2 - predy). This reduces the amount of coding by generating a predicted value with high accuracy using three-dimensional points v0 and v1. This method of predictive coding using the uv coordinates of two or more connected three-dimensional points is called "fine prediction".
[0329] Next, as shown in Figure 51, the device selects a three-dimensional point v3. If the three-dimensional point v3 has connectivity with the encoded three-dimensional points v1 and v2, that is, if there are edges between the three-dimensional point v3 and three-dimensional point v1, and between the three-dimensional point v3 and three-dimensional point v2, the device can generate predicted values using fine prediction with the uv coordinates of three-dimensional points v1 and v2, and differentially encode the uv coordinates of three-dimensional point v3.
[0330] Next, as shown in Figure 52, the device selects a three-dimensional point v4. If the three-dimensional point v4 has connectivity with the encoded three-dimensional points v0, v1, and v3, that is, if there are edges between three-dimensional point v4 and three-dimensional point v3, three-dimensional point v4 and three-dimensional point v1, and three-dimensional point v4 and three-dimensional point v3, the device can generate a predicted value using fine prediction with the uv coordinates of three-dimensional points v0, v1, and v3, and differentially encode the uv coordinate of three-dimensional point v4. In this case, for example, the device may calculate fine prediction value 1 using the uv coordinates of three-dimensional points v0 and v1, calculate fine prediction value 2 using the uv coordinates of three-dimensional points v1 and v3, and use the average of fine prediction value 1 and fine prediction value 2 as the predicted value of the uv coordinate of three-dimensional point v4. When two or more fine prediction values can be calculated in this way, the prediction accuracy can be improved and the amount of coding can be reduced by using the average value as the predicted value.
[0331] Next, as shown in Figure 53, the device selects a three-dimensional point v5. If the three-dimensional point v5 does not have connectivity with any of the previously encoded three-dimensional points, the device selects v5 as the second initial three-dimensional point and can encode the (u5, v5) value directly without prediction, similar to the first process. Similarly thereafter, the device can reduce the code size of the uv coordinates by applying predictive coding using each process.
[0332] Next, we will describe an example of a prediction mode for encoding attribute information.
[0333] First, let's explain the no-prediction mode. The no-prediction mode is an example of the first prediction mode. The no-prediction mode is applied when the three-dimensional point to be encoded is an initial three-dimensional point. Here, an initial three-dimensional point refers to a three-dimensional point to be processed that has already been encoded and for which there are no other encoded three-dimensional points that are connected to that three-dimensional point.
[0334] The device can encode attribute information values directly in no-prediction mode. For example, if the uv coordinates are (100, 60), the device can encode the value of u (100) and the value of v (60) respectively. The device may also assume fixed-length encoding with n bits (where n is a non-negative integer). For example, if the uv coordinates are (100, 60), the device can encode each component with 7 bits.
[0335] If the device applies n-bit fixed-length coding, it may add the value n to the bitstream header. This allows the decoding device to correctly decode the attribute information of the initial three-dimensional points, which has been coded with n-bit fixed-length coding, by decoding the value n in the header.
[0336] The device may vary the bit length used for fixed-length coding for each component of the attribute information. For example, if the uv coordinates are (100, 60), the device can reduce the amount of coding by applying 7-bit fixed-length coding to u and 6-bit fixed-length coding to v. In this case, the device may add the bit length used for fixed-length coding of each component to the bitstream header. This allows the decoding device to correctly decode the attribute information of the initial three-dimensional point, which has been coded with different bit lengths for each component, by decoding the bit length of each component in the header.
[0337] The device can set the bit length of each component to 0 bits. Setting it to 0 bits may indicate that the component is not included in the bitstream. The decoding device can determine that the component is not included in the bitstream if the bit length added to the header is 0 bits, and estimate the value of that component as 0. For example, if the attribute information contains a component whose value is always 0, the device can reduce the amount of code by setting the bit length of the fixed-length encoding of that component to 0 bits and not encoding the value 0. Also, the decoding device can determine that the component is not included in the bitstream if the bit length added to the header is 0 bits, and estimate the value of that component as 0.
[0338] For example, as explained using Figure 46, the device may project a three-dimensional point G(x, y, z) onto the xz plane, and use the plane face0 to which the two-dimensional coordinate (u, v) on the projected xy plane belongs as attribute information for G. In this case, the device can associate G with face0 as uv coordinates by setting (u, v) = (face0, 0). In this case, since v is always valued at 0, the device can reduce the amount of code by setting the bit length of the fixed-length encoding of v to 0 bits and not including the value of v (0) in the bitstream.
[0339] The device may encode the attribute information of the initial three-dimensional points using variable-length coding such as unary coding or exponential Golomb coding. Furthermore, the device can reduce the amount of code by assigning a context to each bit after binarization and applying arithmetic coding.
[0340] It is not always necessary to apply the no-prediction mode to the initial 3D points; the device may apply other prediction methods. For example, the device can calculate representative values for the initial 3D points and use these representative values as predicted values to predictively encode the attribute information of the initial 3D points. As representative values, for example, the average value of the initial 3D points may be used. In this case, the device adds the representative values to the header, allowing the decoding device to decode the header, obtain the predicted values of the initial 3D points, and decode them correctly. Alternatively, the device may use the attribute information of the initial 3D points that was encoded or decoded immediately before as the predicted value of the attribute information of the initial 3D points. This allows the device to reduce the amount of code by predictively encoding the attribute information of the initial 3D points.
[0341] Next, we will explain the coarse prediction mode. The coarse prediction mode is an example of the second prediction mode. The coarse prediction mode is applied when there is one 3D point among the encoded 3D points that is connected to the 3D point to be encoded.
[0342] When applying coarse prediction, the device can use the attribute information of a single encoded and connected three-dimensional point as the predicted value and encode its predicted residual. Since the trend of the predicted residual may differ from that of fine prediction, the device may use different methods for encoding or decoding the predicted residual in coarse and fine prediction.
[0343] For example, when the device binarizes the prediction residuals using unary coding or exponential Golomb coding, assigns a context to each bit, and applies arithmetic coding, different contexts may be used for coarse prediction and fine prediction. This allows for a reduction in the amount of code even when the trends of the prediction residuals differ between coarse and fine predictions. Furthermore, the device can further reduce the amount of code by using different binarization methods for coarse and fine predictions.
[0344] Next, we will explain the fine prediction mode. The fine prediction mode is an example of the third prediction mode. The fine prediction mode is applied when there are two or more three-dimensional points in the encoded three-dimensional points that are connected to the three-dimensional point to be encoded.
[0345] When applying fine prediction, the device can calculate predicted values using attribute information of two or more coded and connected three-dimensional points, and encode the prediction residuals. Furthermore, if a three-dimensional point with coded attribute information belongs to the vertices of two or more triangles composed of coded three-dimensional points (as shown for three-dimensional point v4 in Figure 52), the device can generate two or more fine prediction values using two or more triangles. In this case, the device may use the average of the two or more fine prediction values as the predicted value (average fine prediction).
[0346] For example, if the device can generate count fine prediction values, it may calculate the sum of these count fine prediction values, divide that sum by count to calculate the average fine prediction value, and use that average fine prediction value as the predicted value for the attribute information to be encoded. This allows the device to improve the accuracy of the prediction values obtained through fine prediction and reduce the amount of coding required.
[0347] Furthermore, the device may restrict the value of count to a power of two. For example, by restricting the count of fine predictions used in the average fine prediction to a power of two, such as 1, 2, 4, or 8, the device can use a right bit shift instead of division during averaging arithmetic, thereby reducing the amount of computation. For example, by applying a 1-bit right shift when count = 2, or a 2-bit right shift when count = 4, the average fine prediction value can be generated without using division.
[0348] In this case, the average fine prediction using a count that is not a power of two is prohibited, and the device may omit its calculation. This reduces the processing load required by the device to calculate the average fine prediction. The device may also set a limit on the number of counts used to calculate the average fine prediction. For example, the device can set a maximum value of count, fineMaxcount, and limit the number of fine predictions used to calculate the average fine prediction to fineMaxcount or less. This reduces the processing load required by the device to calculate the average fine prediction.
[0349] The device may add fineMaxcount to the bitstream header. This allows the decoder to decode the fineMaxcount in the header, using the same maximum value of count as the encoder to calculate the average fine prediction value, and thus correctly decode the bitstream.
[0350] Furthermore, by setting fineMaxcount = 1, the device can limit the number of fine predictions used to calculate the average fine prediction value to one, effectively eliminating the need for average calculation. This reduces the processing load for calculating the average and prevents the generation of decimal values that would otherwise occur during average calculation.
[0351] Furthermore, the device can be configured to not perform fine prediction by setting fineMaxcount = 0. In this case, the prediction mode used will be either (1) no prediction mode or (2) coarse prediction mode. For example, if the attribute information of all three-dimensional points following the initial three-dimensional point is the same, the device can apply no prediction mode to the initial three-dimensional point and coarse prediction to the subsequent three-dimensional points, thereby always keeping the prediction residual at 0, and there is no need to apply fine prediction. In such cases, the device can set fineMaxcount = 0 to not apply fine prediction, thereby reducing the processing load required for encoding or decoding.
[0352] Furthermore, the device can reduce the amount of code by setting fineMaxcount = 0 when additional information necessary for fine prediction exists, thereby not applying fine prediction.
[0353] Furthermore, the predictive coding described above can be similarly applied to predictive decoding. In other words, by replacing "coding" with "decoding" above, the explanation of predictive coding above can be used as an explanation of predictive decoding.
[0354] Figure 54 is a diagram showing an example of a flowchart for encoding attribute information of a three-dimensional point according to this embodiment.
[0355] The encoding device acquires multiple three-dimensional points (vertices) of the three-dimensional mesh (S1501).
[0356] The encoding device determines whether or not there are three-dimensional points for which attribute information has not been encoded (S1502). If there are three-dimensional points for which attribute information has not been encoded (Yes in S1502), the encoding device proceeds to step S1503; otherwise (No in S1502), the process ends.
[0357] The encoding device selects a three-dimensional point to be encoded from among the unencoded three-dimensional points of a plurality of three-dimensional points (S1503).
[0358] The encoding device determines whether there are two or more encoded three-dimensional points that are connected to the three-dimensional point to be encoded (S1504). If there are two or more encoded three-dimensional points that are connected to the three-dimensional point to be encoded (Yes in S1504), the device proceeds to step S1505; otherwise (No in S1504), the device proceeds to step S1508.
[0359] The encoding device calculates a predicted value using attribute information of two or more three-dimensional points (S1505).
[0360] The encoding device determines whether count is greater than 0 (S1506). If count is greater than 0 (Yes in S1506), the encoding device proceeds to step S1507; otherwise (No in S1506), it proceeds to step S1508.
[0361] The encoding device encodes the attribute information of the three-dimensional point to be encoded using the calculated predicted value (S1507). In other words, the encoding device calculates the residual between the attribute information of the three-dimensional point to be encoded and the predicted value calculated using the attribute information of two or more three-dimensional points, and encodes the calculated residual. Step S1507 is predictive encoding in the third prediction mode.
[0362] The encoding device determines whether there is one encoded three-dimensional point that is connected to the three-dimensional point to be encoded (S1508). If there is one encoded three-dimensional point that is connected to the three-dimensional point to be encoded (Yes in S1508), the device proceeds to step S1509; otherwise (No in S1508), the device proceeds to step S1510.
[0363] The encoding device encodes the attribute information of the three-dimensional point to be encoded using the attribute information of one three-dimensional point as a predicted value (S1509). In other words, the encoding device calculates the residual between the attribute information of the three-dimensional point to be encoded and the attribute information of one three-dimensional point, and encodes the calculated residual. Step S1509 is predictive encoding in the second prediction mode.
[0364] The encoding device encodes the attribute information of the three-dimensional point to be encoded in non-prediction mode (S1510). In other words, the encoding device encodes the attribute information of the three-dimensional point to be encoded without using predicted values. Step S1510 is predictive encoding in the first prediction mode.
[0365] Figure 55 is a flowchart showing an example of the calculation process for the predicted value in the third prediction mode according to this embodiment.
[0366] The encoding device sets count to 0 (S1511).
[0367] The encoding device determines whether count is less than fineMaxcount (S1512). If count is less than fineMaxcount (Yes in S1512), the encoding device proceeds to step S1513; otherwise (No in S1512), it proceeds to step S1515. The encoding device can turn off fine prediction by setting fineMaxcount=0. The encoding device may also append the set fineMaxcount value to the bitstream.
[0368] Furthermore, if the value of fineMaxcount decoded from the bitstream is 0, the decoder may decode the attribute information without applying fine prediction.
[0369] The encoding device calculates the count-th fine prediction value (S1513).
[0370] The encoding device adds 1 to the count, updates the count (S1514), and returns to step S1512.
[0371] The encoding device determines whether count is greater than 0 (S1515). If count is greater than 0 (Yes in S1515), the encoding device proceeds to step S1516; otherwise (No in S1515), it proceeds to step S1517.
[0372] The encoding device calculates the average fine prediction value (S1516). In other words, the encoding device calculates the average of one or more fine prediction values as the average fine prediction value. If there is only one fine prediction value, the encoding device may calculate that one fine prediction value as the average fine prediction value. The one or more fine prediction values are the count fine prediction values obtained in steps S1511 to S1514.
[0373] The encoding device outputs the count (S1517).
[0374] Figure 56 is a diagram showing an example of a flowchart for decoding attribute information of a three-dimensional point according to this embodiment.
[0375] The decoding device acquires multiple three-dimensional points (vertices) of the three-dimensional mesh (S1521).
[0376] The decoding device determines whether or not there are any three-dimensional points whose attribute information has not been decoded (S1522). If there are three-dimensional points whose attribute information has not been decoded (Yes in S1522), the decoding device proceeds to step S1523; otherwise (No in S1522), the process ends.
[0377] The decoding device selects a three-dimensional point to be decoded from among the undecoded three-dimensional points out of a group of three-dimensional points (S1523).
[0378] The decoding device determines whether there are two or more decoded three-dimensional points that are connected to the three-dimensional point to be decoded (S1524). If there are two or more decoded three-dimensional points that are connected to the three-dimensional point to be decoded (Yes in S1524), the device proceeds to step S1525; otherwise (No in S1524), the device proceeds to step S1528.
[0379] The decoding device calculates a predicted value using attribute information of two or more three-dimensional points (S1525).
[0380] The decoding device determines whether count is greater than 0 (S1526). If count is greater than 0 (Yes in S1526), the decoding device proceeds to step S1527; otherwise (No in S1526), it proceeds to step S1528.
[0381] The decoding device decodes the attribute information of the three-dimensional point to be decoded using the calculated predicted value (S1527). In other words, the decoding device calculates the residual between the attribute information of the three-dimensional point to be decoded and the predicted value calculated using the attribute information of two or more three-dimensional points, and decodes the calculated residual. Step S1527 is predictive decoding in the third prediction mode.
[0382] The decoding device determines whether there is one decoded 3D point that is connected to the 3D point to be decoded (S1528). If there is one decoded 3D point that is connected to the 3D point to be decoded (Yes in S1528), the device proceeds to step S1529; otherwise (No in S1528), the device proceeds to step S1530.
[0383] The decoding device decodes the attribute information of the three-dimensional point to be decoded using the attribute information of one three-dimensional point as a predicted value (S1529). In other words, the decoding device calculates the residual between the attribute information of the three-dimensional point to be decoded and the attribute information of one three-dimensional point, and decodes the calculated residual. Step S1529 is predictive decoding in the second prediction mode.
[0384] The decoding device decodes the attribute information of the three-dimensional point to be decoded in a no-prediction mode (S1530). In other words, the decoding device decodes the attribute information of the three-dimensional point to be decoded without using predicted values. Step S1530 is predictive decoding in the first prediction mode.
[0385] Next, we will describe mesh_attribute_coding, which is a unit for encoding attribute information of three-dimensional points. Figure 57 shows an example of the syntax of mesh_attribute_coding according to this embodiment. Figure 58 shows an example of the syntax of mesh_attribute_header according to this embodiment.
[0386] `mesh_attribute_coding` indicates a unit for encoding attribute information of a three-dimensional point, `mesh_attribute_header` indicates header information related to the attribute information, and `mesh_attribute_data` indicates encoded data related to the attribute information.
[0387] AttributeBitDepthMinus1 is information indicating the bit precision of the attribute information encoded in the bitstream. The bit precision of the attribute information is calculated by adding 1 to AttributeBitDepthMinus1.
[0388] NumComponent indicates the number of components in the attribute information. For example, if the attribute information is a UV coordinate, the NumComponent value is 2, and if the attribute information is a normal vector, the NumComponent value is 3.
[0389] AttributeStartDeltaBitDepth[j] is information indicating the bit precision of the initial three-dimensional point of the attribute information encoded in the bitstream. The bit precision of the j-th component of the initial three-dimensional point, AttributeStartBitDepth[j], is calculated as follows:
[0390] AttributeStartBitDepth[j] = AttributeBitDepthMinus1 + 1 - AttributeStartDeltaBitDepth[j]
[0391] For example, if AttributeBitDepthMinus1 is 9, AttributeStartDeltaBitDepth[0] is 0, and AttributeStartDeltaBitDepth[1] is 2: AttributeStartBitDepth[0] = 9 + 1 - 0 = 10 AttributeStartBitDepth[1] = 9 + 1 - 2 = 8 This shows that the bit precision of the 0th component of the initial 3D point is 10 bits, and the bit precision of the 1st component is 8 bits.
[0392] In this way, by setting AttributeStartDeltaBitDepth for each component of the initial three-dimensional point, different bit precisions can be set for each component. Furthermore, by setting the bit precision of the initial three-dimensional point separately from AttributeBitDepthMinus1, which is the bit precision of the attribute information encoded in the bitstream, if, for example, the range of values that each component of the initial three-dimensional point can take is smaller than the range of values that each component of the attribute information encoded in the bitstream can take, the value of AttributeStartBitDepth may be smaller than the value obtained by adding 1 to AttributeBitDepthMinus1. In such cases, the amount of code can be reduced by encoding AttributeStart, which is the value of each component of the initial three-dimensional point described later, using the value of AttributeStartBitDepth instead of the value obtained by adding 1 to AttributeBitDepthMinus1.
[0393] Furthermore, in this embodiment, the difference between the value obtained by adding 1 to the bit precision of the attribute information encoded in the bitstream, AttributeBitDepthMinus1, and the bit precision of the j-th component of the initial three-dimensional point may be encoded as AttributeStartDeltaBitDepth. This reduces the amount of coding required because, when the bit precision of the attribute information and the bit precision of the initial three-dimensional point are the same, the difference value added to the bitstream becomes 0.
[0394] Furthermore, the value of AttributeStartBitDepth[j] can be set to 0 by setting AttributeStartDeltaBitDepth[j] to the same value as AttributeBitDepthMinus1 plus 1. In this case, the j-th component, for which AttributeStartBitDepth[j] is 0, does not need to be encoded with the value 0 in the bitstream. The decoder can infer that the j-th component for which AttributeStartBitDepth[j] is 0 has the value 0. This reduces the amount of code by not adding the component with AttributeStartBitDepth[j] of 0 to the bitstream.
[0395] In this embodiment, an example is shown in which the difference value, AttributeStartDeltaBitDepth[j], is added to the bitstream. However, this is not the only option; for example, AttributeStartBitDepth[j] may be directly encoded into the bitstream. This method reduces the amount of processing required to calculate the difference value.
[0396] Furthermore, the bit length of AttributeStartDeltaBitDepth[j] can be calculated based on the value of AttributeBitDepthMinus1. Specifically, the bit length of AttributeStartDeltaBitDepth[j] can be calculated as ceil(log2(AttributeBitDepthMinus1 + 2)). Here, ceil(x) is a function that outputs an integer value rounded up from x. In this way, by determining the bit length of AttributeStartDeltaBitDepth[j] according to the value of AttributeBitDepthMinus1, the amount of sign can be reduced.
[0397] AttributePredMode is information indicating how attribute information is encoded. For example, it may indicate a prediction method methodA, which includes fine prediction, coarse prediction, and no prediction mode, as shown in Figure 54. Note that the prediction method is not limited to methodA. For example, it may be a prediction method that combines two of the three prediction modes of methodA, or a prediction method that uses only one prediction mode. Specifically, a prediction method methodB, which includes coarse prediction and no prediction mode, is conceivable. By providing methodB as a prediction method, predicted values can be generated without applying arithmetic operations that require division or decimal precision, thus reducing processing load and enabling lossless encoding.
[0398] fineMaxcount is information indicating the maximum number of fine predictions used when calculating the average fine prediction value. If AttributePredMode includes fine predictions, for example, if AttributePredMode is methodA, fineMaxcount may be added to the bitstream.
[0399] Furthermore, by setting fineMaxcount to 1, the number of fine predictions used to calculate the average fine prediction value can be limited to one, effectively eliminating the need to calculate the average. This reduces the processing load required to calculate the average. It also prevents the generation of decimal values due to the calculation of the average.
[0400] Alternatively, you can set fineMaxcount to 0 to disable fine prediction. In this case, the prediction mode will be either coarse prediction or no prediction mode. For example, if the attribute information of all three-dimensional points following the initial three-dimensional point is the same, you can apply no prediction mode to the initial three-dimensional point and coarse prediction to the subsequent three-dimensional points to make the prediction residual zero, thus eliminating the need to apply fine prediction. In such cases, disabling fine prediction by setting fineMaxcount to 0 can reduce the processing load required for encoding or decoding.
[0401] Furthermore, even if additional information is required for fine prediction, setting fineMaxcount to 0 prevents the application of fine prediction, thereby reducing the amount of additional information and thus the amount of code.
[0402] AttributeBitDepthMinus1, AttributePredMode, and fineMaxcount may be binarized using Exponential Golomb coding or unary coding, and variable-length coding may be applied. Furthermore, coding efficiency can be improved by assigning a context to each bit and updating the probability of occurrence based on the frequency of 0s and 1s. Alternatively, this information may be coded in a fixed length, thereby reducing processing load.
[0403] Furthermore, fineMaxcount may be set to a different value per sequence, per frame, per basemesh, per submesh, or per 3D point. This allows for a reduction in processing load by controlling the number of points used for average fine prediction in appropriate units.
[0404] Figure 59 shows an example of the syntax of mesh_attribute_data according to this embodiment.
[0405] AttributeStartCount is information indicating the number of initial three-dimensional points. AttributeStart[c][j] is information indicating the value of the j-th component of the attribute information for the c-th initial three-dimensional point. For example, if the attribute information is in uv coordinates, AttributeStart[c][0] may indicate the value of u for the c-th initial three-dimensional point, and AttributeStart[c][1] may indicate the value of v for the c-th initial three-dimensional point. Note that AttributeStart[c][j] may be encoded using a fixed-length encoding method with the bit length calculated as AttributeStartBitDepth as explained in Figure 58. For example, if AttributeStartBitDepth[0] is 10 and AttributeStartBitDepth[1] is 8, it can be indicated that AttributeStart[c][0], which is the 0th component of the c-th initial three-dimensional point, is encoded with 10 bits of precision, and AttributeStart[c][1], which is the 1st component, is encoded with 8 bits of precision. The amount of encoding can be reduced by encoding the 0th component with 10 bits and the 1st component with 8 bits of precision.
[0406] Furthermore, if the value of AttributeStartBitDepth[j] is 0, AttributeStart[c][j] does not need to be encoded into the bitstream. The decoder can infer that the value of AttributeStart[c][j] is 0 by checking that the value of AttributeStartBitDepth[j] is 0. In this way, the amount of code can be reduced by not including the component with a value of AttributeStartBitDepth[j] of 0 in the bitstream.
[0407] Furthermore, AttributeStart may apply variable-length coding such as unary coding or exponential Golomb coding. Additionally, the code size can be reduced by assigning a context to each bit after binarization and then applying arithmetic coding.
[0408] Furthermore, if the value of AttributeStartCount is 0, there is no need to encode AttributeStart. In this case, the decoder can check that AttributeStartCount is 0 and estimate that the value of AttributeStart is 0. This method can reduce the amount of code.
[0409] Furthermore, if it is guaranteed that there will always be at least one initial three-dimensional point, AttributeStartCountMinus1 can be added to the bitstream instead of AttributeStartCount. In this case, the decoder can calculate AttributeStartCount by adding 1 to AttributeStartCountMinus1. This makes it possible to further reduce the amount of code.
[0410] AttributeFineResiCount indicates the number of prediction residuals for attribute information predicted and coded by fine prediction. In particular, when the value of fineMaxcount is 0, it is assumed that AttributeFineResiCount will be set to 0 as the most efficient case.
[0411] AttributeFineResiSize represents information indicating the code size of AttributeFineResi. For example, it may indicate the code size after binarizing and arithmetic coding of AttributeFineResi. By notifying the decoder of the code size via AttributeFineResiSize before AttributeFineResi, the decoder can pre-read the bits related to AttributeFineResi from the bitstream. Specifically, while the decoder is performing arithmetic decoding on AttributeFineResi, the decoding process for the subsequent AttributeCoarseResi can be performed in parallel, thus reducing the decoding time.
[0412] AttributeFineResi[c][j] represents the prediction residual for the j-th component of the attribute information predicted and coded by the c-th fine prediction. AttributeFineResi[c][j] can be binarized using methods such as Exponential Golomb coding or unary coding, and variable-length coding can be performed. Furthermore, coding efficiency can be improved by assigning a context to each bit and updating the probability of occurrence based on the frequency of occurrences of 0 and 1 during coding.
[0413] Furthermore, the context used for arithmetic coding of each bit may be shared with the context used in AttributeCoarseResi, which will be described later. This method can reduce the number of contexts. On the other hand, the context used for arithmetic coding of each bit can also be set to be different from the context used in AttributeCoarseResi, which will be described later. By using different contexts in this way, coding efficiency can be improved even when the trends of prediction residuals differ between fine prediction and coarse prediction. In addition, by using different contexts, AttributeFineResi and AttributeCoarseResi can be decoded in parallel, which can further improve the efficiency of the decoding process.
[0414] AttributeCoarseResiCount indicates the number of predicted residuals in attribute information predicted and coded using coarse prediction. AttributeCoarseResiCount plays an important role in the coding of AttributeCoarseResi, which will be discussed later.
[0415] AttributeCoarseResiSize represents information indicating the code size of AttributeCoarseResi. For example, it may indicate the code size when AttributeCoarseResi is binarized and arithmetic coded. By informing the decoder of AttributeCoarseResiSize prior to coding AttributeCoarseResi, the decoder can pre-read the bits related to AttributeCoarseResi from the bitstream. Specifically, while the decoder is performing arithmetic decoding on AttributeCoarseResi, the decoding process of the subsequent bitstream can be performed in parallel, thus reducing decoding time.
[0416] AttributeCoarseResi[c][j] represents the prediction residual for the j-th component of the attribute information predicted and coded by the c-th coarse prediction. AttributeCoarseResi[c][j] can be binarized using methods such as Exponential Golomb coding or unary coding, and variable-length coding can be performed. Furthermore, coding efficiency can be improved by assigning a context to each bit and updating the probability of occurrence based on the frequency of occurrences of 0 and 1 during coding.
[0417] Furthermore, the context used for arithmetic coding of each bit may be shared with the context used in AttributeFineResi as described above. This method can reduce the number of contexts. On the other hand, the context used for arithmetic coding of each bit can also be set to be different from the context used in AttributeFineResi as described above. By using different contexts in this way, coding efficiency can be improved even when the trends of prediction residuals differ between fine prediction and coarse prediction.
[0418] Furthermore, AttributeStartCount, AttributeFineResiCount, AttributeFineResiSize, AttributeFineResi, AttributeCoarseResiCount, AttributeCoarseResiSize, and AttributeCoarseResi may be binarized using Exponential Golomb coding or unary coding, and variable-length coding may be performed. Additionally, coding efficiency can be improved by assigning a context to each bit and updating the probability of occurrence based on the frequency of 0s and 1s. Moreover, this information can also be coded in a fixed length, reducing the amount of processing required.
[0419] This embodiment primarily describes examples relating to the encoding or decoding of attribute information of three-dimensional points, but is not necessarily limited to this. For example, this embodiment may also be applied to the encoding or decoding of positional information of three-dimensional points. An example of the syntax when applied to positional information of three-dimensional points will be described later.
[0420] In this embodiment, for example, as shown in Figure 46, the device projects a three-dimensional point G(x, y, z) onto the xz plane, and the plane face0 to which the two-dimensional coordinate (u, v) on the projected xy plane belongs is used as attribute information for G. In this embodiment, an example is shown in which G and face0 are associated as uv coordinates by setting (u, v) = (face0, 0). However, this is not necessarily the only example. For example, when encoding or decoding face0 as uv coordinates, the attribute information is uv coordinates, but by setting the attribute information to NumComponent=1, only u from the uv coordinates can be encoded. This reduces the amount of code by encoding face0 as u and not encoding v.
[0421] Connected components can be projected onto a two-dimensional image, and the same planar information can be added as attribute information to these three-dimensional components, and this attribute information can be encoded. This allows the decoder to decode the attribute information of the three-dimensional components to obtain the planar information associated with each component, and use that value to obtain, for example, related projection information. The method of generating the UV coordinates of the three-dimensional components on the decoder side using this projection information may be called an "ortho atlas."
[0422] When using an ortho atlas to encode planar information as UV coordinates, the attribute information of connected three-dimensional points will all have the same value (face0, 0), as shown in Figure 60, for example. Figure 60 shows an example of attribute information related to a modified example of this embodiment.
[0423] In this case, the device encodes the initial three-dimensional points in no-prediction mode and applies coarse encoding to subsequent three-dimensional points, thereby reducing the prediction residual to zero. Therefore, when using ortho atlas, setting fineMaxcount=0 as described in this embodiment prevents the application of fine prediction, thereby reducing the processing load for encoding or decoding. Furthermore, if additional information necessary for fine prediction exists, setting fineMaxcount=0 prevents the application of fine prediction, thereby reducing the amount of additional information and thus the amount of encoding required.
[0424] Furthermore, in uv coordinates, v is always 0. Therefore, when using an ortho atlas, by using AttributeBitDepthMinus1 and AttributeStartDeltaBitDepth described in this embodiment, setting AttributeStartBitDepth for the v component of the initial three-dimensional point to 0 and setting the bit length of the fixed-length encoding of the v component to 0 bits, the amount of coding can be reduced by not encoding the value of v which is 0.
[0425] Furthermore, when using an ortho atlas, the prediction residual can be reduced to zero by applying coarse or fine prediction. Therefore, the code size may be reduced by not encoding at least one of the following information related to the prediction residual: AttributeFineResiCount, AttributeFineResiSize, AttributeFineResi, AttributeCoarseResiCount, AttributeCoarseResiSize, and AttributeCoarseResi. The decoder can estimate the value of the unencoded information to be 0.
[0426] Furthermore, when using an ortho atlas and not applying fine prediction, the amount of code can be reduced by not encoding the additional information associated with the fine prediction.
[0427] Figure 61 shows an example of the syntax of mesh_attribute_data according to a modified example of this embodiment.
[0428] If AttributeFineResiCount or AttributeCoarseResiCount is not encoded, the number of 3D points connected to the initial 3D point, OtherAttributeCount, may be added to the bitstream for the same number of 3D points as there are initial 3D points. Figure 60 shows an example of this syntax. In this case, the decoder can correctly decode the uv coordinates by decoding the uv coordinates of the initial 3D point c and copying those uv coordinates to the uv coordinates of the OtherAttributeCount[c] 3D points connected to the initial 3D point c.
[0429] Furthermore, OtherAttributeCount may be binarized using methods such as Exponential Golomb coding or unary coding, and variable-length coding may be performed. By assigning a context to each bit and coding while updating the probability of occurrence based on the frequency of occurrences of 0 and 1, coding efficiency can be improved. Alternatively, it may be coded with a fixed length. In this case, the amount of processing can be reduced.
[0430] Furthermore, when defining the AttributePredMode for the ortho atlas and the encoder encodes the attribute information using the method described above, information indicating that the AttributePredMode is the prediction method for the ortho atlas may be added to the bitstream. Thereby, the decoder can determine whether the bitstream is encoded using the above method by decoding the AttributePredMode. If the determination result is Yes, the decoder can decode the attribute information using the above method.
[0431] In this embodiment, the encoding or decoding method for the attribute information corresponding to three-dimensional points such as uv coordinates is shown, but it is not necessarily limited to this. For example, the encoding or decoding method described in this embodiment may also be applied to the attribute information corresponding to the polygonal surface composed of three-dimensional points.
[0432] Specifically, the face_id, which is the face information corresponding to the polygonal surface, can be encoded or decoded as the attribute information with NumComponent = 1. By doing so, it is possible to reduce the amount of code for the attribute information corresponding to the three-dimensional points or the attribute information corresponding to the polygon composed of the three-dimensional points using this embodiment
[0433] FIG. 62 is a diagram showing an example of the geometry map according to this embodiment. FIG. 63 is a diagram showing an example of the texture map according to this embodiment.
[0434] Texture mapping is a common technique used to add visual appearance by projecting an image onto the surface of a 3D model. This technique can improve performance when rendering is required because it allows changing visual details without directly modifying the model itself. The information stored in the 2D texture image may include color, smoothness, transparency, etc. The relationship between the 3D mesh and the 2D texture image is established using a UV map. The UV map is a transformation and unfolding of the 3D mesh onto a 2D plane. Next, by overlaying the texture image on the UV map, the attribute values from the image are projected onto the 3D model during rendering. Therefore, encoding and decoding techniques that can efficiently store and stream the UV map are required.
[0435] Different prediction schemes are available for encoding the texture coordinates of the UV map. These prediction schemes include those based on the parameterization technique used when projecting the 3D mesh onto the 2D map, or those based on the relationship between the geometric data of the mesh and the texture coordinates. In the case of isometric parameterization, the edge lengths of the triangle maintain approximately the same ratio in both 3D space and UV coordinates. Therefore, the texture coordinates on the UV map can be predicted using the ratio obtained from the corresponding edge in 3D space.
[0436] As shown in FIG. 62, the texture coordinates A2, B2, and C2 respectively correspond to A1, B1, and C1 on the geometry map. The ratio between the edges A1B1 and A2B2 in 3D space is calculated as (A1C1) / (C1B1)=R. Finally, the coordinates of A2 in UV space are predicted as A2C2 = R×C2B2 and A2 = C2 + A2C2.
[0437] As shown in Figure 63, the device can calculate a predicted value for the uv coordinate A2 when a three-dimensional point A1 is projected onto a two-dimensional image (texture map). The device may also calculate this predicted value using already encoded or decoded C2 and B2, and the positional information of three-dimensional points A1, B1, and C1 corresponding to A2, B2, and C2.
[0438] More specifically, the device assumes that triangles A1B1C1 and A2B2C2 are similar, and can calculate A2C2 by multiplying the already determined value of C2B2 by the similarity ratio R, which is calculated as (A1C1) / (C1B1) = R. The device can also predict the position of A2 using C2 + A2C2. As a result, when triangles composed of corresponding three-dimensional points are similar in the geometry map and texture map, the device can accurately estimate the predicted values of the attribute information of the three-dimensional points to be encoded or decoded using the encoded or decoded three-dimensional point information in the geometry map and the encoded or decoded uv coordinates in the texture map, thereby reducing the amount of coding required.
[0439] Furthermore, the method of calculating predicted values of attribute information of a three-dimensional point to be encoded or decoded using the three-dimensional point information in the encoded or decoded geometry map and the uv coordinates in the encoded or decoded texture map, as described above, may also be referred to as the fine prediction in this embodiment.
[0440] Next, we will explain mesh_position_coding, which is a unit for encoding the position information of three-dimensional points. Figure 64 shows an example of the syntax of mesh_position_coding according to this embodiment. Figure 65 shows an example of the syntax of mesh_position_header according to this embodiment.
[0441] `mesh_position_coding` is a unit for encoding the position information of three-dimensional points, and it performs processing related to the encoding and decoding of position information. `mesh_position_header` indicates header information related to the position information, and `mesh_position_data` indicates encoded data related to the position information.
[0442] PositionBitDepthMinus1 is information indicating the bit precision of the position information encoded in the bitstream. The bit precision of the position information can be calculated by adding 1 to PositionBitDepthMinus1.
[0443] NumComponent represents the number of components in the location information. For example, if the location information is in xyz coordinates, the value of NumComponent can be 3.
[0444] PositionStartDeltaBitDepth[j] is information indicating the bit precision of the initial three-dimensional point of the position information encoded in the bitstream. The bit precision of the j-th component of the initial three-dimensional point, PositionStartBitDepth[j], can be calculated as follows:
[0445] PositionStartBitDepth[j] can be obtained by adding 1 to PositionBitDepthMinus1 and then subtracting PositionStartDeltaBitDepth[j]. For example, if PositionBitDepthMinus1 is 9, PositionStartDeltaBitDepth[0] is 0, PositionStartDeltaBitDepth[1] is 2, and PositionStartDeltaBitDepth[2] is 4, then the value of PositionStartBitDepth[0] is 9 + 1 - 0 = 10, the value of PositionStartBitDepth[1] is 9 + 1 - 2 = 8, and the value of PositionStartBitDepth[2] is 9 + 1 - 4 = 6. This shows that the 0th component of the initial 3D point is 10-bit precision, the 1st component is 8-bit precision, and the 2nd component is 6-bit precision.
[0446] In this way, by setting PositionStartDeltaBitDepth for each component of the initial three-dimensional point, it is possible to set different bit precisions for each component. Furthermore, by allowing the bit precision of the initial three-dimensional point to be set separately from PositionBitDepthMinus1, which is the bit precision of the position information encoded in the bitstream, it is possible that, for example, if the range of values that each component of the initial three-dimensional point can take is smaller than the range of values that each component of the position information encoded in the bitstream can take, the value of PositionStartBitDepth may be smaller than the value obtained by adding 1 to PositionBitDepthMinus1.
[0447] In such cases, the amount of code can be reduced by encoding PositionStart, which is the value of each component of the initial three-dimensional point described later, using the value of PositionStartBitDepth instead of PositionBitDepthMinus1 plus 1. Furthermore, this setting improves encoding efficiency by avoiding the specification of unnecessary bit precision.
[0448] In this embodiment, the difference between the value obtained by adding 1 to PositionBitDepthMinus1, which is the bit precision of the position information encoded in the bitstream, and the bit precision of the j-th component of the initial three-dimensional point can be encoded as PositionStartDeltaBitDepth. With this method, if the value obtained by adding 1 to PositionBitDepthMinus1, which is the bit precision of the position information, and the bit precision of the j-th component of the initial three-dimensional point are the same, the difference value added to the bitstream becomes 0, and the amount of coding can be reduced.
[0449] Furthermore, by setting PositionStartDeltaBitDepth[j] to the same value as PositionBitDepthMinus1 plus 1, the PositionStartBitDepth[j] value of the j-th component can be set to 0. In this case, the j-th component, for which PositionStartBitDepth[j] is 0, does not need to be encoded as 0 in the bitstream. The decoder can infer that the j-th component, for which PositionStartBitDepth[j] is 0, has a value of 0. This reduces the amount of code by not adding components with PositionStartBitDepth[j] of 0 to the bitstream.
[0450] This embodiment demonstrates a method for adding the difference value, PositionStartDeltaBitDepth[j], to the bitstream, but it is not necessarily limited to this method. For example, it is also possible to directly encode PositionStartBitDepth[j] into the bitstream. This method can reduce the amount of processing required to calculate the difference value.
[0451] Furthermore, the bit length of PositionStartDeltaBitDepth[j] can be determined based on the value of PositionBitDepthMinus1. Specifically, the bit length of PositionStartDeltaBitDepth[j] can be calculated as ceil(log2(PositionBitDepthMinus1+2)). Here, ceil(x) is a function that outputs an integer value obtained by rounding up the value of x. In this way, by determining the bit length of PositionStartDeltaBitDepth[j] according to the value of PositionBitDepthMinus1, the amount of sign can be efficiently reduced.
[0452] PositionPredMode is information indicating how to encode location information. For example, methodA may be used as a prediction method that includes fine prediction, coarse prediction, and no prediction mode, as shown in the encoding flowchart. The prediction method is not limited to methodA; a prediction method combining two of the three prediction modes, or a prediction method consisting of one prediction mode, can also be adopted. A specific example is methodB, a prediction method that combines coarse prediction and no prediction mode. By using methodB, predicted values can be generated without applying arithmetic operations that require division or decimal precision, thus reducing the amount of processing and enabling lossless encoding. Furthermore, in the case of location information, techniques such as parallelogram prediction may be included as fine prediction.
[0453] `fineMaxcount` indicates the maximum number of fine predictions used when calculating the average fine prediction value. When `PositionPredMode` includes fine predictions, for example, when `PositionPredMode` is `methodA`, `fineMaxcount` can be added to the bitstream. By setting `fineMaxcount` to 1, the number of fine predictions used in calculating the average fine prediction value can be limited to one, effectively eliminating the need for averaging. This reduces the processing load of averaging and prevents the generation of decimal values due to averaging.
[0454] Furthermore, by setting fineMaxcount to 0, it is possible to disable the application of fine prediction. In this case, the prediction mode can be coarse prediction or no prediction mode. For example, if the position information of all 3D points following the initial 3D point is the same, the prediction residual can always be kept at 0 by applying no prediction mode to the initial 3D point and coarse prediction to the subsequent 3D points. In such cases, setting fineMaxcount to 0 disables the application of fine prediction, thereby reducing the processing load required for encoding or decoding.
[0455] Furthermore, if additional information is required for fine prediction, setting fineMaxcount to 0 can reduce the amount of code by not applying fine prediction and not encoding the additional information.
[0456] Figure 66 shows an example of the syntax of mesh_position_data according to this embodiment.
[0457] PositionStartCount indicates the number of initial three-dimensional points. PositionStart[c][j] indicates the value of the j-th component of the position information for the c-th initial three-dimensional point. For example, if the position information is in xyz coordinates, PositionStart[c][0] may indicate the x value of the c-th initial three-dimensional point, PositionStart[c][1] may indicate the y value of the c-th initial three-dimensional point, and PositionStart[c][2] may indicate the z value of the c-th initial three-dimensional point.
[0458] Note that PositionStart[c][j] may also be encoded using a fixed-length encoding method calculated using the aforementioned PositionStartBitDepth. For example, if PositionStartBitDepth[0] is 10, PositionStartBitDepth[1] is 8, and PositionStartBitDepth[2] is 6, then PositionStart[c][0], which is the 0th component of the c-th initial three-dimensional point, is 10-bit precision, PositionStart[c][1], which is the 1st component, is 8-bit precision, and PositionStart[c][2], which is the 2nd component, is 6-bit precision. The 0th component may be encoded using a fixed-length encoding method of 10 bits, the 1st component using a fixed-length encoding method of 8 bits, and the 2nd component using a fixed-length encoding method of 6 bits, respectively. In this way, the amount of encoding can be reduced by setting an appropriate bit length for each component of the initial three-dimensional point and using it for encoding.
[0459] Furthermore, if PositionStartBitDepth[j] is 0, PositionStart[c][j] does not need to be encoded into the bitstream, and the decoder may estimate the value of PositionStart[c][j] as 0 when PositionStartBitDepth[j] is 0. This method reduces the amount of code by not adding the j-th component, where PositionStartBitDepth[j] is 0, to the bitstream.
[0460] PositionStart may use variable-length coding such as unary coding or Exponential Golomb coding. Alternatively, a context may be assigned to each bit after binarization, and arithmetic coding may be applied. This can reduce the amount of code.
[0461] Furthermore, if PositionStartCount is 0, PositionStart does not need to be encoded, and the decoder may estimate the value of PositionStart as 0 when PositionStartCount is 0. This reduces the amount of code required.
[0462] Furthermore, if there is always at least one initial three-dimensional point, PositionStartCountMinus1 may be added to the bitstream instead of PositionStartCount. In this case, the decoder may calculate PositionStartCount by adding 1 to PositionStartCountMinus1.
[0463] PositionFineResiCount indicates the number of predicted residuals of the location information predicted and coded using fine prediction. PositionFineResiSize indicates the code size of PositionFineResi, and may, for example, indicate the code size when PositionFineResi is binarized and arithmetic coded.
[0464] In this way, by informing the decoder of the code size using PositionFineResiSize before PositionFineResi, the decoder can pre-read the bits related to PositionFineResi from the bitstream. For example, while arithmetic decoding related to PositionFineResi is being performed, the decoding process for the subsequent PositionCoarseResi can be performed in parallel, thereby reducing the decoding time.
[0465] PositionFineResi[c][j] represents the prediction residual of the j-th component of the position information predicted and coded by the c-th fine prediction. PositionFineResi[c][j] can be binarized by Exponential Golomb coding, unary coding, etc., and can be coded with variable-length coding. Also, a context can be assigned to each bit, and coding can be performed while updating the occurrence probability based on the occurrence frequencies of 0 and 1. This can improve the coding efficiency.
[0466] Furthermore, the context used for arithmetic coding of each bit may be shared with the context used for PositionCoarseResi described later. By this method, the number of contexts can be reduced. Also, the context used for arithmetic coding of each bit may be different from the context used for PositionCoarseResi described later. By doing so, the coding efficiency in the case where the tendency of the prediction residual is different between fine prediction and coarse prediction can be improved. Also, by using different contexts, it is possible to decode PositionFineResi and PositionCoarseResi in parallel.
[0467] PositionCoarseResiCount represents the information indicating the number of prediction residuals of the position information predicted and coded by coarse prediction. PositionCoarseResiSize represents the information indicating the amount of code of PositionCoarseResi. For example, it may indicate the amount of code when PositionCoarseResi is binarized and arithmetic-coded.
[0468] Thus, by notifying the decoder of the amount of code by PositionCoarseResiSize before PositionCoarseResi, the decoder can pre-read the bits related to PositionCoarseResi from the bit stream. For example, while proceeding with the arithmetic decoding related to PositionCoarseResi, the decoding process of the subsequent bit stream can be performed in parallel, and the decoding time can be reduced.
[0469] PositionCoarseResi[c][j] represents the predicted residual of the j-th component of the position information predicted and coded by the c-th coarse prediction. PositionCoarseResi[c][j] may be binarized by methods such as Exponential Golomb coding or unary coding, and variable-length coding may be performed. Alternatively, a context may be assigned to each bit, and the probability of occurrence may be updated based on the frequency of occurrences of 0 and 1 while coding. This can improve coding efficiency.
[0470] Furthermore, the context used for arithmetic coding of each bit may be shared with the context used in PositionFineResi described above. This method can reduce the number of contexts. Alternatively, the context used for arithmetic coding of each bit may be different from the context used in PositionFineResi described above. This improves coding efficiency when the trends of prediction residuals differ between fine and coarse predictions.
[0471] Furthermore, by using different contexts, it is possible to decode PositionFineResi and PositionCoarseResi in parallel.
[0472] Furthermore, PositionStartCount, PositionFineResiCount, PositionFineResiSize, PositionFineResi, PositionCoarseResiCount, PositionCoarseResiSize, and PositionCoarseResi may be binarized using Exponential Golomb coding or unary coding, and then encoded in variable length. Alternatively, a context may be assigned to each bit, and the probability of occurrence may be updated based on the frequency of occurrences of 0 and 1 while encoding. This can improve encoding efficiency. It is also possible to encode this information in fixed length, which can reduce the amount of processing required.
[0473] (Configuration Example) Figure 67 is a diagram showing an example of the configuration of the encoding device in this embodiment. Figure 68 is a flowchart showing an example of the encoding method by the encoding device in this embodiment.
[0474] The encoding device 1530 includes a circuit 1531 and a memory 1532 connected to the circuit 1531. The encoding device 1530 may perform the processes described in Figures 54 and 55.
[0475] Circuit 1531 performs the following operations.
[0476] Circuit 1531 acquires multiple vertices included in the three-dimensional mesh (S1541). Circuit 1531 encodes the attribute information of one of the multiple vertices using one of at least three prediction modes (S1542). The three prediction modes include a first prediction mode that does not predict, a second prediction mode that performs predictive encoding based on one processed vertex, and a third prediction mode that performs predictive encoding based on two or more processed vertices. If the number of components of the attribute information is set to 1, in encoding (S1542), the attribute information is encoded using a prediction mode different from the third prediction mode.
[0477] According to this approach, the computational load can be reduced by not applying the third prediction mode (predictive coding based on two or more processed vertices) when the number of components is set to 1. Furthermore, the efficiency of the coding process can be improved by eliminating redundant prediction calculations.
[0478] For example, if one vertex has connectivity with other vertices among multiple vertices, the number of components in the attribute information is set to 1.
[0479] According to this, by requiring connectivity when the number of components is 1, it becomes possible to apply an efficient encoding scheme that assumes connectivity. This improves the compression efficiency of encoded data.
[0480] For example, attribute information is specific attribute information.
[0481] According to this method, by encoding specific attribute information, it becomes easier to clarify the target of encoding and to optimize it to improve encoding efficiency.
[0482] For example, specific attribute information is planar information that indicates the region to which a single vertex belongs when that vertex is projected onto a two-dimensional plane.
[0483] According to this method, three-dimensional information can be efficiently represented by encoding based on projection onto a two-dimensional plane. Furthermore, by encoding planar information as attribute information, data simplification and encoding efficiency can be improved.
[0484] For example, the attribute information of one vertex is common to the attribute information of other vertices.
[0485] Therefore, when multiple connected vertices share the same attribute information, redundant data can be omitted, thus reducing the amount of code required.
[0486] For example, the prediction residual is 0 when one vertex is encoded in the second prediction mode.
[0487] Therefore, since the prediction residual is zero when the prediction by the second prediction mode perfectly matches, the coding efficiency can be further improved.
[0488] Figure 69 is a diagram showing an example of the configuration of the decoding device in this embodiment. Figure 70 is a flowchart showing an example of a decoding method by the decoding device in this embodiment.
[0489] The decoding device 1540 includes a circuit 1541 and a memory 1542 connected to the circuit 1541. The decoding device 1540 may perform the processing described in Figure 56.
[0490] Circuit 1541 performs the following operations.
[0491] Circuit 1541 acquires multiple vertices included in the three-dimensional mesh (S1551). Circuit 1541 decodes the attribute information of one of the multiple vertices using one of at least three prediction modes (S1552). The three prediction modes include a first prediction mode that does not predict, a second prediction mode that performs prediction decoding based on one processed vertex, and a third prediction mode that performs prediction decoding based on two or more processed vertices. If the number of components of the attribute information is set to 1, in decoding (S1552), the attribute information is decoded using a prediction mode different from the third prediction mode.
[0492] According to this, the efficiency of the decoding process can be improved by not using the third prediction mode when the number of components is 1 during decoding. In addition, the computational load of the decoding process can be reduced by omitting redundant processing.
[0493] For example, if one vertex has connectivity with other vertices among multiple vertices, the number of components in the attribute information is set to 1.
[0494] According to this, by defining the number of components as 1 when connectivity exists, efficient decoding based on connectivity becomes possible, and the processing load can be reduced.
[0495] For example, attribute information is specific attribute information.
[0496] According to this, by performing a decryption process on specific attribute information, the target of decryption can be clarified, and the efficiency of decryption can be improved.
[0497] For example, specific attribute information is planar information that indicates the region to which a single vertex belongs when that vertex is projected onto a two-dimensional plane.
[0498] According to this, efficient decoding of attribute information for a three-dimensional mesh can be achieved by performing decoding processing based on two-dimensional planar information.
[0499] For example, the attribute information of one vertex is common to the attribute information of other vertices.
[0500] Therefore, when using common attribute information, redundant decoding of attribute information for multiple vertices can be omitted, improving decoding efficiency.
[0501] For example, the prediction residual is 0 when one vertex is decoded using the second prediction mode.
[0502] Therefore, the load on the decoding process can be reduced, and the efficiency of data decoding can be improved.
[0503] Furthermore, the encoding and decoding devices of this disclosure may be implemented by combining at least a part of other embodiments. In addition, the configuration or processing of the encoding and decoding devices of this disclosure may be implemented by combining a part of the processing shown in any of the flowcharts according to this embodiment, a part of the configuration of either device, a part of the syntax, etc., with other embodiments.
[0504] Furthermore, the processing performed by the decoding device of this disclosure may also be performed in the encoding device.
[0505] Furthermore, the circuit 1531 of the encoding device 1530 may perform the following operations. This example of operation is the second example.
[0506] Circuit 1531 acquires multiple vertices included in the three-dimensional mesh (S1541). Circuit 1531 arithmetically encodes one of the multiple vertices using the first prediction mode (second prediction mode in the above embodiment) or the second prediction mode (third prediction mode in the above embodiment) (S1542). The first context used in the arithmetic encoding of the first prediction mode and the second context used in the arithmetic encoding of the second prediction mode are common to both.
[0507] According to this, using a common context across different prediction modes (first prediction mode and second prediction mode) can improve the efficiency of context management. This reduces the processing load required for context initialization or updating, thereby improving the efficiency of the coding process. Furthermore, by sharing the context, coding efficiency may improve even with different parameters if the data shows similar trends in arithmetic coding. Sharing the context reduces the use of redundant context memory, thereby increasing coding efficiency.
[0508] For example, the first prediction mode is a prediction mode that performs predictive coding based on one processed vertex. The second prediction mode is a prediction mode that performs predictive coding based on two or more processed vertices.
[0509] According to this, by using two different prediction modes—a first prediction mode (prediction based on one processed vertex) and a second prediction mode (prediction based on two or more processed vertices)—it is possible to improve prediction accuracy depending on the situation. In particular, prediction based on two or more processed vertices can achieve more accurate coding. Furthermore, by sharing the context for the residual attribute information generated by separate predictions, the coding efficiency of arithmetic coding can be improved while reducing the number of contexts.
[0510] For example, the arithmetic coding of a single vertex includes the arithmetic coding of the position information of that single vertex.
[0511] According to this method, encoding location information allows for highly efficient encoding of the structure or vertex arrangement of a three-dimensional mesh. Applying context sharing to location information further enhances the efficiency of arithmetic coding. In particular, sharing context for residuals of location information generated with different prediction modes enables efficient encoding.
[0512] For example, the arithmetic coding of a single vertex includes the arithmetic coding of the attribute information of that single vertex.
[0513] According to this, including encoding of attribute information (e.g., color, normal vectors, etc.) can contribute to improving the rendering or display quality of 3D meshes. Applying a common context to attribute information can also increase encoding efficiency. Furthermore, sharing the context for residuals of attribute information generated by separate predictions can further improve encoding efficiency.
[0514] For example, multiple vertices belong to a single group based on a three-dimensional mesh. That is, multiple vertices belong to a group of multiple vertices that make up the same three-dimensional mesh, or to a group of multiple vertices that make up the same submesh.
[0515] According to this method, using a common context when encoding vertices within the same group can improve the prediction accuracy between related vertices. Furthermore, by sharing the context at the mesh or submesh level, the number of contexts can be reduced, further improving encoding efficiency.
[0516] Furthermore, the circuit 1541 of the decoding device 1540 may perform the following operations. This example of operation is the second example.
[0517] Circuit 1541 acquires multiple vertices included in the three-dimensional mesh (S1551). Circuit 1541 arithmetically decodes one of the multiple vertices using the first prediction mode (second prediction mode in the above embodiment) or the second prediction mode (third prediction mode in the above embodiment) (S1552). The first context used in the arithmetic decoding of the first prediction mode and the second context used in the arithmetic decoding of the second prediction mode are common to both.
[0518] According to this, the decoding process can be made more efficient by using a common context across different prediction modes. By sharing a common context, memory usage and processing load during arithmetic decoding can be reduced. Furthermore, by using a common context, decoding efficiency can be improved even with different parameters if the trends are the same.
[0519] For example, the first prediction mode is a prediction mode that performs prediction and decoding based on one processed vertex. The second prediction mode is a prediction mode that performs prediction and decoding based on two or more processed vertices.
[0520] According to this approach, by using a first prediction mode (predictive decoding based on one processed vertex) and a second prediction mode (predictive decoding based on two or more processed vertices), efficient decoding corresponding to the prediction method used during encoding becomes possible. Furthermore, by supporting multiple prediction methods, the accuracy of the decoding process can be improved. By sharing the context, the amount of computation and memory consumption in the decoding process can be reduced.
[0521] For example, arithmetic decoding of a single vertex includes arithmetic decoding of the position information of that single vertex.
[0522] According to this method, it is possible to handle cases where location information is encoded during encoding and efficiently restore location information during decoding. By applying a common context to location information, decoding efficiency can be improved. Furthermore, by using a common context with different prediction modes for location information, the processing load can be reduced.
[0523] For example, arithmetic decoding of a single vertex includes arithmetic decoding of the attribute information of that single vertex.
[0524] According to this method, it can handle cases where attribute information is encoded during encoding, and efficiently restore attribute information during decoding. Applying a common context to attribute information can improve decoding efficiency. Furthermore, common context for residuals of attribute information generated in different prediction modes enables efficient decoding.
[0525] For example, multiple vertices belong to a single group based on a three-dimensional mesh. That is, multiple vertices belong to a group of multiple vertices that make up the same three-dimensional mesh, or to a group of multiple vertices that make up the same submesh.
[0526] According to this, using a common context at the mesh or submesh level can improve processing efficiency during decoding. In particular, when processing multiple vertices as a single group, it is effective in efficiently decoding while reducing redundant context.
[0527] Figure 71 is a flowchart showing another example of the calculation process for the predicted value in the third prediction mode according to this embodiment. In other words, it is a modified version of the flowchart in Figure 55.
[0528] The encoding device sets count to 0 (S1561).
[0529] The encoding device determines whether count is less than fineMaxcount (S1562). If count is less than fineMaxcount (Yes in S1562), the encoding device proceeds to step S1563; otherwise (No in S1562), it proceeds to step S1565. The encoding device can turn off fine prediction by setting fineMaxcount=0. The encoding device may also add the set fineMaxcount value to the bitstream. The decoding device may also decode the attribute information without applying fine prediction if the fineMaxcount value decoded from the bitstream is 0.
[0530] The encoding device calculates the count-th fine prediction value with decimal precision (S1563).
[0531] The encoding device adds 1 to the count and updates the count (S1564), then returns to step S1562.
[0532] The encoding device determines whether count is greater than 0 (S1565). If count is greater than 0 (Yes in S1565), the encoding device proceeds to step S1566; otherwise (No in S1565), it proceeds to step S1568.
[0533] The encoding device calculates the average fine prediction value with decimal precision (S1566). In other words, the encoding device calculates the average of one or more fine prediction values as the average fine prediction value. The one or more fine prediction values are the count fine prediction values obtained in steps S1561 to S1564. If there is only one fine prediction value, that single fine prediction value may be calculated as the average fine prediction value. Note that in S1566, it is sufficient to calculate one fine prediction value by integrating two or more fine prediction values, and is not limited to calculating the average fine prediction value.
[0534] The encoding device adjusts the number of bits in the fractional part of the average fine prediction value (integrated value) to be small (S1567). For example, the encoding device adjusts the number of bits in the fractional part to 0. In other words, the encoding device converts the average fine prediction value (integrated value) to integer precision (integer). The encoding device may also perform rounding operations, such as rounding to the nearest integer, when adjusting the number of bits.
[0535] The encoding device outputs the count (S1568).
[0536] Next, we will explain an example of obtaining fine-predicted values in UV coordinates on a two-dimensional plane using information from three-dimensional space.
[0537] Figure 72 shows an example of a vector in three-dimensional space when performing fine prediction according to this embodiment. Figure 73 shows an example of a vector in a two-dimensional plane when performing fine prediction according to this embodiment.
[0538] As shown in Figure 72, two encoded three-dimensional points, G(next) and G(prev), are connected to a three-dimensional point G(curr). Also, as shown in Figure 73, the points obtained by projecting these three-dimensional points G(curr), G(next), and G(prev) onto UV coordinates are defined as UV(curr), UV(next), and UV(prev), respectively. Below, we will explain an example of calculating the predicted value (predUV) when performing fine prediction on the vector corresponding to UV(curr).
[0539] Here, let gCurr, gNext, and gPrev be the three-dimensional vectors corresponding to the three-dimensional points G(curr), G(next), and G(prev), respectively, and let uvCurr, uvNext, and uvPrev be the two-dimensional vectors corresponding to the two-dimensional points UV(curr), UV(next), and UV(prev), respectively. Furthermore, let gNgP be the vector from gNext to gPrev, gNgC be the vector from gNext to gCurr, and uvNuvP be the vector from UV(next) to UV(prev). In addition, let dot(gNgP, gNgC) be the dot product of gNgP and gNgC, and let dot(gNgP, gNgP) be the dot product of two gNgPs.
[0540] Let G(proj) and UV(proj) be the projection points from a three-dimensional point G(curr) and a two-dimensional point UV(curr) onto their respective opposite sides. Vectors gProj and uvProj can be obtained by assuming that the triangle formed by the three-dimensional points G(curr), G(next), and G(prev) is similar to the triangle formed by the two-dimensional points UV(curr), UV(next), and UV(prev). Specifically, vectors gProj and uvProj are expressed as follows.
[0541] uvProj = uvNext + uvNuvP × (dot(gNgP, gNgC) / dot(gNgP, gNgP)) gProj = gNext + gNgP × (dot(gNgP, gNgC) / dot(gNgP, gNgP))
[0542] Furthermore, the dot product of the vectors from projection point G(proj) to G(curr) is dot(gCurr - gProj, gCurr - gProj). Here, let orthogonal(uvNuvP) be the vector orthogonal to uvNuvP.
[0543] In this case, the distance between projection points G(proj) and G(curr) corresponds to the distance between projection points UV(proj) and UV(curr), assuming that the triangles are similar. Therefore, the vector uvProjuvCurr from projection point UV(proj) to UV(curr) can be expressed as follows.
[0544] uvProjuvCurr = orthogonal(uvNuvP) * sqrt(dot(gCurr - gProj, gCurr - gProj) / dot(gNgP, gNgP))
[0545] However, sqrt is the square root.
[0546] Using these results, the prediction vector predUV can be obtained as follows.
[0547] predUV = uvProj + uvProjuvCur
[0548] Alternatively, the prediction vector can be denoted as predUV1 and calculated as follows.
[0549] predUV1 = uvProj - uvProjuvCur
[0550] Thus, by considering a prediction vector in the opposite direction to predUV, prediction accuracy can be improved and the amount of code can be reduced.
[0551] The above shows a specific example of calculating fine prediction, but for example, if the range of the vector components is extremely wide, setting a large number of bits in the fractional part to ensure calculation precision may cause an overflow. On the other hand, reducing the number of bits in the fractional part or performing calculations with integer precision to avoid overflow may worsen the calculation precision.
[0552] To solve this problem, the number of bits in the fractional part can be adjusted based on the magnitude of the dot product of vectors gNgP and gNgC, or the magnitude of the absolute value of the dot product.
[0553] The device may determine whether the magnitude of the absolute value of dot(gNgP, gNgC) exceeds the threshold TH and perform calculations according to the determination result. Figure 74A shows an example of calculating the predicted value when the magnitude of the absolute value of dot(gNgP, gNgC) exceeds the threshold TH according to this embodiment. Figure 74B shows an example of calculating the predicted value when the magnitude of the absolute value of dot(gNgP, gNgC) is less than or equal to the threshold TH according to this embodiment.
[0554] As a specific example, if the magnitude of the absolute value of dot(gNgP, gNgC) exceeds the threshold TH = (1 << (9 * 2)) - 1, the device may perform integer arithmetic without preserving the fractional part, as shown in Figure 74A.
[0555] On the other hand, if the magnitude of the absolute value of dot(gNgP, gNgC) is less than or equal to TH = (1 << (9 * 2)) - 1, the device may perform the calculation to allocate a predetermined number of bits (for example, 9 bits) for the fractional part, as shown in Figure 74B. This allows for lower calculation precision when the vector is large, and higher calculation precision when the vector is small. In this way, the device can broaden the range of vector values it can handle by adjusting the number of bits in the fractional part.
[0556] The above describes a method for switching the number of bits in the fractional part in two stages, but the device may also determine the number of bits in the fractional part according to the number of bits in the magnitude of the absolute value of dot(gNgP, gNgC). For example, the device may determine the number of bits in the fractional part using a base-2 logarithm based on the magnitude of the absolute value of dot(gNgP, gNgC). This allows the device to adjust the number of bits in the fractional part more accurately. Alternatively, the number of bits in the fractional part may be a fixed value of only one type. This reduces the processing load on the device.
[0557] Furthermore, the device may determine the number of bits in the fractional part based on a value other than the magnitude of the absolute value of dot(gNgP, gNgC). For example, it may use the magnitude of the vector itself as the basis for determination. This allows the device to reduce the amount of processing required.
[0558] In the example above, the number of bits in the fractional part is set to 9 bits, but it is not limited to being set to 9 bits. For example, the device may set the number of bits in the fractional part based on the number of bits in the hardware registers or arithmetic units. Alternatively, the number of bits may be set based on the performance of the software or system. Furthermore, the optimal number of bits may be set considering the trade-off between precision and processing power. In this way, the device can improve arithmetic precision and reduce the amount of code by setting an appropriate number of bits according to the execution environment.
[0559] Furthermore, the device may add a threshold value TH to the bitstream header or the like. By adding the threshold value TH used by the encoding device to the header and encoding, the decoding device can obtain the threshold value TH from the header and appropriately decode the bitstream using the same value as the number of fractional bits used by the encoding device.
[0560] Figure 75 shows another example of calculating the predicted value according to this embodiment.
[0561] After calculating the predicted values, the device may adjust the number of bits in the fractional part to 8 bits to ensure accuracy in the division operation, for example, when calculating the average of the fine predicted values, as shown in Figure 75. This adjustment may be performed regardless of the magnitude of the absolute value of dot(gNgP, gNgC). The device may also convert the average of the fine predicted values to integer precision, or adjust the number of bits in the fractional part. In this case, the device may perform rounding operations such as rounding to the nearest integer. Rounding operations may be performed in the same direction for positive and negative numbers, or in different directions. In some cases, the device may not perform rounding operations at all.
[0562] Figure 76 is a flowchart showing an example of the calculation process for predicted values according to this embodiment.
[0563] The device calculates the dot product dot(gNgP, gNgC) between vectors gNgP and gNgC (S1571).
[0564] The device determines whether the magnitude of the absolute value of the dot product dot(gNgP, gNgC) exceeds the threshold TH (S1572). If the magnitude of the absolute value of the dot product dot(gNgP, gNgC) exceeds the threshold TH (Yes in S1572), the device proceeds to step S1573; otherwise (No in S1572), it proceeds to step S1575.
[0565] The device calculates the predicted value (S1573). A specific example of how the predicted value is calculated is shown in Figures 72 and 73.
[0566] The device adjusts the fractional part of the predicted value to 8 bits (S1574). In other words, the device converts the fractional part of the predicted value to 8 bits.
[0567] The device adjusts the fractional part of the predicted value to 9 bits (S1575). In other words, the device converts the fractional part of the predicted value to 9 bits.
[0568] The device calculates the predicted value (S1576). A specific example of how the predicted value is calculated is shown in Figures 72 and 73.
[0569] The device adjusts the fractional part of the predicted value to 8 bits (S1577). In other words, the device converts the fractional part of the predicted value to 8 bits.
[0570] When step S1574 or S1577 is executed, the process of calculating the predicted value is completed.
[0571] Next, we will show a more specific example of how to calculate predicted values.
[0572] The device can calculate the following, using the dot product of vectors gNgP and gNgC as dot(gNgP, gNgC), and the x, y, and z components of each vector as .x, .y, and .z, respectively.
[0573] dot(gNgP, gNgC) = gNgP.x * gNgC.x + gNgP.y * gNgC.y + gNgP.z * gNgC.z
[0574] Similarly, the device can be calculated by defining the dot product of two vectors gNgP as dot(gNgP, gNgP) as follows.
[0575] dot(gNgP, gNgP) = gNgP.x * gNgP.x + gNgP.y * gNgP.y + gNgP.z * gNgP.z
[0576] Next, the device defines the magnitude of the absolute value of the dot product dot(gNgP, gNgC) as |dot(gNgP, gNgC)| and determines the number of bits in the fractional part representing the bit shift amount, shiftBits, as follows.
[0577] if (| dot(gNgP, gNgC) | > (1 << (9 * 2)) - 1) { shiftBits = 0} else { shiftBits = 9}}
[0578] The device can use this bit shift amount, shiftBits, to calculate the vector to the projection point as follows:
[0579] uvProj = (uvNext << shiftBits) + (uvNuvP * (dot(gNgP, gNgC) << shiftBits)) / dot(gNgP, gNgP) gProj = (gNext << shiftBits) + (gNgP * (dot(gNgP, gNgC) << shiftBits)) / dot(gNgP, gNgP)
[0580] Furthermore, the device defines the dot product of the vectors from projection point G(proj) to point G(curr) as dot(gCurr - gProj, gCurr - gProj) and calculates it as follows.
[0581] dot(gCurr - gProj, gCurr - gProj) = ((gCurr.x << shiftBits) - gProj.x) * ((gCurr.x << shiftBits) - gProj.x) + ((gCurr.y << shiftBits) - gProj.y) * ((gCurr.y << shiftBits) - gProj.y) + ((gCurr.z << shiftBits) - gProj.z) * ((gCurr.z << shiftBits) - gProj.z)
[0582] Furthermore, the device defines an orthogonal vector to the two-dimensional vector uvNuvP as orthogonal(uvNuvP), and calculates the vector uvProjuvCurr as follows.
[0583] uvProjuvCurr = ((orthogonal(uvNuvP) << shiftBits) * sqrt(dot(gCurr - gProj, gCurr - gProj))) / sqrt(dot(gNgP, gNgP) << (shiftBits * 2))
[0584] However, sqrt is the square root. The device may also use an approximate integer of the square root for sqrt. This allows operations to be performed using integer types.
[0585] The device can calculate candidate prediction vectors, predUV0 or predUV1, as follows:
[0586] predUV0 = uvProj + uvProjuvCurr
[0587] Furthermore, if the direction of uvProjuvCurr is reversed, it can be calculated as follows.
[0588] predUV1 = uvProj - uvProjuvCurr
[0589] If the device wants to set the number of bits in the fractional part to 8, it may perform the following steps.
[0590] if (shiftBits >= 8) { predUV0 >>= (shiftBits - 8) predUV1 >>= (shiftBits - 8)} else { predUV0 <<= (8 - shiftBits) predUV1 <<= (8 - shiftBits)}
[0591] Through the above process, the device can change the number of bits in the fractional part from shiftBits to 8 bits. Although this embodiment shows an example of changing to 8 bits, the device is not necessarily limited to this, and can also change the number of bits in the fractional part to any n bits (where n is a non-negative integer).
[0592] Furthermore, if the device changes the number of bits in the fractional part from shiftBits to n bits, the following process may be performed.
[0593] if (shiftBits >= n) { predUV0 >>= (shiftBits - n) predUV1 >>= (shiftBits - n)} else { predUV0 <<= (n - shiftBits) predUV1 <<= (n - shiftBits)}
[0594] The device can change the number of bits in the fractional part from shiftBits to any n bits through this process.
[0595] Either predUV0 or predUV1, which are candidate prediction vectors, is selected, and the prediction vector predUV is set as follows.
[0596] (1) If predUV0 is selected: predUV = predUV0 (2) If predUV1 is selected: predUV = predUV1
[0597] Furthermore, when calculating predUV from predUV0 or predUV1, you may add or subtract some offset value, for example, considering that the fractional part is 8 bits. An example is shown below.
[0598] predUV = (offset0 << 8) + predUV0 predUV = (offset1 << 8) - predUV1
[0599] This process adjusts the prediction vector using an offset.
[0600] Furthermore, several prediction vectors may be added together and the average value calculated. For example, the prediction vector obtained by cumulatively adding count prediction vectors may be named predUVcount, and then rounded and added together before averaging.
[0601] if(predUVcount.x >= 0) { rounding.x = 1 << (8 - 1)} else { rounding.x = -1 * (1 << (8 - 1))} if(predUVcount.y >= 0) { rounding.y = 1 << (8 - 1)} else { rounding.y = -1 * (1 << (8 - 1))} predUV = (predUVcount / count + rounding) >> 8
[0602] In this process, the direction of rounding and addition is adjusted by the sign of predUVcount. This improves the accuracy of the predicted vector while increasing the efficiency of encoding.
[0603] In the example above, the direction of rounding and addition is adjusted by the sign of predUVcount, but you may also use a method that rounds in the same direction, or a method that does not perform rounding and addition.
[0604] Furthermore, the above calculation formula "predUV = (predUVcount / count + rounding) >> 8" represents the following individual processing.
[0605] predUV.x = (predUVcount.x / count + rounding.x) >> 8 predUV.y = (predUVcount.y / count + rounding.y) >> 8
[0606] This ensures that each component is treated appropriately.
[0607] Next, we will explain an example of obtaining fine-predicted values in UV coordinates using information from three-dimensional space.
[0608] Figure 77 shows an example of a vector in three-dimensional space when making a fine prediction according to a modified example of this embodiment. Figure 78 shows an example of a vector in a two-dimensional plane when making a fine prediction according to a modified example of this embodiment.
[0609] As shown in Figure 77, let G(Next) and G(Prev) be two encoded three-dimensional points connected to a three-dimensional point G(Curr). Also, as shown in Figure 78, let UV(Next) and UV(Prev) be the points obtained by projecting these three-dimensional points G(curr), G(next), and G(prev) onto the UV coordinate system, respectively. Furthermore, let UV(Curr) be the point obtained by projecting G(Curr) onto the UV coordinate system. Here, we show an example of fine prediction of the vector corresponding to UV(Curr) (derivation of predUV).
[0610] Here, the three-dimensional vectors corresponding to known three-dimensional points G(Curr), G(Next), and G(Prev) are denoted as three-dimensional vectors gCurr, gNext, and gPrev, respectively, and the two-dimensional vectors corresponding to known two-dimensional points UV(Next) and UV(Prev) are denoted as two-dimensional vectors uvNext and uvPrev, respectively. Furthermore, the vector from three-dimensional vector gNext to three-dimensional vector gPrev is denoted as three-dimensional vector gNgP, the vector from three-dimensional vector gNext to three-dimensional vector gCurr is denoted as three-dimensional vector gNgC, and the vector from two-dimensional vector uvNext to two-dimensional vector uvPrev is denoted as two-dimensional vector uvNuvP.
[0611] Let dot(gNgP, gNgC) be the dot product of three-dimensional vector gNgP and three-dimensional vector gNgC, and dot(gNgP, gNgP) be the dot product of two three-dimensional vectors gNgP. Here, let G(Proj) be the projection point from three-dimensional point G(Curr) and two-dimensional point UV(Curr) onto their respective opposite edges.
[0612] The vector gProj for a three-dimensional point G(Proj) and the vector uvProj for a two-dimensional point UV(Proj) can be expressed using the ratio of the dot product as follows, assuming that the triangle formed by the three-dimensional points G(Curr), G(Next), and G(Prev) is similar to the triangle formed by the two-dimensional points UV(Curr), UV(Next), and UV(Prev).
[0613] uvProj = uvNext + uvNuvP * (dot(gNgP, gNgC) / dot(gNgP, gNgP)) gProj = gNext + gNgP * (dot(gNgP, gNgC) / dot(gNgP, gNgP))
[0614] Furthermore, let dot(gCurr - gProj, gCurr - gProj) be the dot product of the vectors from the projection point 3D point G(Proj) to the 3D point G(Curr). Let orthogonal(uvNuvP) be the vector orthogonal to the 2D vector uvNuvP. Similarly, by using similar triangles, the 2D vector uvProjuvCurr from the projection point 2D point UV(Proj) to the 2D point UV(Curr) can be expressed as follows.
[0615] uvProjuvCurr = orthogonal(uvNuvP) * sqrt(dot(gCurr - gProj, gCurr - gProj) / dot(gNgP, gNgP))
[0616] However, sqrt represents the square root.
[0617] Therefore, the two-dimensional prediction vector predUV can be calculated, for example, as follows.
[0618] predUV = uvProj + uvProjuvCur
[0619] Furthermore, depending on the direction of the two-dimensional vector uvProjuvCurr, the prediction vector can be set to predUV0 or predUV1 (the opposite direction), and calculated as follows.
[0620] predUV0 = uvProj + uvProjuvCur predUV1 = uvProj - uvProjuvCur
[0621] In this way, by considering predUV as a bidirectional prediction vector, prediction accuracy can be improved and the amount of code can be reduced.
[0622] Figure 79 shows another example of calculating the predicted value according to this embodiment.
[0623] In this embodiment, a method for fixedly securing the number of bits N in the fractional part in order to ensure integer arithmetic precision in the calculation of fine predictions in three-dimensional space is described. For example, as shown in Figure 79, the fractional part may be set to N bits. N bits may be set to, for example, N=1, N=2, N=3, ..., N=5, N=6, N=7, N=8. Alternatively, N=16 or N=0 (treated as an integer without a fractional part) may also be used.
[0624] The number of bits in the fractional part N may be changed depending on the processing of the fine prediction. For example, when calculating the square root, the number of bits in the fractional part is halved, so by doubling the number of bits in the fractional part to 2N bits before the calculation, the fractional part of the calculation result can be maintained at N bits.
[0625] Furthermore, to prevent value overflow, the number of bits N in the fractional part may be right-shifted to reduce it. Conversely, to prevent underflow, the number of bits N in the fractional part may be left-shifted beforehand to increase it. In addition, to improve the precision of calculations such as integer division, the number of bits in the fractional part can be left-shifted before the calculation to increase it, and then right-shifted after the calculation to decrease it. In this case, calculation errors can be suppressed by performing rounding. This makes it possible to handle a wide range of input values.
[0626] The number of bits N in the fractional part is not limited to the above-mentioned setting. For example, it can be determined according to the number of bits in the hardware registers or arithmetic units. It can also be set according to the performance of the software or system, or the optimal number of bits can be determined by considering the trade-off with arithmetic precision. In this way, by setting an appropriate number of bits according to the execution environment, arithmetic precision can be improved and the amount of code can be reduced.
[0627] The number of bits N in the fractional part may be added to the bitstream header, etc. By this method, the decoding device can obtain the number of bits N used by the encoding device from the header by including the number of bits N used by the encoding device in the header. As a result, the decoding device can perform the decoding process using the same number of bits in the fractional part as used by the encoding device, enabling proper decoding of the bitstream.
[0628] Furthermore, the number of bits N in the fractional part can be switched to a different value depending on the profile or level of the standard. For example, in profile A, which is intended for high image quality, a larger value (e.g., N=16) can be used to increase calculation precision, while in profile B, which is intended for low processing load, a smaller value (e.g., N=5) can be used to reduce the number of bits required. In this way, by defining the number of bits N in the fractional part according to the profile or level of the standard, the decoder can appropriately switch the number of bits N based on the profile or level information contained in the stream and correctly decode the bitstream.
[0629] Furthermore, the number of bits N in the fractional part can also be determined by the input value or conditions. For example, in the example of calculating the count-th fine prediction value, one could determine the number of bits N in the fractional part according to the number of bits in the magnitude of the absolute value of dot(gNgP, gNgC). More specifically, one could determine the number of bits N in the fractional part using the number of bits in dot(gNgP, gNgC) or the base-2 logarithm of the magnitude of the absolute value. This ensures an appropriate number of bits N in the fractional part according to the range of the input value, thereby improving calculation precision.
[0630] Figure 80 is a flowchart showing another example of the calculation process for predicted values according to this embodiment.
[0631] The device converts the fractional part of the predicted value to N bits (S1581). In other words, the device calculates the predicted value with fractional precision.
[0632] The device calculates the dot product value and dot product ratio of the vectors (S1582).
[0633] The device converts the fractional part to 2*N bits (S1583).
[0634] The device calculates the square root (S1584). In this way, the fractional part may be converted to 2*N bits in step S1583, before calculating the square root.
[0635] The device converts the fractional part to N bits (S1585). Step S1585 may be incorporated into step S1584. In other words, the device may convert the fractional part to N bits when calculating the square root.
[0636] The device calculates a predicted value (vector) (S1586).
[0637] As shown in this flowchart, when calculating the square root, the number of bits in the fractional part is also halved. Therefore, by doubling the number of bits N in the fractional part beforehand, it is possible to maintain the same precision as the number of bits N in the fractional part of the output.
[0638] Note that the input for the square root calculation process may be the original number of bits (N), or the fractional part may be converted to N bits after the output of the square root calculation process.
[0639] Next, we will show a specific example of how to calculate predicted values.
[0640] First, let dot(gNgP, gNgC) be the dot product of vectors gNgP and gNgC, and let .x, .y, and .z be the symbols representing their x, y, and z components, respectively. Then, we can calculate it as follows.
[0641] dot(gNgP, gNgC) = gNgP.x * gNgC.x + gNgP.y * gNgC.y + gNgP.z * gNgC.z
[0642] Similarly, the device can be calculated by defining the dot product of two vectors gNgP as dot(gNgP, gNgP) as follows.
[0643] dot(gNgP, gNgP) = gNgP.x * gNgP.x + gNgP.y * gNgP.y + gNgP.z * gNgP.z
[0644] Here, we will assume that the number of bits in the fractional part N = 8, and that shiftBits = N = 8. Note that shiftBits can also be set to values such as 5, 6, or 7.
[0645] Next, we will explain how to find the vector to the projection point. In this calculation, " / " signifies integer division, truncating the decimal part. The vector to the projection point is calculated using the following formula.
[0646] uvProj = (uvNext << shiftBits) + uvNuvP * ((dot(gNgP, gNgC) << shiftBits) / dot(gNgP, gNgP)) gProj = (gNext << shiftBits) + gNgP * ((dot(gNgP, gNgC) << shiftBits) / dot(gNgP, gNgP))
[0647] Next, the dot product of the vectors from projection point G(Proj) to G(Curr) is defined as dot(gCurr - gProj, gCurr - gProj), and is calculated as follows.
[0648] dot(gCurr - gProj, gCurr - gProj) = ((gCurr.x << shiftBits) - gProj.x) * ((gCurr.x << shiftBits) - gProj.x) + ((gCurr.y << shiftBits) - gProj.y) * ((gCurr.y << shiftBits) - gProj.y) + ((gCurr.z << shiftBits) - gProj.z) * ((gCurr.z << shiftBits) - gProj.z)
[0649] Furthermore, if we define orthogonal(uvNuvP) as a vector orthogonal to the two-dimensional vector uvNuvP, then the vector uvProjuvCurr can be expressed as follows.
[0650] uvProjuvCurr = orthogonal(uvNuvP) * sqrt(dot(gCurr - gProj, gCurr - gProj) / dot(gNgP, gNgP))
[0651] However, sqrt() refers to the square root operation.
[0652] Note that in this case, since dot(gCurr - gProj, gCurr - gProj) has 2N bits in the fractional part (2 * shiftBits), the input to sqrt() also has 2N bits in the fractional part (2 * shiftBits).
[0653] Therefore, the candidate prediction vector predUV0 or predUV1 can be obtained as follows.
[0654] predUV0 = uvProj + uvProjuvCurr predUV1 = uvProj - uvProjuvCurr (when the direction of uvProjuvCurr is reversed)
[0655] Either predUV0 or predUV1, obtained in this way, can be used as the prediction vector.
[0656] Either the prediction vector candidate predUV0 or predUV1 is selected, and the prediction vector predUV is set as follows:
[0657] (1) If predUV0 is selected: predUV = predUV0 (2) If predUV1 is selected: predUV = predUV1
[0658] Furthermore, when calculating predUV from predUV0 or predUV1, you may add or subtract some offset value, for example, considering that the decimal part is shiftBits. An example is shown below.
[0659] predUV = (offset0 << shiftBits) + predUV0 predUV = (offset1 << shiftBits) - predUV1
[0660] Furthermore, several prediction vectors may be added together and the average value calculated. For example, the prediction vector obtained by cumulatively adding count prediction vectors may be named predUVcount, and then rounded and added together before averaging.
[0661] if (predUVcount.x >= 0) { rounding.x = 1 << (shiftBits - 1);} else { rounding.x = -1 * (1 << (shiftBits - 1));} if (predUVcount.y >= 0) { rounding.y = 1 << (shiftBits - 1);} else { rounding.y = -1 * (1 << (shiftBits - 1));} predUV = (predUVcount / count + rounding) >> shiftBits;
[0662] In this way, rounding errors can be reduced by setting the direction of rounding to differ based on the sign of predUVcount. Alternatively, a method of rounding in the same direction may be used, or rounding addition may be omitted altogether.
[0663] Furthermore, the above formula predUV = (predUVcount / count + rounding) >> shiftBits indicates that the x and y components are processed separately, as follows:
[0664] predUV.x = (predUVcount.x / count + rounding.x) >> shiftBits predUV.y = (predUVcount.y / count + rounding.y) >> shiftBits
[0665] Next, we will show a concrete example of the square root calculation process (sqrt). Let a be the input value for the square root calculation, and the square root of a is expressed as follows.
[0666]
[0667] The square root calculation process may also be performed using Newton's method based on the following recurrence relation.
[0668]
[0669] After convergence, root = x n This is the result.
[0670] More specifically, it is calculated as follows:
[0671]
[0672] The device may also switch the number of iterations k in the recurrence relation to a different value depending on the standard profile or level. For example, in the case of profile A, which targets high image quality, the device can use a larger value for the number of iterations k in the recurrence relation to improve calculation accuracy. Specifically, a large value such as k = 5 could be used.
[0673] On the other hand, in the case of profile B, which targets low processing volume, the device can use a small value for the number of iterations k in the recurrence relation in order to reduce the processing volume. For example, a small value such as k = 2 can be used.
[0674] In this way, the device can correctly decode the bitstream by defining the value of the number of iterations k of the recurrence relation according to the profile or level of the standard, and by appropriately switching the number of iterations of the recurrence relation based on the profile or level information contained in the stream.
[0675] In addition, in a certain profile C, the device can set the number of iterations k in the recurrence relation to 1. In this case, the device can perform one iteration by right bit shifting, eliminating the need to perform division and allowing for the definition of a profile with reduced processing load.
[0676] Furthermore, the device may add the number of iterations k of the recurrence relation to the bitstream header or the like. This allows the device (encoder side) to encode by adding the number of iterations k used to the header. In addition, the decoder can obtain the number of iterations k from the header and perform the decoding process using the same value as the number of iterations used by the encoding device, thereby appropriately decoding the bitstream.
[0677] In the calculations for Math 3, "≪" indicates a left shift, "≫" indicates a right shift, " / " indicates integer division that truncates the decimal part, "++" indicates an increment, and "≫=" indicates assigning the shifted value.
[0678] Initial value x of the recurrence relation 0 x0 can be calculated using the bit count value bits(a) of the input value a as an approximation of the true square root. Specifically, x0 can be set to x0 = 2 ((bits(a) ≫ 1)). For example, if the input value a = 8, the bit count value bits(a) = 4, and the initial value x 0 x0 can be set to 4. Also, if the input value a = 40, the bit count value bits(a) = 6, and the initial value x 0 x0 can be set to 8.
[0679] In this way, approximate initial values can be easily calculated without performing complex operations including lookup tables or division. Furthermore, by using an approximate value of the true square root as the initial value, it is possible to obtain a highly accurate square root while reducing the number of iterations required for the recurrence relation to converge. Moreover, by reducing the number of iterations, the square root can be calculated quickly.
[0680] For example, by using an approximate value of the true square root as the initial value, it is possible to find the square root with high accuracy even with only two iterations of the recurrence relation, thereby achieving reduced processing load.
[0681] Also, the initial value x of the recurrence relation 0By setting x0 to a power of 2, for example, x0 = 2 ((bits(a)≫1)), the division operation in the first iteration of the recurrence relation can be performed by a right shift. This reduces the number of computationally expensive division operations.
[0682] For example, if the number of iterations of a recurrence relation is set to two, the initial value of the recurrence relation can be set to a power of two. This allows the division operation in the first iteration to be performed by a right shift, and the division operation to be performed in the second iteration. In this way, the number of division operations can be reduced by one.
[0683] Since the input value 'a' has a 2*N bit fractional part, sufficient precision can be obtained by performing integer division directly in the second iteration of the recurrence relation. Therefore, there is no need to separately reserve the fractional part. This reduces the amount of processing required.
[0684] If high precision in finding the square root is not required, the first iteration using the recurrence relation can be terminated, and the value representing the square root can be set to root = x1. On the other hand, if even higher precision in finding the square root is required, the iteration process can be performed k times, and the subsequent iterations using the recurrence relation can be performed (k - 1) times, with the value representing the square root being set to root = x n That is also acceptable.
[0685] In this case, the subsequent iterations using the recurrence relation may be repeated as follows:
[0686]
[0687] Alternatively, the initial iteration using a recurrence relation with shifts may be omitted, and the shifts may be included in subsequent iterations, ensuring that all iterations are calculated in the same way.
[0688] Next, we will show another specific example of the square root calculation process (sqrt). Let a be the input value for the square root calculation, and the square root of a is expressed as follows.
[0689]
[0690] The square root calculation process may be calculated by Newton's method using the following recurrence formula for obtaining the reciprocal of the square root.
[0691]
[0692] After convergence, root = 1 / x n will be the case.
[0693] More specifically, it is calculated as follows.
[0694]
[0695] In the above process, "<<" means left shift, ">>" means right shift, "++" means increment, and ">>=" means substituting the value after the shift.
[0696] The initial value x of the recurrence formula 0 may be calculated using the bit count value bits(a) of the input value a as an approximation of the reciprocal of the square root which is the true value. In this case, x 0 may be calculated as x0 = 2 - ((bits(a) >> 1)).
[0697] Thereby, the device can easily calculate an approximation value without performing complex operations including a lookup table and division. Also, by using the approximation value of the reciprocal of the square root which is the true value as the initial value, the device can obtain a high-precision square root while suppressing the number of iteration processes until the recurrence formula converges. Furthermore, by suppressing the number of iteration processes, the device can obtain the square root at high speed.
[0698] For example, by using the approximation value of the reciprocal of the square root which is the true value as the initial value, the device can obtain a high-precision square root even when the number of iterations of the recurrence formula is set to two, and realize low processing quantization.
[0699] Also, since x 0 is a power-of-two value, the device can perform the multiplication process by right shift in the first iterative process by the recurrence formula. Thus, it is not necessary to multiply a value having a fractional part, and it is possible to easily avoid an overflow due to insufficient bits. Also, since it can be realized by a shift operation, the implementation of the device becomes easy.
[0700] Next, we will describe in more detail other specific examples of the square root calculation process (sqrt), including its effects. Figures 81 to 85 are diagrams illustrating other specific examples of the square root calculation process (sqrt) according to this embodiment.
[0701] As shown in Figure 81, the input 'a' for calculating the square root is assumed to have a maximum total of 48 bits, with an integer part of 32 bits and a fractional part of 8*2=16 bits (where N=8).
[0702] By the way, the recurrence relation in the first iteration is shown in equation 6. If the initial value x0 = 2 - ((bits(a)≫1)), then input a with a large number of bits and x which contains a large fractional part. 0 Multiplication by x is necessary. 0 Since it needs to be multiplied three times, a considerable number of bits will be required.
[0703] Here, we add a new fractional part of (N × 2 + (bits(a)≫1)) bits to input a, which is (16 + (bits(a)≫1)) bits when N = 8. In this case, the (N × 2) part is a fixed-point part that does not depend on the number of bits in input a, and the (bits(a)≫1) part is a variable-point part that depends on the number of bits in input a (Figure 82).
[0704] As described above, simply adding the fractional part requires more than 80 bits, which could cause an overflow in systems with, for example, 64-bit registers or arithmetic units. Therefore, x 0 By utilizing the fact that is a negative power of 2, we add a variable fraction (left shift) and x 0 The multiplication with is treated as truncation by a right shift of (bits(a)≫1). Here, the addition of the variable fraction is performed as a left shift operation, x 0 By treating the reduction in the number of bits resulting from multiplication as a right shift, the left shift and right shift cancel each other out. In this way, the process of adding the variable-point part with a left shift and x 0 By simultaneously performing a right shift operation associated with multiplication, the number of bits can be prevented from increasing. Therefore, the device is x 0This can be performed by adding a fixed-point part to input a using a left shift operation (Figure 83). This process allows for the addition of a necessary and sufficient fractional part for square root calculation while avoiding overflow in a 64-bit system.
[0705] Next, in the recurrence relation for the first iteration, the device is x 0 The multiplication operation can be performed by a right shift of (bits(a)≫1) bits. This is represented as shown in Figures 84 and 85.
[0706] This process ensures that even after the first iteration is complete, x 1 This can maintain 16-bit precision.
[0707] Next, we will explain in more detail the second iteration of the square root calculation process (sqrt), including its effects. Figures 86 to 93 are diagrams illustrating other specific examples of the square root calculation process (sqrt) according to this embodiment.
[0708] First, a and x 1 When the operation is performed, the result becomes 64 bits, as shown in Figure 86. However, after this, a * x 1 2 Performing the operation may result in an overflow. Therefore, (N × 2 = 16) bits are truncated by a right shift operation as shown in Figure 87, and then a * x 1 2 By performing the calculation (Figure 88), overflow can be prevented.
[0709] Next, a・x 1 2 Performing the operation results in 64 bits again. However, after this, x1 × (3 - a・x1 2 Performing the operation ) may result in an overflow. Therefore, (bits(a)≫1) bits are truncated by a right shift operation as shown in Figure 89, and then x1 × (3 - a・x1 2The operation is performed as shown in Figure 90. Note that since (N × 2 = 16) bits have already been truncated by a right shift operation, only (bits(a) ≫ 1) bits are truncated by the right shift operation in this process. This makes it possible to perform accurate calculations while preventing overflow without reducing the number of bits more than necessary.
[0710] Furthermore, x1 × (3 - a・x1 2 Performing the operation ) results in a 56-bit value. Finally, the reciprocal of the square root x 2 Multiplying by the input a (48 bits) to calculate the square root (root) may cause an overflow. Therefore, as shown in Figure 91, it is necessary to truncate (16 + (bits(a)≫1)) bits by a right shift operation. Note that in this process, (x1 / 2) × (3 - a・x1 2 To calculate this, (16 + (bits(a)≫1) + 1) bits are truncated by a right shift.
[0711] The second iteration is complete, and the reciprocal of the converged square root x is obtained. 2 Once this is found, in order to calculate the square root, as shown in Figure 92, the reciprocal x 2 Multiply this by input a.
[0712] Finally, as shown in Figure 93, the final square root output, root, can be obtained by truncating the fractional part (16 + (bits(a)≫1)) bits added during the square root operation using a right shift operation.
[0713] By performing this process, the square root can be calculated without using computationally expensive division. Furthermore, even with a wide input range and large input values, the square root can be calculated without causing an overflow, even with registers and arithmetic units of a limited number of bits.
[0714] In this explanation, we have shown an example where N=8 and a 16-bit fixed-point part is added. However, N=7, N=6, N=5, etc., and the fixed-point part can also be 14 bits, 12 bits, 10 bits, or any other number of bits. Furthermore, for systems with registers or arithmetic units exceeding 64 bits, such as 128 bits, N=9 (18 bits for the fixed-point part) or more is also acceptable. The variable-point part is not limited to (bits(a)≫1) bits, but can also have any other number of bits.
[0715] Furthermore, if the input 'a' has a small number of bits and the multiplication does not overflow even after repeated multiplications, the number of shift operations can be reduced by performing a shift only at the end of the process, rather than performing a right shift after each multiplication. In this case, the fractional part added by the left shift can be reduced or omitted altogether.
[0716] Furthermore, the lower bits of input a, or the fractional part originally present in input a, may be used as part or all of the fixed-point part, as shown in other specific examples of square root calculation (sqrt). This allows processing to be performed with fewer bit registers and arithmetic units, and reduces the number of shift operations.
[0717] If a highly accurate square root cannot be obtained, the process can be terminated after the first iteration using the recurrence relation, and the square root can be calculated as root=(a・x1)≫(16+(bits(a)≫1)). This allows the square root to be calculated without additional multiplication, thus reducing the amount of processing required.
[0718] Conversely, to obtain a square root with even higher accuracy, the iteration process can be performed k times, and the second and subsequent iterations using the recurrence relation can be performed (k-1) times, with the final square root being root=(a・xn)≫(16+(bits(a)≫1)). In that case, a highly accurate square root can be obtained by repeating the second iteration using the recurrence relation in other specific examples of square root calculation (sqrt) as follows.
[0719]
[0720] Furthermore, the number of iterations k in the recurrence relation may be switched to a different value depending on the standard profile or level. For example, in the case of profile A, which targets high image quality, the device may use a large value for the number of iterations in the recurrence relation to improve calculation accuracy. Specifically, a large value such as k=5 could be used.
[0721] On the other hand, in the case of profile B, which targets low processing loads, the device may use a small value for the number of iterations in the recurrence relation to reduce the processing load. For example, a small value such as k=2 may be used. In this way, by specifying the number of iterations in the recurrence relation according to the profile or level of the standard, the decoding device can correctly decode the bitstream by selecting an appropriate number of iterations based on the profile or level information contained in the bitstream.
[0722] In addition, in a certain profile C, the number of iterations k=1 may be used. In this case, one iteration can be performed by a right bit shift, reducing the amount of processing without performing multiplication.
[0723] Furthermore, the number of iterations k in the recurrence relation may be added to the bitstream header or similar. This allows the decoder to obtain the number of iterations k used by the encoder from the header by adding it to the header during encoding. The decoder can then properly decode the bitstream by using the same value as the number of iterations used by the encoder to perform the decoding process.
[0724] Furthermore, the specific examples of the square root calculation process (sqrt) described in this embodiment, as well as other specific examples, may be applied not only to the calculation of fine prediction values, but also to other processes that require the calculation of square roots. In such cases, as shown in these examples of square root calculation, by setting the initial value of the recurrence relation to a power of 2, it becomes possible to perform the division or multiplication operation in the first iteration of the recurrence relation using a shift operation, thereby reducing the amount of processing required.
[0725] Figure 94 is a diagram showing another example of the configuration of the encoding device according to this embodiment. Figure 95 is a flowchart showing another example of the encoding method by the encoding device according to this embodiment.
[0726] The encoding device 1550 includes a circuit 1551 and a memory 1552 connected to the circuit 1551. The encoding device 1550 may perform the processes described in Figures 71 to 93.
[0727] Circuit 1551 performs the following operations.
[0728] Circuit 1551 acquires multiple vertices included in the three-dimensional mesh (S1561). Circuit 1551 predictively encodes one of the multiple vertices (S1562). In predictive encoding, a first predicted value for one vertex is calculated with fractional precision, a second predicted value is calculated by reducing the number of bits in the fractional part of the first predicted value, and the vertex is predictively encoded using the second predicted value.
[0729] According to this method, by reducing the number of bits in the fractional part of the predicted value calculated with fractional precision, it is possible to reduce the amount of data while controlling the precision during calculations. Since the amount of encoding processing is reduced by reducing the number of bits, processing efficiency can be improved.
[0730] For example, in calculating the first predicted value, two or more third predicted values are calculated with decimal precision based on each of the two or more processed vertices, and the first predicted value is calculated by combining these two or more third predicted values.
[0731] According to this method, integrating multiple predicted values can yield highly accurate predictions. This reduces prediction errors and improves coding efficiency. Furthermore, integrating multiple predicted values allows for greater flexibility in the method for selecting the appropriate prediction.
[0732] For example, in calculating the first predicted value, two or more third predicted values are combined with decimal precision.
[0733] According to this method, by performing the integration process with decimal precision and then converting it to an integer later, encoding can be performed while maintaining the precision of the integration process. By performing high-precision integration and then converting to an integer, the increase in prediction error can be suppressed.
[0734] For example, two or more third predicted values and the first predicted value have a predetermined number of bits in their fractional part. In calculating the second predicted value, the second predicted value is calculated by reducing the predetermined number of bits in the fractional part of the first predicted value.
[0735] According to this method, the integer conversion process can be simplified by reducing the fractional part by a predetermined number of bits. Reducing the number of bits reduces the amount of code, thereby improving the efficiency of arithmetic processing.
[0736] For example, the first predicted value is the average of two or more third predicted values.
[0737] According to this method, using the average of multiple predicted values can reduce the overall prediction error. Averaging can potentially cancel out noise components and prediction errors, thus improving coding accuracy.
[0738] For example, in predictive coding, the predetermined number of bits is not reduced until the calculation of the second predicted value.
[0739] According to this method, the accuracy of the predicted value can be maintained by not reducing the number of bits before converting to an integer. Since no loss of precision occurs during calculation, high-precision encoding can be achieved.
[0740] For example, in predictive coding of a single vertex using a second predicted value, the second predicted value is rounded to calculate a fourth predicted value, and this fourth predicted value is then used to predictively code the single vertex.
[0741] According to this method, rounding can reduce errors in decimal places. This improves the accuracy of the encoded data and minimizes errors when reducing the number of bits.
[0742] For example, in rounding, the second predicted value is rounded in different directions depending on whether it is a positive or negative number.
[0743] According to this method, the coding accuracy can be improved by switching the rounding direction depending on whether the predicted value is positive or negative. In particular, it can achieve optimal rounding processing according to the magnitude and sign of the predicted value.
[0744] Figure 96 is a diagram showing another example of the configuration of the decoding device according to this embodiment. Figure 97 is a flowchart showing another example of the decoding method by the decoding device according to this embodiment.
[0745] The decoding device 1560 includes a circuit 1561 and a memory 1562 connected to the circuit 1561. The decoding device 1560 may perform the processes described in Figures 71 to 93.
[0746] Circuit 1561 performs the following operations.
[0747] Circuit 1561 acquires multiple vertices included in the three-dimensional mesh (S1571). Circuit 1561 predicts and decodes one of the multiple vertices (S1572). In predictive decoding, a first predicted value for one vertex is calculated with fractional precision, a second predicted value is calculated by reducing the number of bits in the fractional part of the first predicted value, and the second predicted value is used to predict and decode one vertex.
[0748] According to this method, by reducing the number of bits in the fractional part of the predicted value calculated with fractional precision, it is possible to reduce the amount of data while controlling the precision during calculations. Since the amount of processing required for decoding is reduced by reducing the number of bits, processing efficiency can be improved.
[0749] For example, in calculating the first predicted value, two or more third predicted values are calculated with decimal precision based on each of the two or more processed vertices, and the first predicted value is calculated by combining these two or more third predicted values.
[0750] According to this method, by integrating multiple predicted values, a more accurate predicted value can be obtained. This reduces the decoding error and improves the decoding accuracy.
[0751] For example, in calculating the first predicted value, two or more third predicted values are combined with decimal precision.
[0752] According to this method, more accurate prediction values can be calculated by integrating the results using decimal precision. Later, by converting the results to integers, the decoding process can be performed efficiently.
[0753] For example, two or more third predicted values and the first predicted value have a predetermined number of bits in their fractional part. In calculating the second predicted value, the second predicted value is calculated by reducing the predetermined number of bits in the fractional part of the first predicted value.
[0754] According to this method, reducing the number of bits in the fractional part and converting it to an integer reduces the burden of the decoding process while enabling the use of highly accurate predicted values.
[0755] For example, the first predicted value is the average of two or more third predicted values.
[0756] According to this method, averaging multiple predicted values can reduce prediction errors. Averaging suppresses the variability of predicted values, enabling highly accurate decoding.
[0757] For example, in predictive decoding, the predetermined number of bits is not reduced until the calculation of the second predicted value.
[0758] According to this method, since no reduction in precision due to bit reduction occurs until the data is converted to an integer during the decoding process, high-precision decoding can be maintained.
[0759] For example, in predicting and decoding a single vertex using the second predicted value, the second predicted value is rounded to obtain a fourth predicted value, and the fourth predicted value is used to predict and decode the single vertex.
[0760] According to this method, applying rounding minimizes errors during bit reduction, allowing for highly accurate predictions in the decoding process.
[0761] For example, in rounding, the second predicted value is rounded in different directions depending on whether it is a positive or negative number.
[0762] According to this method, the accuracy of the predicted value can be improved by adjusting the direction of rounding depending on whether the predicted value is positive or negative. Furthermore, the occurrence of rounding errors can be suppressed.
[0763] Furthermore, not all of the components described in this embodiment are always necessary, and only some of the components of the encoding and decoding devices of this disclosure may be included.
[0764] (Embodiment 2) Figure 98 is a block diagram showing the configuration of a three-dimensional data encoding device according to this embodiment. Note that in Figure 98, the encoding unit that encodes position information, which is included in the three-dimensional data encoding device, is not shown.
[0765] The three-dimensional data encoding device 2400 comprises a conversion unit 2410 and an encoding unit 2420.
[0766] The conversion unit 2410 performs a conversion process on the input attribute information before inputting it to the encoding unit 2420. The attribute information is, for example, the projected coordinates of three-dimensional coordinates onto the UV plane (video texture coordinates) as shown in the above embodiment. The conversion process is, for example, at least one of offset (offset processing) and scaling (scaling processing) described later.
[0767] The conversion unit 2410 includes a scale unit 2411 and an offset unit 2412. However, the conversion unit 2410 only needs to have at least one of the offset unit 2412 and the scale unit 2411. For example, if the conversion unit 2410 only performs an offset on the attribute information, the scale unit 2411 may not be necessary.
[0768] The scaling unit 2411 performs scaling (multiplication or division), which is an example of a conversion process, on the input attribute information and outputs a scale value (more specifically, scale information, which is information indicating the scale value, which is the value used for scaling).
[0769] The offset unit 2412 performs an offset (subtraction), which is another example of the conversion process, on the scaled attribute information and outputs an offset value (more specifically, offset information, which is information indicating the offset value, which is the value used for the offset).
[0770] The encoding unit 2420 encodes the attribute information (post-converted attribute information) converted by the conversion unit 2410, and also encodes the conversion information such as offset values or scale values as additional information (metadata).
[0771] The encoding unit 2420 includes an attribute information encoding unit 2421 and an additional information encoding unit 2422.
[0772] The attribute information encoding unit 2421 encodes the converted attribute information, which is the attribute information converted by the conversion unit 2410.
[0773] The additional information encoding unit 2422 encodes additional information, including conversion information such as scale values and offset values output by the conversion unit 2410.
[0774] For example, if the encoding unit 2420 does not support encoding negative values, or if it is specified that the encoding unit 2420 does not support encoding negative values, the conversion unit 2410 adds an offset value to the attribute information and converts the attribute information to a positive value when the format of the input attribute information has a negative value.
[0775] For example, if the encoding unit 2420 does not support decimals and floating-point numbers but supports integers, or if the encoding unit 2420 is specified to support integers but not decimals and floating-point numbers, the scale unit 2411 multiplies the input attribute information (more specifically, the numerical value indicated by the input attribute information) by a scale value to convert the attribute information into a positive numerical value when the format of the input attribute information is not an integer.
[0776] For example, if the encoding unit 2420 corresponds to encoding attribute information of a 12-bit unsigned integer type (positive integer), and the input attribute information is a 32-bit signed floating-point number in the range [-1, 1], then first, the attribute information (input_attribute) is converted to scaled_attribute, which is a 12-bit signed integer value in the range [-2048, 2047], through processes such as scaling, rounding, truncation, and rounding up.
[0777] Note that scaled_attribute = round(input_attribute × scale).
[0778] Here, `scale` is an example of a scale value, which is multiplied by the value indicated by the attribute information, for example, 2^(12 bits - 1 bit), that is, 2 to the power of 11 = 2048.
[0779] Next, the scaled attribute information is converted by the offset to a 12-bit unsigned integer type in the range [0, 4095].
[0780] Furthermore, offset_attribute = scaled_attribute - offset.
[0781] Here, offset is an example of an offset value, which is the value to subtract from the value indicated by the attribute information, for example, -1 * 2^(12 bits - 1 bit) - 1 = -2048.
[0782] The conversion information used for the conversion, namely the offset value and / or scale value, is input to the encoding unit 2420 and encoded as additional information.
[0783] The additional information encoding unit 2422 may encode the converted information as additional information as is, or it may encode information from which the converted information can be derived as additional information.
[0784] Furthermore, the information from which the offset value can be derived and the information from which the scale value can be derived may be shown independently, or they may be shown using common information.
[0785] For example, in the above example, the offset value and scale value are predetermined as offset = -1 * 2^(N-1) and scale = 2^(N-1). In this case, for example, the encoding unit 2420 stores the value of N (an integer greater than or equal to 1) in the additional information and encodes it, that is, it encodes additional information that indicates the value of N as conversion information.
[0786] Furthermore, if the encoding unit 2420 corresponds to a 12-bit unsigned integer type, N may be defined as the number of bits in the unsigned integer type, and N = 12.
[0787] Alternatively, N may be predetermined to represent the number of bits in an unsigned integer. In such a case, when information indicating that N is the number of bits in an unsigned integer is stored in the encoded stream (bitstream), it is not necessary to include information indicating that N is the number of bits in an unsigned integer as additional information.
[0788] Furthermore, the conversion unit 2410 may determine offset and scale (offset value and scale value) based on the values and characteristics indicated by the attribute information constituting the three-dimensional data.
[0789] The above example illustrates how the scale value is calculated by multiplication and the offset value by subtraction, but this is not the only way. The scale value may also be calculated by division, and the offset value may also be calculated by addition.
[0790] Furthermore, the scaling unit 2411 may round the values of the scaled attribute information after scaling by processes such as rounding, truncation, and rounding up.
[0791] Furthermore, the conversion unit 2410 does not need to convert the attribute information if the value indicated by the attribute information is not a positive integer. In this case, the conversion unit 2410 does not need to output the scale value and offset value, and may output information indicating that it was not converted as conversion information. Also in this case, for example, the encoding unit 2420 encodes the attribute information that has not been converted by the conversion unit 2410.
[0792] Figure 99 is a block diagram showing the configuration of a three-dimensional data decoding device according to this embodiment. Note that in Figure 99, the decoding unit that decodes the encoded position information, which is included in the three-dimensional data decoding device, is not shown.
[0793] The three-dimensional data decoding device 2430 comprises a decoding unit 2440 and an inverse conversion unit 2450.
[0794] The decoding unit 2440 receives encoded attribute information (encoded attribute information) and encoded additional information (encoded additional information) as input and decodes the encoded attribute information and encoded additional information. The decoding unit 2440 includes an attribute information decoding unit 2441 and an additional information decoding unit 2442.
[0795] The attribute information decoding unit 2441 generates decoded attribute information by decoding the encoded attribute information.
[0796] The additional information decoding unit 2442 decodes the encoded additional information to extract conversion information such as offset values and scale values.
[0797] The inverse conversion unit 2450 performs an inverse conversion process on the decoded attribute information based on the conversion information. The inverse conversion process is at least one of the inverse offset (inverse offset processing) and inverse scaling (inverse scaling processing) described later. The inverse conversion unit 2450 has an inverse offset unit 2451 and an inverse scaling unit 2452.
[0798] The reverse offset unit 2451 reverse offsets the decoded attribute information using the offset value extracted from the converted information, which is an example of reverse conversion processing. In other words, the reverse offset unit 2451 performs the reverse conversion on the decoded attribute information compared to the conversion performed on the attribute information by the conversion unit 2410 (more specifically, the offset unit 2412). For example, if the conversion unit 2410 adds the offset value to the value indicated by the attribute information, the reverse offset unit 2451 subtracts the offset value from the value indicated by the decoded attribute information.
[0799] The inverse scaling unit 2452 inversely scales the inversely offset decoded attribute information using the scale value extracted from the additional information, which is another example of the inverse transformation process. In other words, the inverse scaling unit 2452 performs the reverse transformation on the decoded attribute information compared to the transformation performed on the attribute information by the transformation unit 2410 (more specifically, the scaling unit 2411). For example, if the transformation unit 2410 multiplied the value indicated by the attribute information by the scale value, the inverse scaling unit 2452 divides the value indicated by the decoded attribute information by the scale value.
[0800] For example, if the offset value extracted from the additional information is offset and the scale value is scale, then the attribute information with the inverse offset is derived as offset_attribute = decoded_value + offset, and the attribute information with the inverse scale is derived as scaled_attribute = offset_attribute / scale.
[0801] Furthermore, in scaling and descaling, the processing load may be reduced by using shift operations (bit shifts) instead of multiplication and / or division, by expressing the scale value as a power of two, etc. In other words, scaling and descaling are processes that perform at least one of multiplication, division, and shift operations on the value indicated by the attribute information.
[0802] With the above configuration, the inverse conversion unit 2450 of the three-dimensional data decoding device 2430 can perform an inverse conversion process based on the conversion information contained in the encoded data, thereby reproducing the attribute information before it was converted by the conversion unit 2410 of the three-dimensional data encoding device 2400.
[0803] Furthermore, the three-dimensional data decoding device 2430 does not necessarily have to perform inverse transformation processing; it may choose whether or not to perform inverse transformation processing based on the application or use case.
[0804] The above examples illustrate the use of addition for inverse offset and division for inverse scaling, but this is not the only example. Subtraction may be used for inverse offset and multiplication for inverse scaling.
[0805] Furthermore, while the conversion unit 2410 is configured such that the offset unit 2412 is located after the scale unit 2411 (in a later stage), and the inverse conversion unit 2450 is configured such that the inverse scale unit 2452 is located after the inverse offset unit 2451, the configuration is not limited to these. For example, the conversion unit 2410 may be configured such that the scale unit 2411 is located after the offset unit 2412, and the inverse conversion unit 2450 may be configured such that the inverse offset unit 2451 is located after the inverse scale unit 2452.
[0806] Alternatively, the three-dimensional data encoding device 2400 may, based on the type of attribute information (attribute_type), select which configuration to use, that is, the order in which to perform scaling and offsetting on the attribute information, and store information indicating which configuration was used, that is, information indicating the order in which scaling and offsetting were performed (order information), as additional information, for example as a flag, and transmit it to the three-dimensional data decoding device 2430. The three-dimensional data decoding device 2430 may, based on the order information, select the order in which to perform inverse scaling and inverse offsetting, and perform inverse conversion processing on the decoded attribute information in the selected order.
[0807] The scale section 2411 and offset section 2412 of the conversion section 2410 and the inverse scale section 2452 and inverse offset section 2451 of the inverse conversion section 2450 may be arranged in different order, for example, as follows.
[0808] [Specific example of sequence 1] In sequence 1, the conversion unit 2410 performs an offset after scaling, and the inverse conversion unit 2450 performs inverse scaling after inverse offset.
[0809] Here, let org be the value indicated by the attribute information input to the encoding device, conv be the value indicated by the converted attribute information output from the conversion unit 2410, scale be the scale value, and offset be the offset value. At this time, the processing performed by the conversion unit 2410 can be expressed, for example, by the following equation A1. However, round() indicates the process of rounding a real number to the nearest integer value.
[0810] conv = round(org × scale) - offset (Equation A1)
[0811] Furthermore, if we denote the value indicated by the decoded attribute information obtained by decoding the encoded attribute information input to the three-dimensional data decoding device 2430 as conv, the value indicated by the attribute information output from the inverse transform unit 2450 as org, the scale value used for inverse scaling as 1 / scale, and the offset value used for inverse offset as -offset, then the processing performed by the inverse transform unit 2450 can be expressed, for example, by the following equation A2.
[0812] org = (conv + offset) / scale (Equation A2)
[0813] [Specific example of sequence 2] In sequence 2, scaling is performed after the offset in the conversion unit 2410, and the reverse offset is performed after the reverse scaling in the reverse conversion unit 2450.
[0814] If we let org be the value indicated by the attribute information input to the encoding device, conv be the value indicated by the converted attribute information output from the conversion unit 2410, scale be the scale value, and offset be offset, then the processing performed by the conversion unit 2410 can be expressed, for example, by the following equation A3. However, round() indicates the process of rounding a real number to the nearest integer value.
[0815] conv = round((org - offset) × scale) (Equation A3)
[0816] Furthermore, if we denote the value indicated by the decoded attribute information as conv, the value indicated by the attribute information output from the inverse transform unit 2450 as org, the scale value used for inverse scaling as 1 / scale, and the offset value used for inverse offset as -offset, then the processing performed by the inverse transform unit 2450 can be represented, for example, by the following equation A4.
[0817] org = conv / scale + offset (Equation A4)
[0818] Thus, even if the order of the scale unit 2411 and offset unit 2412 is switched with that of the inverse scale unit 2452 and inverse offset unit 2451, the attribute information before it was input to the three-dimensional data encoding device 2400 can be restored by setting appropriate scale and offset values.
[0819] Figure 100 is a flowchart showing the processing procedure of the three-dimensional data encoding device according to this embodiment.
[0820] First, the three-dimensional data encoding device 2400 determines whether or not to convert the input attribute information (S2401).
[0821] If the three-dimensional data encoding device 2400 determines that the input attribute information should be converted (Yes in S2401), it executes a conversion process on the input attribute information (S2402). For example, the three-dimensional data encoding device 2400 performs an offset and a scale on the input attribute information.
[0822] Next, the three-dimensional data encoding device 2400 stores the conversion information in the additional information and sets transform_flag=1 (S2403). For example, the additional information includes information indicating the offset value used for the offset and the scale value used for the scale as conversion information.
[0823] Next, the three-dimensional data encoding device 2400 encodes the additional information including the conversion information and the attribute information on which the conversion process has been performed (S2404). After step S2404, for example, the three-dimensional data encoding device 2400 generates a bitstream containing this encoded information as encoded data and transmits it to the three-dimensional data decoding device 2430.
[0824] On the other hand, if the three-dimensional data encoding device 2400 determines that the input attribute information does not need to be converted (No in S2401), it does not perform the conversion process on the input attribute information, does not store the converted information in the additional information, and sets transform_flag = 0 (S2405).
[0825] Next, the three-dimensional data encoding device 2400 encodes the additional information that does not contain the conversion information and the attribute information that has not undergone conversion processing, i.e., the input attribute information (S2406). After step S2406, for example, the three-dimensional data encoding device 2400 generates a bitstream containing this encoded information as encoded data and transmits it to the three-dimensional data decoding device 2430.
[0826] Figure 101 is a flowchart showing the processing procedure of the three-dimensional data decoding device according to this embodiment.
[0827] First, the three-dimensional data decoding device 2430 receives, for example, a bitstream transmitted by the three-dimensional data encoding device 2400, decodes the encoded data contained in the received bitstream, and analyzes the additional information contained in the decoded encoded data (S2411).
[0828] Next, the three-dimensional data decoding device 2430 decodes the attribute information of the encoded data contained in the bitstream (S2412).
[0829] Next, the three-dimensional data decoding device 2430 determines whether the transform_flag included in the bitstream is set to 1 (S2413). In other words, by determining whether transform_flag = 1, the three-dimensional data decoding device 2430 determines whether the attribute information included in the bitstream has been transformed.
[0830] If the three-dimensional data decoding device 2430 determines that the transform_flag included in the bitstream is set to 1 (Yes in S2413), it extracts the transformation information from the additional information and performs an inverse transformation process to attribute information based on the extracted transformation information (S2414).
[0831] On the other hand, if the three-dimensional data decoding device 2430 determines that the transform_flag included in the bitstream is not set to 1 (No in S2413), that is, if transform_flag = 0, it terminates processing without performing the inverse transform process.
[0832] The above example shows how to convert the data format of attribute information input to the encoding unit 2420 and include the converted information in the bitstream, but the method is not limited to this.
[0833] Figure 102 shows an example of the syntax for a VPS (Volumetric Parameter Set) according to this embodiment.
[0834] In VPS, first, the number of attributes (num_attribute) is signaled as information indicating the total number of attributes included in the stream. Subsequently, through a loop process (for (i=0; i<num_attribute; i++)), the attribute identifier (attribute_type[i]) corresponding to each attribute i is sequentially signaled.
[0835] attribute_type[i] is an identifier that indicates the type of attribute i. For example, in the example shown in the figure, the 0th attribute indicates "video texture coordinate (VDDC)", the 1st attribute indicates "color", and the 2nd attribute indicates "reflectance". In this way, the VPS syntax shown in Figure 102 allows the number of attributes included in the stream and the type of each attribute to be defined as header information.
[0836] Figure 103 shows an example of the syntax of the sequence parameter set for VDDC (SPS_VDDC) according to this embodiment.
[0837] For a 3D mesh encoding (VMMC) stream, the SPS_VMMC (Sequence parameter set for VMMC) parameter indicates the number of VMMC attributes included in the stream (vdmc_num_attribute), and for each VMMC attribute, it indicates the attribute identifier (vdmc_attribute_type), width (attribute_frame_width), and height (attribute_frame_height). In this example, this information is sequentially signaled for each VMMC attribute i through a loop process (for (i=0; i < vdmc_num_attribute; i++)). For example, if a VDDC attribute includes multiple video textures, it will also have multiple video attribute counts.
[0838] Furthermore, transform_flag is a flag indicating whether the VDDC attribute information includes transformation information. If transformation information is included, transform_flag[i] is set to 1, and the transformation information offset[i] or scale[i] is set. If transformation information is not included, transform_flag[i] is set to 0. It is also possible to omit transform_flag from the syntax and always set the transformation information offset or scale, or to provide separate transform_flags for offset and scale, respectively.
[0839] In the example shown in the figure, the 0th attribute of vdmc_attribute_type is "texture", the 1st attribute is "transparency", and the 2nd attribute is "reflectance". In this way, the SPS_VMMC syntax shown in Figure 103 allows the type, frame size, presence or absence of conversion information, and specific conversion information to be defined as header information for each VMMC attribute.
[0840] The above describes a configuration in which attribute information and conversion information are included in the SPS_VMMC for VMMC when the stream header has a hierarchical structure and the VMMC attribute is a video texture. However, it is not limited to this, and attribute information and conversion information may also be included in the higher-level header, VPS. Conversely, even if the stream header does not have a hierarchical structure, the configuration may include attribute information and conversion information in the header corresponding to SPS. Furthermore, in any case, attribute information and conversion information may also be included in an APS dedicated to attributes.
[0841] Figure 104 is a block diagram illustrating another example of the processing of the three-dimensional data encoding device according to this embodiment. Figure 105 is a block diagram illustrating another example of the processing of the three-dimensional data decoding device according to this embodiment.
[0842] For example, the three-dimensional data encoding device 2400 may store format information indicating the data format of the attribute information input to the conversion unit 2460 and the format of the attribute information to be encoded after conversion in additional information, and the encoding unit 2470 may encode the additional information including the format information.
[0843] The three-dimensional data decoding device 2430 can reproduce the attribute information before the conversion process in the conversion unit 2460 of the three-dimensional data encoding device 2400 by performing an inverse conversion on the decoded attribute information in the inverse conversion unit 2490 based on the format information extracted in the decoding unit 2480.
[0844] Format information (data format information) is information that indicates, for example, the data type, number of bits, whether it is signed or unsigned, etc. For example, format information is information such as int8, uint16, float16, etc. The number 8 in int8, for example, indicates the number of bits.
[0845] Alternatively, for example, the format information may indicate the file format of the three-dimensional mesh data before the conversion process is performed (e.g., obj file, mtl file, png file, etc.) as extended information SEI, or it may indicate the file format of the three-dimensional mesh data after the conversion process is performed.
[0846] Furthermore, SEI may include information indicating the header information contained in each file format. These file formats or header information may also be included in user_data.
[0847] As a result, the three-dimensional data encoding device 2400 can encode the format information extracted by the conversion unit 2460 as additional information in the encoding unit 2470, and the three-dimensional data decoding device 2430 can perform an inverse conversion on the attribute information decoded by the inverse conversion unit 2490 based on the format information obtained by the decoding unit 2480. Furthermore, the three-dimensional data decoding device 2430 can reconstruct header information and other information that are not to be encoded on the decoding side based on the format information and header information.
[0848] Furthermore, while the offset and scale of attribute information were described above, the method described in this embodiment can also be applied to positional information.
[0849] Before encoding the location information, a conversion process such as offset and scale may be performed, and after decoding the location information, a reverse conversion process may be performed. In that case, the conversion information or format information may be stored in additional information such as SPS.
[0850] Furthermore, the three-dimensional data encoding device may have a conversion unit that performs conversion on at least one or both of the location information and the attribute information. Similarly, the three-dimensional data decoding device may have an inverse conversion unit that performs conversion on either one or both (i.e., at least one) of the location information and the attribute information. In such cases, either or both (i.e., at least one) of the converted location information and the converted attribute information may be included in the additional information.
[0851] Furthermore, while offsetting and scaling were described above as methods for transforming the input three-dimensional mesh data, other transformation methods may be used. For example, the transformation process may employ predetermined linear or nonlinear transformation means, such as using a transformation or approximation using a predetermined function.
[0852] Furthermore, the additional information may include not only information indicating the format of attribute information, but also information indicating the order of the three-dimensional mesh data, information indicating the sort order, timestamp information, and so on.
[0853] As described above, the three-dimensional data encoding device according to this embodiment performs the processing shown in Figure 106.
[0854] Figure 106 is a flowchart showing the processing procedure of the three-dimensional data encoding device according to this embodiment.
[0855] First, the three-dimensional data encoding device acquires attribute information of the three-dimensional points (S2421).
[0856] Next, the three-dimensional data encoding device performs a conversion process on the numerical value indicated by the acquired attribute information, which involves scaling by performing at least one of multiplication / division and / or shift operations, and offset by performing at least one of addition / subtraction operations, before encoding the attribute information, or it performs an encoding process that encodes the attribute information without performing the conversion process (S2422). For example, the three-dimensional data encoding device determines whether the format of the attribute information is a predetermined format, and if it is a predetermined format, it converts the attribute information to a positive value using a predetermined offset value and a predetermined scale value.
[0857] Next, the three-dimensional data encoding device generates a bitstream containing encoded attribute information and transformation identification information indicating whether or not the transformation process has been performed (S2423). The transformation identification information is, for example, the transform_flag and transform_information_type information mentioned above. Alternatively, the three-dimensional data encoding device may encode the transformation identification information and generate a bitstream containing the encoded attribute information and the encoded transformation identification information.
[0858] Depending on the encoding unit that encodes the attribute information (for example, encoding unit 2420 or 2470), it may not be able to process decimal points and / or negative numbers. Therefore, for example, by performing a conversion process on the attribute information using scaling and offset, the value indicated by the attribute information can be converted to a positive number. Thus, according to the three-dimensional data encoding method of this disclosure, even if it is not possible to encode decimal points and / or negative numbers, for example, the attribute information can be appr...
Claims
1. An encoding method that obtains first attribute information which is attribute information associated with a point in the base mesh of a three-dimensional mesh, performs a first transformation process which involves performing an offset process that performs addition or subtraction on the numerical value indicated by the first attribute information and a scaling process which performs at least one of multiplication, division, and shift operations on the numerical value indicated by the first attribute information, encodes the first attribute information, and generates a bitstream which includes encoded first attribute information which includes the encoded first attribute information and an offset value and a scale value used in a first inverse transformation process to obtain the first attribute information from the encoded first attribute information.
2. The encoding method according to claim 1, wherein the acquisition further acquires second attribute information which is attribute information of the video, performs a second transformation process on the second attribute information and then encodes the second attribute information, and the bitstream includes the encoded first attribute information, encoded second attribute information which includes the encoded second attribute information, first metadata which includes first parameters which include the offset value and the scale value used in the first inverse transformation process, and second metadata which includes second parameters which are used in the second inverse transformation process which corresponds to the second transformation process.
3. The encoding method according to claim 1, wherein the first conversion process performs the scaling process after the offset process.
4. The encoding method according to any one of claims 1 to 3, wherein the offset value is expressed as a fraction with a fixed denominator, and the scale value is expressed as a fraction with a fixed denominator.
5. The encoding method according to claim 4, wherein the denominator of the fraction representing the offset value is 2 to the power of 16, and the denominator of the fraction representing the scale value is 2 to the power of 16.
6. The encoding method according to any one of claims 1 to 3, wherein the offset processing is performed with decimal precision.
7. A decoding method for obtaining the first attribute information by obtaining encoded first attribute information, which is attribute information associated with points in the base mesh of a three-dimensional mesh, and an offset value and a scale value used for the inverse transformation of the first attribute information obtained from the encoded first attribute information; decoding the encoded first attribute information from the bitstream to obtain the first attribute information, and obtaining the offset value and the scale value; and performing a first inverse transformation process on the numerical value indicated by the first attribute information, which includes a scaling process that performs at least one of multiplication / division and shift operations based on the scale value and an offset process that performs addition / subtraction based on the offset value.
8. The bitstream further includes second encoded data in which second attribute information, which is attribute information of a video, is encoded; first metadata including first parameters used in the first inverse transformation process; and second metadata including second parameters used in the second inverse transformation process on the second attribute information, wherein the first parameters include the offset value and the scale value; and the decoding method comprises decoding the second encoded data from the bitstream to decode the second attribute information, executing the first inverse transformation process based on the first parameters included in the first metadata, and executing the second inverse transformation process on the decoded second attribute information based on the second parameters included in the second metadata, according to claim 7.
9. The decoding method according to claim 7, wherein the first inverse transformation process performs the offset process after the scaling process.
10. The decoding method according to any one of claims 7 to 9, wherein the offset value is expressed as a fraction with a fixed denominator, and the scale value is expressed as a fraction with a fixed denominator.
11. The decoding method according to claim 10, wherein the denominator of the fraction representing the offset value is 2 to the power of 16, and the denominator of the fraction representing the scale value is 2 to the power of 16.
12. The decoding method according to any one of claims 7 to 9, wherein the offset processing is performed with decimal precision.
13. An encoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires first attribute information which is attribute information associated with a point in the base mesh of a three-dimensional mesh, performs a first transformation process which involves performing an offset process that performs addition or subtraction on the numerical value indicated by the first attribute information and a scaling process which performs at least one of multiplication, division and shift operations, and then encodes the first attribute information, and generates a bitstream which includes encoded first attribute information which includes the encoded first attribute information and an offset value and a scale value used in a first inverse transformation process to obtain the first attribute information from the encoded first attribute information.
14. A decoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, acquires a bitstream including encoded first attribute information, which is attribute information associated with points in the base mesh of a three-dimensional mesh, encoded in first attribute information, and offset values and scale values used for the inverse transformation of the first attribute information obtained from the encoded first attribute information; decodes the encoded first attribute information from the bitstream to obtain first attribute information, and acquires the offset values and scale values; and performs a first inverse transformation process on the numerical value indicated by the first attribute information, which includes a scaling process that performs at least one of multiplication / division and shift operations based on the scale value, and an offset process that performs addition / subtraction based on the offset value, thereby acquiring the first attribute information.