Three-dimensional data generating method, three-dimensional data decoding method, rasterization method, three-dimensional data generating device, three-dimensional data decoding device, and rasterization device

WO2026204361A1PCT designated stage Publication Date: 2026-10-01PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/009303
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-03-11
Publication Date
2026-10-01

Smart Images

  • Figure JP2026009303_01102026_PF_FP_ABST
    Figure JP2026009303_01102026_PF_FP_ABST
Patent Text Reader

Abstract

This three-dimensional data generating method, which is executed by a three-dimensional data generating device, includes: generating one or more pieces of three-dimensional Gaussian data on the basis of a plurality of two-dimensional images (S1701); executing rasterization processing corresponding to at least one viewpoint in the processing for generating the one or more pieces of three-dimensional Gaussian data (S1702); and outputting a rasterization parameter for use in the rasterization processing (S1703).
Need to check novelty before this filing date? Find Prior Art

Description

Three-dimensional data generation method, three-dimensional data decoding method, rasterization method, three-dimensional data generation apparatus, three-dimensional data decoding apparatus, and rasterization apparatus

[0001] The present disclosure relates to a three-dimensional data generation method, a three-dimensional data decoding method, a rasterization method, a three-dimensional data generation apparatus, a three-dimensional data decoding apparatus, and a rasterization apparatus.

[0002] In a wide range of fields such as computer vision for autonomous operation of automobiles or robots, map information, monitoring, infrastructure inspection, and video distribution, it is expected that devices or services utilizing three-dimensional data will become widespread in the future. Three-dimensional data is acquired by various methods, such as distance sensors like range finders, stereo cameras, or combinations of a plurality of monocular cameras.

[0003] As one method for representing three-dimensional data, there is a representation method called point cloud, which represents the shape of a three-dimensional structure by means of a point group in a three-dimensional space. In a point cloud, the positions and colors of the point group are stored. Although point cloud is expected to become a mainstream representation method for three-dimensional data, the data amount of a point group is extremely large. Therefore, in storage or transmission of three-dimensional data, compression of the data amount by encoding is essential, similarly to two-dimensional moving images (for example, MPEG-4 AVC (Advanced Video Coding) and HEVC (High Efficiency Video Coding) standardized by MPEG).

[0004] Furthermore, compression of point clouds is partially supported by, for example, a public library (Point Cloud Library) that performs point cloud-related processing.

[0005] There is also known a technology of searching for and displaying facilities located around a vehicle using three-dimensional map data (see, for example, Patent Document 1).

[0006] International Publication No. 2014 / 020663

[0007] In such generation of three-dimensional data or the like, it is desired that the reproducibility of the generation result can be improved.

[0008] This disclosure aims to provide a three-dimensional data generation method, a three-dimensional data decoding method, a rasterization method, a three-dimensional data generation apparatus, a three-dimensional data decoding apparatus, and a rasterization apparatus that can improve the reproducibility of the generated results.

[0009] A three-dimensional data generation method according to one aspect of the present disclosure is a three-dimensional data generation method performed by a three-dimensional data generation device, which generates one or more three-dimensional Gaussian data based on a plurality of two-dimensional images, and in the process of generating the one or more three-dimensional Gaussian data, performs a rasterization process corresponding to at least one viewpoint and outputs the rasterization parameters used in the rasterization process.

[0010] A three-dimensional data decoding method according to one aspect of the present disclosure is a three-dimensional data decoding method performed by a three-dimensional data decoding device, wherein the method decodes rasterization parameters used in a rasterization process performed in a process of generating one or more three-dimensional Gaussian data from a bitstream, the rasterization process corresponding to at least one viewpoint, and outputs the decoded rasterization parameters.

[0011] A rasterization method according to one aspect of the present disclosure is a rasterization method performed by a rasterization apparatus that generates a two-dimensional image based on one or more three-dimensional Gaussian data, wherein the rasterization method obtains rasterization parameters used in the rasterization process corresponding to at least one viewpoint, obtains one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane, and controls the blending process of the one or more two-dimensional Gaussian data based on the rasterization parameters.

[0012] This disclosure provides a method for generating three-dimensional data, a method for decoding three-dimensional data, a rasterization method, a three-dimensional data generation apparatus, a three-dimensional data decoding apparatus, and a rasterization apparatus that can improve the reproducibility of the generated results.

[0013] Figure 1 is a diagram showing an example configuration of a three-dimensional data encoding and decoding system according to an embodiment. Figure 2 is a diagram showing an example configuration of point cloud data according to an embodiment. Figure 3 is a diagram showing an example configuration of a data file describing point cloud data information according to an embodiment. Figure 4 is a diagram showing an example configuration of mesh data according to an embodiment. Figure 5 is a diagram showing an example configuration of a data file describing mesh data information according to an embodiment. Figure 6 is a diagram showing the types of three-dimensional data according to an embodiment. Figure 7 is a diagram showing the configuration of the first encoding unit according to an embodiment. Figure 8 is a block diagram of the first encoding unit according to an embodiment. Figure 9 is a diagram showing the configuration of the first decoding unit according to an embodiment. Figure 10 is a block diagram of the first decoding unit according to an embodiment. Figure 11 is a diagram showing the configuration of the second encoding unit according to an embodiment. Figure 12 is a block diagram of the second encoding unit according to an embodiment. Figure 13 is a diagram showing the configuration of the second decoding unit according to an embodiment. Figure 14 is a block diagram of the second decoding unit according to an embodiment. Figure 15 is a block diagram of the position information encoding unit according to an embodiment. Figure 16 is a block diagram of the position information decoding unit according to an embodiment. Figure 17 is a block diagram of the octree coding unit according to the embodiment. Figure 18 is a diagram showing an example of location information according to the embodiment. Figure 19 is a diagram showing an example of the octree representation of location information according to the embodiment. Figure 20 is a block diagram of the octree decoding unit according to the embodiment. Figure 21 is a block diagram of the attribute information coding unit according to the embodiment. Figure 22 is a block diagram of the attribute information decoding unit according to the embodiment. Figure 23 is a block diagram of the attribute information coding unit according to the embodiment. Figure 24 is a block diagram of the attribute information decoding unit according to the embodiment. Figure 25 is a block diagram of the three-dimensional data coding device according to the embodiment. Figure 26 is a block diagram of the three-dimensional data decoding device according to the embodiment. Figure 27 is a diagram showing the relationship between tiles and slices according to the embodiment. Figure 28 is a diagram showing an example of the configuration of a bitstream according to the embodiment. Figure 29 is a block diagram showing an example of the configuration of a three-dimensional data coding device according to the embodiment. Figure 30 is a diagram showing an example of the configuration of coded data and a NAL unit according to the embodiment.Figure 31 is a diagram showing an example of the semantics of pcc_nal_unit_type according to the embodiment. Figure 32 is a diagram showing a part of the processing in the three-dimensional data generation system according to the embodiment. Figure 33 is a diagram showing the components of Gaussian data according to the embodiment. Figure 34 is a diagram showing the rendering process of Gaussian data according to the embodiment. Figure 35 is a diagram showing an example of the number of SH coefficients of a spherical harmonic function according to the embodiment. Figure 36 is a diagram showing an example of input and output of a spherical harmonic function according to the embodiment. Figure 37 is a diagram showing the rendering process according to the embodiment. Figure 38 is a block diagram of the encoding device according to the embodiment. Figure 39 is a block diagram of the decoding device according to the embodiment. Figure 40 is a diagram showing the detailed processing of the three-dimensional data generation device 380 according to the embodiment. Figure 41 is a block diagram of the encoding device according to the embodiment. Figure 42 is a block diagram of the decoding device according to the embodiment. Figure 43 is a diagram showing an example of Gaussian data mapping according to the embodiment. Figure 44 is a diagram showing an example of mapping information according to the embodiment. Figure 45 is a diagram showing an example of the configuration of conversion information according to the embodiment. Figure 46 is a diagram showing an example of the configuration of encoded data according to the embodiment. Figure 47 shows an example of the G-PCC attribute type of Gaussian data according to the embodiment. Figure 48 is a flowchart of the processing by the decoding device according to the embodiment. Figure 49 shows an example of the configuration of the rendering unit. Figure 50 is a flowchart of an example of the rendering process shown in Figure 49. Figure 51 shows 3DGS extraction and projection within the viewing frustum. Figure 52 shows the syntax (example structure) of the 2DGS information gs2d_info. Figure 53 shows how to identify tiles where two-dimensional Gaussian data intersects. Figure 54 shows the syntax of the 2DGS list (Tile_GS_list) for each tile. Figure 55 shows an example of the arrangement of identifiers for two-dimensional Gaussian data. Figure 56 shows the syntax of the 2DGS list for each tile sorted by depth. Figure 57 shows an example of determining pixel values ​​by blending multiple 2DGS within a tile. Figure 58 shows the configuration of the three-dimensional data generation device. Figure 59 shows the configuration of the three-dimensional data decoding device.Figure 60 is a diagram showing the functional blocks of the rendering unit. Figure 61 is a flowchart showing an example of the processing procedure in the rendering unit. Figure 62 is a diagram showing an example of the syntax of rasterization information SEI. Figure 63 is a diagram showing an example of the syntax of camera parameter information. Figure 64 is a diagram showing an example of pose information. Figure 65 is a diagram showing an example of camera parameter information. Figure 66 is a diagram showing an example of the configuration of the rendering unit after decoding. Figure 67 is a flowchart showing an example of the rasterization process after decoding. Figure 68 is a diagram showing an example of 2DGS information. Figure 69 is a diagram showing an example of a 2DGS list for each tile. Figure 70 is a diagram showing an example of the syntax of rasterization information SEI. Figure 71 is a diagram showing the functional blocks of the rendering unit that generate a two-dimensional image. Figure 72 is a flowchart showing the flow of the rasterization process after the decoding process. Figure 73 is a diagram showing an example of 2DGS information. Figure 74 is a diagram showing an example of a 2DGS depth list for each tile. Figure 75 is an example of the syntax of sort key type SEI. Figure 76 shows an example of SEI syntax for switching and notifying multiple sort keys. Figure 77 shows syntax for notifying a recommended re-sorting method. Figure 78 shows an example of View_information_SEI syntax that includes viewpoint information from multiple viewports. Figure 79 shows the syntax for sort_by_depth_fix_intersect_SEI(). Figure 80 shows an example of generating a two-dimensional bounding box based on two-dimensional Gaussian data. Figure 81 shows an example of a scene composed of multiple frames. Figure 82 shows how tiles change with viewport movement and zooming. Figure 83 shows the syntax for reuse_prev_frame_sorting_SEI(). Figure 84 shows an example of a user manipulating a viewport. Figure 85 shows the left-right movement and zooming of the viewport shown in Figure 84 on a tiled two-dimensional plane. Figure 86 is a diagram showing an example of the configuration of an encoding device. Figure 87 is a flowchart showing an example of an encoding method using an encoding device. Figure 88 is a diagram showing an example of the configuration of a decoding device.Figure 89 is a flowchart illustrating an example of a decoding method by a decoding device. Figure 90 is a diagram illustrating an example of the configuration of a rasterization device. Figure 91 is a flowchart illustrating an example of a rasterization method by a rasterization device. Figure 92 is a block diagram illustrating an example of a device that generates format data from data. Figure 93 is a block diagram illustrating an example of a device that restores the original data from the format data. Figure 94 is a conceptual diagram illustrating an example of how an encoded data bitstream is stored in a system format. Figure 95 is a diagram illustrating an example of the box structure of an ISOBMFF.

[0014] A three-dimensional data generation method according to a first aspect of this disclosure is a three-dimensional data generation method performed by a three-dimensional data generation device, which generates one or more three-dimensional Gaussian data based on a plurality of two-dimensional images, and in the process of generating the one or more three-dimensional Gaussian data, performs a rasterization process corresponding to at least one viewpoint and outputs the rasterization parameters used in the rasterization process.

[0015] This allows subsequent processing units to reference the rasterization conditions used in the generation of three-dimensional Gaussian data. This facilitates the generation of two-dimensional images based on the same conditions and improves the reproducibility of the generation results.

[0016] A three-dimensional data generation method according to a second aspect of the present disclosure is a three-dimensional data generation method according to a first aspect, wherein the rasterization parameter includes at least one of the following: viewpoint information indicating the at least one viewpoint; generation information for generating one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane; and distance-related information relating to one or more distances between the at least one viewpoint and the one or more two-dimensional Gaussian locations indicated by the one or more two-dimensional Gaussian data.

[0017] This allows for the clear sharing of rasterization processing conditions and two-dimensional Gaussian data generation conditions for each viewpoint as viewpoint information and generation information. This enables the appropriate determination of the processing order or estimation of the processing range according to the viewpoint using distance-related information, thereby stabilizing the generation of two-dimensional images.

[0018] A three-dimensional data generation method according to a third aspect of this disclosure is a three-dimensional data generation method according to a second aspect, wherein the generated information includes two-dimensional bounding box information associated with each of the one or more two-dimensional Gaussian data.

[0019] This allows us to understand the affected region on a two-dimensional plane as two-dimensional bounding box information. This reduces the computational cost required for selecting two-dimensional Gaussian data and makes the rasterization process more efficient.

[0020] A three-dimensional data generation method according to a fourth aspect of this disclosure is a three-dimensional data generation method according to a second or third aspect, wherein the distance-related information includes the order of the one or more two-dimensional Gaussian data based on the one or more distances.

[0021] This allows for determining the processing order of distance-based two-dimensional Gaussian data without the need for additional sorting. This reduces the processing load during rasterization and speeds up the generation of two-dimensional images.

[0022] A three-dimensional data generation method according to a fifth aspect of this disclosure is a three-dimensional data generation method according to a second or third aspect, wherein the distance-related information includes identification information for each of the one or more two-dimensional Gaussian data.

[0023] This allows the target two-dimensional Gaussian data to be identified based on the identification information contained in the distance-related information. This preserves the correspondence between distance information and the two-dimensional Gaussian data, making it easier to reference in subsequent processing.

[0024] A three-dimensional data generation method according to the sixth aspect of this disclosure is a three-dimensional data generation method according to the second or third aspect, wherein the distance-related information includes distance information contained in the one or more two-dimensional Gaussian data.

[0025] This allows distance information contained in two-dimensional Gaussian data to be directly used as distance-related information. As a result, the processing order or weighting can be determined according to the viewpoint without repeatedly calculating distances.

[0026] A three-dimensional data generation method according to the seventh aspect of this disclosure is a three-dimensional data generation method according to any one aspect of the first to sixth aspects, further comprising encoding the one or more three-dimensional Gaussian data and the rasterization parameters into a bitstream.

[0027] This allows for the integrated transmission or recording of three-dimensional Gaussian data and rasterization parameters within the same bitstream. This suppresses errors in the mapping between three-dimensional Gaussian data and rasterization parameters, ensuring reliable use on the decoding side.

[0028] The three-dimensional data generation method according to the eighth aspect of this disclosure is the three-dimensional data generation method according to the seventh aspect, wherein the rasterization parameters are included in the metadata of the bitstream.

[0029] This allows rasterization parameters to be treated as metadata and referenced independently of the three-dimensional Gaussian data. This enables the decryption process to easily extract the rasterization parameters and apply them to post-decryption processing.

[0030] A three-dimensional data decoding method according to the ninth aspect of this disclosure is a three-dimensional data decoding method performed by a three-dimensional data decoding device, wherein the method decodes rasterization parameters used in a rasterization process performed in a process of generating one or more three-dimensional Gaussian data from a bitstream, the rasterization process corresponding to at least one viewpoint, and outputs the decoded rasterization parameters.

[0031] This allows the rasterization parameters decoded from the bitstream to be passed to the processing unit after decoding. This enables the generation of a two-dimensional image based on the conditions used on the encoding side, improving the reproducibility of the decoding result.

[0032] A three-dimensional data decoding method according to a tenth aspect of the present disclosure is a three-dimensional data decoding method according to a ninth aspect, wherein the rasterization parameter includes at least one of the following: viewpoint information indicating the at least one viewpoint; generation information for generating one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane; and distance-related information relating to one or more distances between the at least one viewpoint and the one or more two-dimensional Gaussian locations indicated by the one or more two-dimensional Gaussian data.

[0033] This allows the processing entity after decoding to understand the processing conditions based on viewpoint information and the generation information of the two-dimensional Gaussian data. As a result, the processing order can be determined according to the viewpoint using distance-related information, and the computational complexity of the rasterization process can be reduced.

[0034] A three-dimensional data decoding method according to an eleventh aspect of this disclosure is a three-dimensional data decoding method according to a tenth aspect, wherein the generated information includes two-dimensional bounding box information associated with each of the one or more two-dimensional Gaussian data.

[0035] This allows us to estimate the affected regions based on two-dimensional bounding box information. This reduces the computational cost required for selecting two-dimensional Gaussian data and makes the rasterization process more efficient.

[0036] A three-dimensional data decoding method according to a twelfth aspect of this disclosure is a three-dimensional data decoding method according to a tenth or eleventh aspect, wherein the distance-related information includes the order of the one or more three-dimensional Gaussian data based on the one or more distances.

[0037] This allows the processing order after decoding to be determined by utilizing the order of three-dimensional Gaussian data based on distance. This reduces the computational cost required for sorting and speeds up the generation of two-dimensional images.

[0038] A three-dimensional data decoding method according to a thirteenth aspect of this disclosure is a three-dimensional data decoding method according to a tenth or eleventh aspect, wherein the distance-related information includes identification information for each of the one or more two-dimensional Gaussian data.

[0039] This allows for the identification of target Gaussian data based on identification information and the maintenance of a correspondence with distance-related information. This simplifies the process by which the decoded processing entity references the Gaussian data and improves processing stability.

[0040] A three-dimensional data decoding method according to a fourteenth aspect of this disclosure is a three-dimensional data decoding method according to a tenth or eleventh aspect, wherein the distance-related information includes distance information corresponding to one or more two-dimensional Gaussian data.

[0041] This allows distance information corresponding to two-dimensional Gaussian data to be referenced after decoding. This enables the determination or weighting of the processing order according to the viewpoint while omitting the calculation of distance.

[0042] A three-dimensional data decoding method according to a 15th aspect of this disclosure is a three-dimensional data decoding method according to any one aspect of the 9th to 14th aspects, wherein the rasterization parameters are included in the metadata of the bitstream.

[0043] This allows for easy extraction of rasterization parameters as metadata from the bitstream. This ensures flexibility in how rasterization parameters are communicated while facilitating their application in post-decryption processing.

[0044] The rasterization method according to the 16th aspect of the present disclosure is a rasterization method executed by a rasterization apparatus that generates a two-dimensional image based on one or more pieces of three-dimensional Gaussian data, the method comprising: acquiring a rasterization parameter used for rasterization processing corresponding to at least one viewpoint; acquiring one or more pieces of two-dimensional Gaussian data obtained by projecting the one or more pieces of three-dimensional Gaussian data onto a two-dimensional plane; and controlling blending processing of the one or more pieces of two-dimensional Gaussian data based on the rasterization parameter.

[0045] Accordingly, the rasterization parameter used on the encoding side can be acquired, and the blending processing of two-dimensional Gaussian data can be controlled based on the rasterization parameter. Thereby, a blending condition corresponding to a viewpoint can be appropriately applied, and the reproducibility of a two-dimensional image generation result can be improved.

[0046] The rasterization method according to the 17th aspect of the present disclosure is the rasterization method according to the 16th aspect, further comprising: determining whether to use the rasterization parameter; when the rasterization parameter is used, executing blending processing of the one or more pieces of two-dimensional Gaussian data based on the rasterization parameter; and when the rasterization parameter is not used, executing blending processing of the one or more pieces of two-dimensional Gaussian data based on at least depth information of the one or more pieces of two-dimensional Gaussian data.

[0047] Accordingly, whether to use the rasterization parameter is determined, and execution conditions for blending processing can be switched between a case where the rasterization parameter is used and a case where the rasterization parameter is not used. Thereby, when the rasterization parameter is available, processing based on the rasterization parameter can be applied, and when the rasterization parameter is not available, generation of a two-dimensional image can be continued by processing based on depth information.

[0048] A three-dimensional data generation apparatus according to an eighteenth aspect of the present disclosure includes a circuit and a memory connected to the circuit. In operation, the circuit generates one or more three-dimensional Gaussian data based on a plurality of two-dimensional images, executes rasterization processing corresponding to at least one viewpoint in the processing of generating the one or more three-dimensional Gaussian data, and outputs rasterization parameters used for the rasterization processing.

[0049] Accordingly, a subsequent processing entity can refer to the rasterization processing conditions used in the three-dimensional Gaussian data generation processing. This facilitates generation of two-dimensional images based on the same conditions and improves the reproducibility of generation results.

[0050] A three-dimensional data decoding apparatus according to a nineteenth aspect of the present disclosure includes a circuit and a memory connected to the circuit. In operation, in processing of generating one or more three-dimensional Gaussian data from a bitstream, the circuit decodes rasterization parameters used for rasterization processing corresponding to at least one viewpoint, and outputs the decoded rasterization parameters.

[0051] Accordingly, the rasterization parameters decoded from the bitstream can be delivered to a processing entity after decoding. This enables generation of two-dimensional images based on the conditions used on the encoding side, and improves the reproducibility of decoding results.

[0052] A rasterization apparatus according to a twentieth aspect of the present disclosure is a rasterization apparatus that generates a two-dimensional image based on one or more three-dimensional Gaussian data, and includes a circuit and a memory connected to the circuit. In operation, the circuit acquires rasterization parameters used for rasterization processing corresponding to at least one viewpoint, acquires one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane, and controls blending processing of the one or more two-dimensional Gaussian data based on the rasterization parameters.

[0053] This allows the system to determine whether or not to use rasterization parameters and switch the execution conditions for blending depending on whether or not rasterization parameters are used. As a result, if rasterization parameters are available, processing based on those rasterization parameters is applied, and if rasterization parameters are not available, the generation of a two-dimensional image can continue using processing based on depth information.

[0054] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, and recording medium.

[0055] The embodiments will be described in detail below with reference to the drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, materials, components, arrangement and connection configurations of components, steps, and the order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.

[0056] (Embodiment) [Three-Dimensional Data Encoding and Decoding System] First, the configuration of the three-dimensional data encoding and decoding system according to this embodiment will be described. Figure 1 is a diagram showing an example of the configuration of the three-dimensional data encoding and decoding system according to this embodiment. As shown in Figure 1, the three-dimensional data encoding and decoding system includes a three-dimensional data encoding system 101, a three-dimensional data decoding system 102, a sensor terminal 103, and an external connection unit 104.

[0057] The three-dimensional data encoding system 101 generates encoded data or multiplexed data by encoding three-dimensional data such as three-dimensional point cloud data or three-dimensional mesh data. The three-dimensional data encoding system 101 may be a three-dimensional data encoding device implemented by a single device, or it may be a system implemented by multiple devices. Furthermore, the three-dimensional data encoding device may include some of the multiple processing units included in the three-dimensional data encoding system 101.

[0058] The three-dimensional data encoding system 101 includes a three-dimensional data generation system 111, a presentation unit 112, an encoding unit 113, a multiplexing unit 114, an input / output unit 115, and a control unit 116. The three-dimensional data generation system 111 also includes a sensor information acquisition unit 117 and a three-dimensional data generation unit 118.

[0059] The sensor information acquisition unit 117 acquires sensor information (sensor signals) from the sensor terminal 103 and outputs the sensor information to the three-dimensional data generation unit 118. The three-dimensional data generation unit 118 generates three-dimensional data from the sensor information and outputs the three-dimensional data to the encoding unit 113.

[0060] The display unit 112 presents sensor information or three-dimensional data to the user. For example, the display unit 112 displays information or images based on sensor information or three-dimensional data.

[0061] The encoding unit 113 encodes (compresses) the three-dimensional data and outputs the resulting encoded data, control information obtained during the encoding process, and other additional information to the multiplexing unit 114. The additional information includes, for example, sensor information.

[0062] The multiplexing unit 114 generates multiplexed data by multiplexing the encoded data input from the encoding unit 113, control information, and additional information. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.

[0063] The input / output unit 115 (for example, the communication unit or interface) outputs the multiplexed data to the outside. Alternatively, the multiplexed data is stored in a storage unit such as internal memory. The control unit 116 (or application execution unit) controls each processing unit. In other words, the control unit 116 performs control such as encoding and multiplexing.

[0064] The sensor information may also be input to the encoding unit 113 or the multiplexing unit 114. Furthermore, the input / output unit 115 may output the three-dimensional data or encoded data directly to the outside.

[0065] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 101 is input to the three-dimensional data decoding system 102 via the external connection unit 104.

[0066] The three-dimensional data decoding system 102 generates three-dimensional data, such as three-dimensional point cloud data or three-dimensional mesh data, by decoding encoded data or multiplexed data. The three-dimensional data decoding system 102 may be a three-dimensional data decoding device implemented by a single device, or it may be a system implemented by multiple devices. Furthermore, the three-dimensional data decoding device may include some of the multiple processing units included in the three-dimensional data decoding system 102.

[0067] The three-dimensional data decoding system 102 includes a sensor information acquisition unit 121, an input / output unit 122, a demultiplexing unit 123, a decoding unit 124, a presentation unit 125, a user interface 126, and a control unit 127.

[0068] The sensor information acquisition unit 121 acquires sensor information (sensor signals) from the sensor terminal 103.

[0069] The input / output unit 122 acquires the transmission signal, decodes the multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the demultiplexing unit 123.

[0070] The demultiplexing unit 123 acquires encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 124.

[0071] The decoding unit 124 reconstructs the three-dimensional data by decoding the encoded data.

[0072] The presentation unit 125 presents three-dimensional data to the user. For example, the presentation unit 125 displays information or images based on the three-dimensional data. The user interface 126 acquires instructions based on user operations. The control unit 127 (or application execution unit) controls each processing unit. In other words, the control unit 127 performs control such as demultiplexing, decoding, and presentation.

[0073] The input / output unit 122 may acquire three-dimensional data or encoded data directly from an external source. The presentation unit 125 may acquire additional information such as sensor information and present information based on that additional information. The presentation unit 125 may also make presentations based on user instructions acquired through the user interface 126.

[0074] The sensor terminal 103 generates sensor information, which is information obtained from the sensor. The sensor terminal 103 is a terminal equipped with a sensor or camera, and may be, for example, a mobile object such as an automobile, an aerial object such as an airplane, a mobile terminal, or a camera.

[0075] The sensor information that can be acquired by the sensor terminal 103 includes, for example, (1) the distance between the sensor terminal 103 and the object, or the reflectivity of the object, obtained from a LiDAR, millimeter-wave radar, or infrared sensor, and (2) the distance between the camera and the object, or the reflectivity of the object, obtained from multiple monocular camera images or stereo camera images. The sensor information may also include the sensor's attitude, orientation, gyroscope (angular velocity), position (GPS information or altitude), speed, or acceleration. The sensor information may also include temperature, atmospheric pressure, humidity, or magnetism.

[0076] The external connection unit 104 is realized by an integrated circuit (LSI or IC), an external storage unit, communication with a cloud server via the internet, or broadcasting, etc.

[0077] Next, we will explain three-dimensional point cloud data (hereinafter also referred to as point cloud data). Figure 2 is a diagram showing the structure of point cloud data. Figure 3 is a diagram showing an example of the structure of a data file containing information about point cloud data.

[0078] Point cloud data contains data for multiple points. Each point's data includes location information (three-dimensional coordinates) and attribute information related to that location. A collection of these points is called a point cloud. For example, a point cloud represents the three-dimensional shape of an object.

[0079] Position information, such as three-dimensional coordinates, is sometimes called geometry. Furthermore, the data for each point may include attribute information of multiple attribute types. Attribute types include, for example, color or reflectance.

[0080] One location information may be associated with one attribute information, or multiple attribute information of different attribute types may be associated with one location information. Furthermore, multiple attribute information of the same attribute type may be associated with one location information.

[0081] The example data file structure shown in Figure 3 represents a case where location information and attribute information correspond one-to-one, and it shows the location information and attribute information of the N points that make up the point cloud data.

[0082] Location information includes, for example, information for the three axes: x, y, and z. Attribute information includes, for example, RGB color information. A typical data file is a ply file.

[0083] Next, we will explain three-dimensional mesh data (hereinafter also referred to as mesh data). Figure 4 shows an example of the structure of mesh data. Figure 5 shows an example of the structure of a data file containing mesh data information.

[0084] Mesh data is a data format used in computer graphics (CG). Mesh data represents the three-dimensional shape of an object through a collection of surface information. Surface information consists of polygons such as triangles or quadrilaterals, and is also referred to as a polygon or polygon mesh.

[0085] The components of mesh data include a three-dimensional point cloud (a set of points that have three-dimensional positional information and attribute information corresponding to that positional information), as well as a set of three-dimensional points as vertices, edges connecting two vertices, and faces enclosed by edges.

[0086] Vertices (also expressed as vertex or position) may have attribute information such as color information, reflectivity, or normal vectors. Information indicating the relationships between vertices that constitute an edge or face is also called connectivity. The direction of the normal vector to a point can represent the front and back of a face. Mesh data may also contain attribute information for faces.

[0087] One example of a mesh data file format is an object file. The data file contains the position information G(1) to G(N) and vertex attribute information A(1) to A(N) for the N vertices that make up the mesh. Note that the data file does not have to include attribute information. Also, the attribute information does not have to correspond one-to-one with the vertices. Figure 5 shows an example where the data file has 1 to M attribute information A2.

[0088] Face information is represented by a combination of vertex indices. n[1, 3, 4] indicates that the face is a triangular face composed of three vertices n=1, n=3, and n=4. m[2, 4, 6] indicates that attribute information m=1, m=4, and m=6 correspond to the three vertices, respectively.

[0089] Furthermore, attribute information may be recorded in a separate file from the data file, with the data file indicating its pointer information. For example, attribute information may be stored in a two-dimensional attribute map file, and the file name of the attribute map and its two-dimensional coordinates may be recorded in attribute information A2 of the data file. In any of these methods, it is possible to specify attribute information for a point.

[0090] Next, we will explain the types of three-dimensional data (point cloud data or mesh data). Figure 6 is a diagram showing the types of three-dimensional data. As shown in Figure 6, three-dimensional data includes static objects and dynamic objects.

[0091] A static object is three-dimensional data for any given time (a specific moment). A dynamic object is three-dimensional data that changes over time. Hereafter, three-dimensional point cloud data for a given time will be referred to as a PCC (Point Cloud Compression) frame, or simply a frame. Similarly, three-dimensional mesh data for a given time will be referred to as a mesh frame, or simply a frame.

[0092] The multiple points that make up an object may have limitations in terms of area, pixels, or the number of points, similar to regular video data. Alternatively, the multiple points that make up an object may not have any area limitations, similar to map information.

[0093] Furthermore, there may be point cloud data or mesh data of various densities, including both sparse and dense point cloud data or mesh data.

[0094] The details of each processing unit are described below. Sensor information is acquired by various methods, such as distance sensors like LIDAR or rangefinders, stereo cameras, or combinations of multiple monocular cameras. The three-dimensional data generation unit 118 generates three-dimensional data based on the sensor information obtained by the sensor information acquisition unit 117. The three-dimensional data generation unit 118 generates position information as three-dimensional data and adds attribute information to the position information.

[0095] The three-dimensional data generation unit 118 may process the three-dimensional data when generating position information or adding attribute information. For example, the three-dimensional data generation unit 118 may reduce the amount of data by deleting point clouds with overlapping positions. The three-dimensional data generation unit 118 may also transform the position information (e.g., position shift, rotation, or normalization). Furthermore, the three-dimensional data generation unit 118 may generate mesh data from the point cloud data. Finally, the three-dimensional data generation unit 118 may render attribute information.

[0096] In Figure 1, the three-dimensional data generation system 111 is included in the three-dimensional data encoding system 101, but it may also be provided independently outside of the three-dimensional data encoding system 101.

[0097] The encoding unit 113 generates encoded data by encoding three-dimensional data. There are the following encoding methods. The first is an encoding method using positional information, which will be referred to as the first encoding method hereafter. The second is an encoding method using a video codec, which will be referred to as the second encoding method hereafter.

[0098] The decoding unit 124 decodes the point cloud data by decoding the encoded data. The multiplexing unit 114 generates multiplexed data by multiplexing the encoded data using an existing multiplexing method. The generated multiplexed data is transmitted or stored. In addition to the PCC encoded data, the multiplexing unit 114 multiplexes other media such as video, audio, subtitles, applications, files, or reference time information. Furthermore, the multiplexing unit 114 may also multiplex sensor information or attribute information related to the three-dimensional data.

[0099] Multiplexing methods or file formats include ISOBMFF, ISOBMFF-based transmission methods such as MPEG-DASH, MMT, MPEG-2 TS Systems, and RTP.

[0100] The demultiplexing unit 123 extracts encoded data, other media, and time information from the multiplexed data.

[0101] The input / output unit 115 transmits the multiplexed data using a method appropriate to the transmission medium or storage medium, such as broadcasting or communication. The input / output unit 115 may communicate with other devices via the Internet, or with storage units such as cloud servers.

[0102] Communication protocols such as HTTP, FTP, TCP, or UDP can be used. A pull-type communication method or a push-type communication method may be used.

[0103] Either wired or wireless transmission may be used. Wired transmission methods include Ethernet®, USB, RS-232C, HDMI®, or coaxial cable. Wireless transmission methods include wireless LAN, Wi-Fi®, Bluetooth®, or millimeter wave.

[0104] Furthermore, broadcasting formats such as DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3 may be used.

[0105] [First Encoding Method] Hereafter, encoding and decoding methods for three-dimensional point clouds or three-dimensional meshes will be described. When the apparatus, processing, or syntax in this disclosure relates to the encoding or decoding of point cloud data, it can also be applied to the encoding or decoding of vertices in a three-dimensional mesh. Furthermore, disclosures relating to the encoding or decoding of vertices in a three-dimensional mesh can also be applied to the encoding or decoding of point cloud data. In addition, disclosures relating to the encoding or decoding of attribute information of point cloud data may be applied to the encoding or decoding of face information, or attribute information for faces or vertices in a three-dimensional mesh. Moreover, the processing may be unified between point cloud encoding and mesh encoding. By unifying the processing, it may be possible to reduce the size of the circuit or software.

[0106] Figure 7 shows the configuration of a first encoding unit 130, which is an example of an encoding unit 113 that performs encoding using the first encoding method. Figure 8 is a block diagram of the first encoding unit 130. The first encoding unit 130 generates encoded data (encoded stream) by encoding point cloud data using the first encoding method. This first encoding unit 130 includes a location information encoding unit 131, an attribute information encoding unit 132, an additional information encoding unit 133, and a multiplexing unit 134.

[0107] The first encoding unit 130 is characterized by performing encoding while being aware of the three-dimensional structure. Furthermore, the first encoding unit 130 is characterized in that the attribute information encoding unit 132 performs encoding using information obtained from the position information encoding unit 131. The first encoding method is also called G-PCC (Geometry-based PCC).

[0108] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information, attribute information, and other additional information (MetaData). The position information is input to the position information encoding unit 131, the attribute information is input to the attribute information encoding unit 132, and the additional information is input to the additional information encoding unit 133.

[0109] When mesh data is encoded, the position information of the vertex information is input to the position information encoding unit 131, and the face information, or attribute information for faces and vertices, is input to the attribute information encoding unit 132.

[0110] The location information encoding unit 131 generates encoded location information (Compressed Geometry), which is encoded data, by encoding the location information. For example, the location information encoding unit 131 encodes the location information using an N-tree structure such as an octree. Specifically, in an octree, the target space is divided into eight nodes (subspaces), and eight bits of information (occupancy code) are generated to indicate whether or not a point cloud is contained in each node. Furthermore, nodes containing point clouds are further divided into eight nodes, and eight bits of information are generated to indicate whether or not a point cloud is contained in each of these eight nodes. This process is repeated until the number of point clouds contained in a predetermined hierarchy or node falls below a threshold.

[0111] The attribute information encoding unit 132 generates encoded attribute information (Compressed Attribute), which is encoded data, by encoding it using the configuration information generated by the location information encoding unit 131. For example, the attribute information encoding unit 132 determines the reference point (reference node) to be referenced in encoding the target point (target node) to be processed, based on the octave tree structure generated by the location information encoding unit 131. For example, the attribute information encoding unit 132 references a surrounding node or adjacent node whose parent node in the octave tree is the same as the target node. However, the method for determining the reference relationship is not limited to this.

[0112] Furthermore, the attribute information encoding process may include at least one of the following: quantization, prediction, and arithmetic encoding. In this case, a reference means using a reference node to calculate the predicted value of the attribute information, or using the state of a reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the encoding parameters. For example, encoding parameters may be quantization parameters in the quantization process, or context in arithmetic encoding.

[0113] The additional information encoding unit 133 generates encoded additional information (Compressed MetaData), which is encoded data, by encoding the compressible data among the additional information.

[0114] The multiplexing unit 134 generates a compressed stream, which is encoded data, by multiplexing encoded position information, encoded attribute information, encoded additional information, and other additional information. The generated compressed stream is output to a processing unit of the system layer (not shown).

[0115] Next, we will describe a first decoding unit 140, which is an example of a decoding unit 124 that performs decoding of the first encoding method. Figure 9 is a diagram showing the configuration of the first decoding unit 140. Figure 10 is a block diagram of the first decoding unit 140. The first decoding unit 140 generates point cloud data by decoding the encoded data (encoded stream) encoded by the first encoding method using the first encoding method. This first decoding unit 140 includes a demultiplexing unit 141, a location information decoding unit 142, an attribute information decoding unit 143, and an additional information decoding unit 144.

[0116] A encoded stream, which is encoded data, is input to the first decoding unit 140 from a processing unit of the system layer (not shown).

[0117] The demultiplexing unit 141 separates encoded location information (Compressed Geometry), encoded attribute information (Compressed Attribute), encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0118] The location information decoding unit 142 generates location information by decoding the encoded location information. For example, the location information decoding unit 142 reconstructs the location information of a point cloud represented by three-dimensional coordinates from encoded location information represented by an N-tree structure such as an octree.

[0119] The attribute information decoding unit 143 decodes the encoded attribute information based on the configuration information generated by the location information decoding unit 142. For example, the attribute information decoding unit 143 determines the reference point (reference node) to be referenced in the decoding of the target point (target node) to be processed, based on the octave tree structure obtained by the location information decoding unit 142. For example, the attribute information decoding unit 143 references a surrounding node or adjacent node whose parent node in the octave tree is the same as the target node. Note that the method for determining the reference relationship is not limited to this.

[0120] Furthermore, the attribute information decoding process may include at least one of the following: inverse quantization, prediction, and arithmetic decoding. In this case, "reference" means using a reference node to calculate the predicted value of the attribute information, or using the state of the reference node (for example, occupancy information indicating whether or not the reference node contains a point cloud) to determine the decoding parameters. For example, decoding parameters may be quantization parameters in the inverse quantization process, or context in arithmetic decoding.

[0121] The additional information decoding unit 144 generates additional information by decoding the encoded additional information. The first decoding unit 140 uses the additional information necessary for decoding location information and attribute information during decoding and outputs the additional information necessary for the application to the outside.

[0122] [Second Encoding Method] Next, a second encoding unit 150, which is an example of an encoding unit 113 that performs encoding using the second encoding method, will be described. Figure 11 is a diagram showing the configuration of the second encoding unit 150. Figure 12 is a block diagram of the second encoding unit 150.

[0123] The second encoding unit 150 generates encoded data (encoded stream) by encoding the point cloud data using a second encoding method. This second encoding unit 150 includes an additional information generation unit 151, a position image generation unit 152, an attribute image generation unit 153, a video encoding unit 154, an additional information encoding unit 155, and a multiplexing unit 156.

[0124] The second encoding unit 150 is characterized by generating a position image and an attribute image by projecting a three-dimensional structure onto a two-dimensional image, and encoding the generated position image and attribute image using an existing video encoding scheme. The second encoding method is also called VPCC (Video-based PCC).

[0125] The point cloud data is PCC point cloud data such as a PLY file, or PCC point cloud data generated from sensor information, and includes position information, attribute information, and other additional information (MetaData).

[0126] The additional information generation unit 151 generates map information for multiple two-dimensional images by projecting a three-dimensional structure onto a two-dimensional image.

[0127] The location image generation unit 152 generates a location image (Geometry Image) based on location information and map information generated by the additional information generation unit 151. This location image is, for example, a depth image in which the distance is indicated as a pixel value. This depth image may be an image of multiple point clouds viewed from one viewpoint (an image of multiple point clouds projected onto a single two-dimensional plane), or multiple images of multiple point clouds viewed from multiple viewpoints, or a single image formed by integrating these multiple images.

[0128] The attribute image generation unit 153 generates an attribute image based on attribute information and map information generated by the additional information generation unit 151. This attribute image is, for example, an image in which attribute information (e.g., color (RGB)) is shown as pixel values. This image may be an image of multiple point clouds viewed from one viewpoint (an image of multiple point clouds projected onto a single two-dimensional plane), or multiple images of multiple point clouds viewed from multiple viewpoints, or a single image formed by integrating these multiple images.

[0129] When mesh data is encoded, the position information of the vertex information is input to the position image generation unit 152, and the face information, or attribute information for faces and vertices, is input to the attribute image generation unit 153.

[0130] The video encoding unit 154 generates encoded data, namely a encoded geometry image and an encoded attribute image, by encoding the position image and attribute image using a video encoding scheme. Any known encoding method may be used as the video encoding scheme. For example, the video encoding scheme may be AVC or HEVC.

[0131] The additional information encoding unit 155 generates encoded additional information (Compressed MetaData) by encoding additional information and map information included in the point cloud data.

[0132] The multiplexing unit 156 generates a compressed stream, which is encoded data, by multiplexing the encoded position image, encoded attribute image, encoded additional information, and other additional information. The generated compressed stream is output to a processing unit of a system layer (not shown).

[0133] Next, we will describe a second decoding unit 160, which is an example of a decoding unit 124 that performs decoding of the second encoding method. Figure 13 is a diagram showing the configuration of the second decoding unit 160. Figure 14 is a block diagram of the second decoding unit 160. The second decoding unit 160 generates point cloud data by decoding the encoded data (encoded stream) encoded by the second encoding method using the second encoding method. This second decoding unit 160 includes a demultiplexing unit 161, a video decoding unit 162, an additional information decoding unit 163, a location information generation unit 164, and an attribute information generation unit 165.

[0134] A encoded stream (Compressed Stream), which is encoded data, is input to the second decoding unit 160 from a processing unit of the system layer (not shown).

[0135] The demultiplexing unit 161 separates the encoded location image (Compressed Geometry Image), encoded attribute image (Compressed Attribute Image), encoded additional information (Compressed MetaData), and other additional information from the encoded data.

[0136] The video decoding unit 162 generates a position image and an attribute image by decoding the encoded position image and the encoded attribute image using a video encoding scheme. Any known encoding scheme may be used as the video encoding scheme. For example, the video encoding scheme may be AVC or HEVC.

[0137] The additional information decoding unit 163 generates additional information, including map information, by decoding the encoded additional information.

[0138] The location information generation unit 164 generates location information using the location image and map information. The attribute information generation unit 165 generates attribute information using the attribute image and map information.

[0139] The second decoding unit 160 uses the additional information necessary for decoding during decoding and outputs the additional information necessary for the application to the outside.

[0140] [Location Information Coding in the First Coding Method] Figure 15 is a block diagram showing an example configuration of the location information coding unit 131. The location information coding unit 131 comprises an octree coding unit 171 and a prediction tree coding unit 172. The octree coding unit 171 generates coded location information and metadata by coding location information using an octree coding method (octree coding). The prediction tree coding unit 172 generates coded location information and metadata by coding location information using a prediction tree coding method (prediction tree coding).

[0141] The location information encoding unit 131 encodes location information using either octree encoding or predictive tree encoding, or both. The location information encoding unit 131 may switch between these two methods, use an encoding method other than these two, or include the location information in the bitstream as raw data without encoding it. Information indicating the encoding method of the encoded data is stored in metadata and notified to the decoding device.

[0142] Figure 16 is a block diagram showing an example configuration of the location information decoding unit 142. The location information decoding unit 142 comprises an octree decoding unit 173 and a predictive tree decoding unit 174. The octree decoding unit 173 generates location information by decoding the encoded location information using an octree decoding method (octree decoding). The predictive tree decoding unit 174 generates location information by decoding the encoded location information using a predictive tree decoding method (predictive tree decoding). The location information decoding unit 142 also performs decoding using the encoding method notified in the metadata.

[0143] Next, an example of the configuration of the location information coding unit will be described. Figure 17 is a block diagram of the octree coding unit 171 according to this embodiment. The octree coding unit 171 comprises an octree generation unit 181, a geometric information calculation unit 182, a coding table selection unit 183, and an entropy coding unit 184.

[0144] The octree generation unit 181 generates an octree from the input position information and generates occupancy codes for each node in the octree. The geometric information calculation unit 182 obtains information indicating whether the adjacent nodes of the target node are occupied nodes or not. For example, the geometric information calculation unit 182 calculates the occupancy information of adjacent nodes (information indicating whether the adjacent node is an occupied node or not) from the occupancy code of the parent node to which the target node belongs. Alternatively, the geometric information calculation unit 182 may store the encoded nodes in a list and search for adjacent nodes from that list. The geometric information calculation unit 182 may also switch adjacent nodes depending on the position of the target node within the parent node.

[0145] The coding table selection unit 183 selects a coding table to be used for entropy coding of the target node using the occupancy information of adjacent nodes calculated by the geometric information calculation unit 182. For example, the coding table selection unit 183 may generate a bit sequence using the occupancy information of adjacent nodes and select a coding table with an index number generated from that bit sequence.

[0146] The entropy coding unit 184 generates coded location information and metadata by performing entropy coding on the occupancy code of the target node using the coding table of the selected index number. The entropy coding unit 184 may also add information indicating the selected coding table to the coded location information.

[0147] The following describes the octree representation and the scanning order of location information. Location information (location data) is converted into an octree structure (octreeization) and then encoded. An octree structure consists of nodes and leaves. Each node has eight nodes or leaves, and each leaf has voxel (VXL) information. Figure 18 shows an example of the structure of location information containing multiple voxels. Figure 19 shows an example of the location information shown in Figure 18 converted into an octree structure. Here, among the leaves shown in Figure 19, leaves 1, 2, and 3 represent the voxels VXL1, VXL2, and VXL3 shown in Figure 18, respectively, and represent a VXL containing a point cloud (hereinafter referred to as effective VXL).

[0148] Specifically, node 1 corresponds to the overall space encompassing the positional information in Figure 18. The overall space corresponding to node 1 is divided into eight nodes, and of these eight nodes, the node containing the valid VXL is further divided into eight nodes or leaves, and this process is repeated for each level of the tree structure. Here, each node corresponds to a subspace and holds information (occupancy code) indicating the position of the next node or leaf after the division as node information. In addition, the lowest-level block is set as a leaf, and leaf information such as the number of points contained within the leaf is held.

[0149] Next, an example of the configuration of the location information decoding unit will be described. Figure 20 is a block diagram of the octree decoding unit 173 according to this embodiment. The octree decoding unit 173 comprises an octree generation unit 191, a geometric information calculation unit 192, an encoding table selection unit 193, and an entropy decoding unit 194.

[0150] The octane tree generation unit 191 generates an octane tree of a certain space (node) using the header information or metadata of the bitstream. For example, the octane tree generation unit 191 generates a large space (root node) using the size of the x, y, and z axis directions of a certain space attached to the header information, and generates an octane tree by dividing that space into two in each direction of the x, y, and z axes to generate eight small spaces A (nodes A0 to A7). Also, nodes A0 to A7 are set in order as the target nodes.

[0151] The geometric information calculation unit 192 obtains occupancy information indicating whether an adjacent node to the target node is an occupied node. For example, the geometric information calculation unit 192 calculates the occupancy information of an adjacent node from the occupancy code of the parent node to which the target node belongs. Alternatively, the geometric information calculation unit 192 may store the decoded nodes in a list and search for adjacent nodes from that list. The geometric information calculation unit 192 may also switch adjacent nodes depending on the position of the target node within its parent node.

[0152] The coding table selection unit 193 selects a coding table (decoding table) to be used for entropy decoding of the target node using the occupancy information of adjacent nodes calculated by the geometric information calculation unit 192. For example, the coding table selection unit 193 may generate a bit sequence using the occupancy information of adjacent nodes and select a coding table with an index number generated from that bit sequence.

[0153] The entropy decoding unit 194 generates location information by entropy decoding the occupancy code of the target node using the selected coding table. Alternatively, the entropy decoding unit 194 may decode and obtain information from the bitstream regarding the selected coding table, and then entropy decode the occupancy code of the target node using the coding table indicated by this information.

[0154] [Attribute Information Encoding in the First Encoding Scheme] The configuration of the attribute information encoding unit and the attribute information decoding unit will be described below. Figure 21 is a block diagram showing an example configuration of the attribute information encoding unit 132. The attribute information encoding unit may include multiple encoding units that perform different encoding methods. For example, the attribute information encoding unit may switch between the following two methods depending on the use case.

[0155] The attribute information encoding unit 132 includes an LoD attribute information encoding unit 201 and a conversion attribute information encoding unit 202. The LoD attribute information encoding unit 201 uses the positional information of the three-dimensional points to classify each three-dimensional point into multiple layers, predicts the attribute information of the three-dimensional points belonging to each layer, and encodes the predicted residual. Here, each classified layer is called LoD (Level of Detail).

[0156] The attribute information encoding unit 202 encodes attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the attribute information encoding unit 202 generates high-frequency and low-frequency components of each layer by applying RAHT or Haar transform to each attribute information based on the position information of three-dimensional points, and encodes these values ​​using quantization and entropy coding, etc.

[0157] Figure 22 is a block diagram showing an example configuration of the attribute information decoding unit 143. The attribute information decoding unit may include multiple decoding units that perform different decoding methods. For example, the attribute information decoding unit may decode by switching between the following two methods based on the information contained in the header and metadata.

[0158] The attribute information decoding unit 143 includes an LoD attribute information decoding unit 203 and a converted attribute information decoding unit 204. The LoD attribute information decoding unit 203 classifies each three-dimensional point into multiple layers using the position information of the three-dimensional points, and decodes the attribute values ​​while predicting the attribute information of the three-dimensional points belonging to each layer.

[0159] The attribute information decoding unit 204 decodes attribute information using RAHT (Region Adaptive Hierarchical Transform). Specifically, the attribute information decoding unit 204 decodes attribute values ​​by applying inverse RAHT or inverse Haar transform to the high-frequency and low-frequency components of each attribute value based on the position information of the three-dimensional points.

[0160] Figure 23 is a block diagram of a transformation attribute information encoding unit 202, which is an example of the transformation attribute information encoding unit 202. The transformation attribute information encoding unit 202 comprises a sorting unit 211, a Haar transformation unit 212, a quantization unit 213, an inverse quantization unit 214, an inverse Haar transformation unit 215, a memory 216, and an arithmetic encoding unit 217.

[0161] The sorting unit 211 generates Morton codes using the position information of three-dimensional points and sorts the multiple three-dimensional points in Morton code order. The Haar transform unit 212 generates coding coefficients by applying the Haar transform to the attribute information. The quantization unit 213 quantizes the coding coefficients of the attribute information.

[0162] The inverse quantization unit 214 inversely quantizes the encoded coefficients after quantization. The inverse Haar transform unit 215 applies the inverse Haar transform to the encoded coefficients. The memory 216 stores the attribute information values ​​of the decoded three-dimensional points. For example, the attribute information of the decoded three-dimensional points stored in the memory 216 may be used for predicting unencoded three-dimensional points.

[0163] The arithmetic coding unit 217 calculates ZeroCnt from the quantized coding coefficients and arithmetically codes ZeroCnt. The arithmetic coding unit 217 also arithmetically codes the non-zero coding coefficients after quantization. The arithmetic coding unit 217 may also binarize the coding coefficients before arithmetic coding. The arithmetic coding unit 217 may also generate and code various header information.

[0164] Figure 24 is a block diagram of a conversion attribute information decoding unit 204, which is an example of a conversion attribute information decoding unit 204. The conversion attribute information decoding unit 204 comprises an arithmetic decoding unit 221, an inverse quantization unit 222, an inverse Haar conversion unit 223, and a memory 224.

[0165] The arithmetic decoding unit 221 arithmetically decodes the ZeroCnt and coding coefficients contained in the bitstream. The arithmetic decoding unit 221 may also decode various header information.

[0166] The inverse quantization unit 222 inversely quantizes the arithmetic-decoded coding coefficients. The inverse Haar transform unit 223 applies the inverse Haar transform to the coding coefficients after inverse quantization. The memory 224 stores the attribute information values ​​of the multiple decoded three-dimensional points. For example, the attribute information of the decoded three-dimensional points stored in the memory 224 may be used to predict the undecoded three-dimensional points.

[0167] [Slicing] The encoding device may divide the three-dimensional data into one or more three-dimensional data points and encode them. A three-dimensional point cloud divided into multiple parts is called a slice. A slice is a set of points that have geometry information and attribute information.

[0168] Figure 25 is a block diagram showing an example configuration of a three-dimensional data encoding device in this case. The three-dimensional data encoding device comprises a data division unit 231 and an encoding unit 232.

[0169] The data division unit 231 generates multiple divided three-dimensional data by dividing the three-dimensional data. Each divided three-dimensional data corresponds to a slice. The encoding unit 232 generates encoded data by encoding the positional information and attribute information of each of the multiple divided three-dimensional data (slices).

[0170] Figure 26 is a block diagram showing an example configuration of a three-dimensional data decoding device in this case. The three-dimensional data decoding device comprises a decoding unit 233 and a data merging unit 234. The decoding unit 233 generates multiple divided three-dimensional data by decoding encoded data. Each divided three-dimensional data includes slice position information and attribute information. The data merging unit 234 restores the three-dimensional data by merging the multiple divided three-dimensional data.

[0171] In encoding, there may or may not be dependencies between slices. If there are no dependencies, slices can be encoded or decoded independently. Therefore, processing time can be reduced by processing multiple slices in parallel. Also, the amount of processing can be reduced by partial decoding, which decodes only some of the slices.

[0172] If dependencies exist, an identifier indicating the dependency is stored in the bitstream, and encoding or decoding is performed sequentially, starting with the dependent data (the referenced data).

[0173] The number of divisions and the division method can be any method. The 3D data encoding device may determine the shape of an object and divide the point cloud for each object, or it may divide it based on the number of points included in a slice. Alternatively, the 3D data encoding device may use map information or location information to divide the point cloud based on whether or not it is included in 3D space (tile information).

[0174] Figure 27 shows the relationship between tiles and slices. As shown in Figure 27, tiles correspond to three-dimensional space, and slices correspond to divided three-dimensional point clouds. Multiple tiles may overlap. This enables parallel processing in adaptive encoding or decoding depending on the content or object, improving the flexibility of the point cloud encoding or decoding system.

[0175] [Configuration of Encoded Data] Positional and attribute information for each slice are encoded. At least a portion of the resulting encoded data is stored in the data unit's payload. A header is also added to the payload.

[0176] If multiple attribute pieces of information exist for each point, each of the multiple attribute pieces of information is stored in a data unit. Figure 28 shows an example of the structure of a bitstream (encoded data). Figure 28 shows encoded data of a point cloud of two frames with two types of attribute pieces of information. The bitstream has a geometry (position information) data unit (Geometry Data Unit: Geom) and two attribute (attribute information) data units (Attribute Data Unit: Attr(0), Attr(1)) for each slice. For example, if each point has two attribute pieces of information, color and reflectance, the encoded color data is stored in Attr(0) and the encoded reflectance data is stored in Attr(1). The header of the geometry data unit stores the slice identifier (slice_id). The header of the attribute data unit shows the slice_id of the corresponding (referenced) geometry data unit. Note that data units are sometimes called slices, and data unit headers are sometimes called slice headers.

[0177] Furthermore, metadata related to location information coding is stored in the Geometry Parameter Set (GPS). Metadata related to attribute information coding is stored in the Attribute Parameter Set (APS). Metadata common to multiple PCC frames (PCC sequences) is stored in the Sequence Parameter Set (SPS).

[0178] Each data unit and parameter set is converted to either a NAL (Network Access Layer) unit or a TLV (Type Length Value) unit, and the data stream of the unit is output.

[0179] Figure 29 is a block diagram showing an example configuration of a three-dimensional data encoding device that outputs a stream of TLV units. The three-dimensional data encoding device comprises an encoding unit 241 and a TLV storage unit 242. The encoding unit 241 generates encoded data by encoding point cloud data. The TLV storage unit 242 generates a TLV stream (a single stream of TLV units) by storing (encapsulating) the encoded data in multiple TLV units.

[0180] A unit includes a Type field indicating the data type, a Length field indicating length information, and a Value field that stores the data indicated by the Type field. Note that a unit may also be in a different format, such as one that does not include length information.

[0181] Figure 28 shows a sequence of units that store data units and parameter sets. The arrows in Figure 28 indicate dependencies related to the decoding of encoded data. The source of an arrow depends on the data at the end of the arrow, and the header of the source data shows the identifier of the data at the end of the arrow (referenced). The three-dimensional data decoding device decodes the data at the end of the arrow and uses that decoded data to decode the source data of the arrow. For example, the geometry data unit (Geom) shows the GPS ID and SPS ID corresponding to that geometry data unit. The attribute data unit (Attr) shows the APS ID corresponding to that attribute data unit.

[0182] GPS and APS may be provided one per frame. Alternatively, if the encoding scheme is changed for each slice, GPS and APS may be provided for each slice. Furthermore, GPS and APS may be shared across multiple frames (sequences).

[0183] Furthermore, APS may be shared across multiple attribute information sets. The parameter set of the referenced object is sent before the referenced object.

[0184] When interframe prediction is used, or when there are dependencies between data units between frames, a Group of Frames (GOF) containing multiple frames may be constructed. A GOF is a random access unit, and the first slice of a GOF is independent of dependencies and is the slice where decoding begins. A stream may have delimiters that indicate frame boundaries or GOF boundaries, or TLV units that indicate boundaries.

[0185] Next, we will explain the structure of the encoded data and how to store the encoded data in the NAL unit.

[0186] For example, a data format is defined for each type of encoded data. Figure 30 shows an example of encoded data and a NAL unit.

[0187] For example, as shown in Figure 30, the encoded data includes a header and a payload. The encoded data may also include length information indicating the length (data volume) of the encoded data, header, or payload. The encoded data may also be in TLV format including length information. Furthermore, the encoded data does not necessarily have to include a header.

[0188] The header includes, for example, identification information to identify the data. This identification information may indicate, for example, the data type or frame number.

[0189] The header contains, for example, identification information indicating a reference relationship. This identification information is stored in the header when there is a dependency between data, and it is information used to reference the referenced data from the source. For example, the header of the referenced data contains identification information to identify that data. The header of the referenced data contains identification information indicating the referenced data.

[0190] Furthermore, if the referenced or source can be identified or derived from other information, the identifying information for identifying the data or identifying information indicating the reference relationship may be omitted.

[0191] The three-dimensional data encoding device stores the encoded data in the payload of the NAL unit. The NAL unit header contains pcc_nal_unit_type, which is the identification information for the encoded data. Figure 31 shows an example of the semantics of pcc_nal_unit_type.

[0192] As shown in Figure 31, when pcc_codec_type is codec 1 (Codec1: first encoding method), the values ​​of pcc_nal_unit_type from 0 to 10 are the encoded position data (Geometry), encoded attribute X data (AttributeX), encoded attribute Y data (AttributeY), position PS (Geom.PS), attribute XPS (AttriX.PS), attribute YPS (AttriY.PS), position SPS (Geometry Sequence PS), attribute XSPS (AttributeX Sequence PS), attribute YSPS (AttributeY Sequence PS), AU header (AU Header), GOF header (GOF It is assigned to the Header. Values ​​11 and above are also assigned to the backup of Codec 1.

[0193] The following describes modifications of this embodiment. PS has levels, such as frame-level PS, sequence-level PS, and PCC sequence-level PS. If the PCC sequence level is considered a higher level and the frame level a lower level, the following method may be used to store the parameters.

[0194] The default PS value is shown in the higher-level PS. If the value of the lower-level PS differs from the value of the higher-level PS, the PS value is shown in the lower-level PS. Alternatively, the PS value is not listed in the higher-level PS, but is listed in the lower-level PS. Alternatively, information on whether the PS value is shown in the lower-level PS, the higher-level PS, or both is shown in either the lower-level PS or the higher-level PS, or both. Alternatively, the lower-level PS may be merged with the higher-level PS. Alternatively, if the lower-level PS and the higher-level PS overlap, the 3D data encoding device may omit the transmission of one of them.

[0195] The three-dimensional data encoding device may divide the data into slices or tiles and send out the divided data. The divided data contains information for identifying the divided data, and the parameters used for decoding the divided data are included in the parameter set. In this case, pcc_nal_unit_type is defined as an identifier indicating that it is data that stores data or parameters related to tiles or slices.

[0196] [Gaussian Data] The following describes Gaussian data (Gaussian Splatting 3D data). First, we will explain the configuration of a system that encodes or decodes three-dimensional data generated using the Gaussian Splatting method.

[0197] Figure 32 shows a part of the processing in the three-dimensional data generation system shown in Figure 1. The sensor information input unit 301 and the three-dimensional data generation unit 302 shown in Figure 32 correspond, for example, to the sensor information acquisition unit 117 and the three-dimensional data generation unit 118 shown in Figure 1.

[0198] The sensor information input unit 301 acquires point cloud data (three-dimensional point cloud data) and outputs the acquired point cloud data to the three-dimensional data generation unit 302. The three-dimensional data generation unit 302 generates Gaussian data (Gaussian Splatting data) from the point cloud data and outputs the generated Gaussian data.

[0199] Specifically, the three-dimensional data generation unit 302 first generates a grid according to the density of the input point cloud data. Each point in the point cloud data is projected onto the nearest grid. At this time, attribute information such as color, reflectance, or normal associated with the point is also linked to the grid. The three-dimensional data generation unit 302 derives parameters such as the mean and variance of the Gaussian function from the information projected onto the grid.

[0200] The method for generating Gaussian data is not limited to the above. For example, Gaussian data may be generated by machine learning using point cloud data generated from multiple two-dimensional images using Structure from Motion (SfM), etc.

[0201] Gaussian data is a collection of multiple 3D Gaussian (ellipsoid) data. Each 3D Gaussian contains data such as three-dimensional point coordinates, a 3x3 covariance matrix, color, and transparency. It is also possible to represent the above data of a 3D Gaussian by converting it into components described in the PLY format.

[0202] Figure 33 shows the components of Gaussian data described in PLY format. The Gaussian data shown in Figure 33 includes Position, Orientation, Scale, SH Coefficient, and Transparency. Transparency is also called transparency, clarity, or opacity.

[0203] Position is the three-dimensional coordinate of a three-dimensional point. Scale and Orientation are parameters of the three-dimensional covariance matrix. Furthermore, the SH coefficient is the coefficient of the spherical harmonics when color information is represented using spherical harmonics.

[0204] In other words, Gaussian data includes a reference three-dimensional coordinate (position) and information associated with that three-dimensional coordinate (rotation, scale, SH coefficient, transparency) for each 3D Gaussian.

[0205] Note that the structure of Gaussian data is not limited to the above. For example, Gaussian data does not have to include a covariance matrix. For example, Gaussian data may include a position (Position) indicating the three-dimensional coordinates of a three-dimensional point, a scale (Scale) indicating the scale along each axis, a color (Color) indicating RGB, and a transmittance (Transparency).

[0206] Furthermore, Gaussian data may include multiple cubes (voxels) that divide 3D space, instead of multiple ellipsoids. Each voxel has attribute information. For example, each voxel may include a position indicating the three-dimensional coordinates of the voxel's center point, attribute information such as color, reflectivity, or material, and a scale indicating the size of the voxel.

[0207] Furthermore, the position may be expressed not in three-dimensional coordinates, but in two-dimensional point coordinates (x, y). In this case, the scale may indicate the size or shape of the ellipsoid on a two-dimensional plane. Specifically, the Gaussian data may include a position indicating the location on a two-dimensional plane, a scale vector indicating the scale (Su, Sv) along each axis, and rotation information indicating the orientation of the ellipsoid.

[0208] Note that the format for describing Gaussian data is not limited to the PLY format. For example, the SPZ format or other formats can also be used to describe Gaussian data.

[0209] [Rendering of Gaussian Data] Next, the rendering process of Gaussian data will be explained. Figure 34 shows the rendering process of Gaussian data. As shown in Figure 34, the rendering unit 303 generates a 3D model by rendering the Gaussian data. This makes it possible to present the generated 3D model in applications such as 3D display devices. A 3D display device is, for example, a device that displays a stereoscopic image in space like a hologram, or a glasses-free 3D display device.

[0210] Figure 35 shows an example of the number of SH coefficients in a spherical harmonic function. The number of elements (number of SH coefficients) in a spherical harmonic function is determined according to the level of resolution it represents. For example, a higher level (a larger numerical value of the level) indicates higher resolution. For example, as the level increases, more information about high-frequency components is added.

[0211] For example, to represent level 2, the spherical harmonics have a total of nine elements (SH coefficients) from levels 0 to 2. For example, to represent level 3, the spherical harmonics have a total of sixteen elements (SH coefficients) from levels 0 to 3.

[0212] For example, the example shown in Figure 33 represents the spherical harmonics of a level 3 color, where the three elements of color (R, G, B) are each represented by 16 SH coefficients. Therefore, the spherical harmonics have 16 × 3 = 48 elements. It is also possible to define levels 4 and above.

[0213] By representing color using spherical harmonics, it is possible to represent three-dimensional color information. Figure 36 shows an example of input and output of spherical harmonics. For example, as shown in Figure 36, by inputting viewpoint information into the spherical harmonics 304 of the color, the color information (R, G, B) as seen from that viewpoint is output.

[0214] Figure 37 shows the rendering process. As shown in Figure 37, the rendering unit 305 generates a 3D model or 2D image based on Gaussian data and a specified viewpoint, viewed from that viewpoint. The generated 3D model can be displayed in applications such as VR display. The generated 2D image can also be displayed in applications such as 2D image display devices.

[0215] Furthermore, other attribute information (such as reflectance or infrared information) may be represented using spherical harmonics, not just color. In that case, the Gaussian data will include the SH coefficients of the spherical harmonics for each attribute information. Also, the level may be changed depending on the attribute information or resolution.

[0216] [Gaussian Data Encoding and Decoding] Next, the encoding and decoding processes for Gaussian data will be described. Figure 38 is a block diagram of the encoding device (encoding system) according to this embodiment. This encoding device encodes and multiplexes Gaussian data. Figure 39 is a block diagram of the decoding device (decoding system) according to this embodiment. This decoding device decodes the Gaussian data and presents a 3D model or 2D image in the application. Note that the encoding and decoding devices shown in Figures 38 and 39 are included, for example, in the three-dimensional data generation system shown in Figure 1.

[0217] The encoding device shown in Figure 38 includes an encoding unit 311 and a multiplexing unit 312. The encoding unit 311 generates encoded data (bitstream) by encoding the input Gaussian data using a predetermined encoding scheme, and outputs the generated encoded data to the multiplexing unit 312.

[0218] The multiplexing unit 312 generates multiplexed data by multiplexing the input encoded data using a predetermined multiplexing method, and outputs the generated multiplexed data. This multiplexed data is stored or transmitted.

[0219] The decoding device shown in Figure 39 includes a demultiplexing unit 321, a decoding unit 322, and an application unit 323. The demultiplexing unit 321 generates encoded data by demultiplexing the input multiplexed data using a predetermined multiplexing method, and outputs the generated encoded data to the decoding unit 322.

[0220] The decoding unit 322 generates Gaussian data by decoding the input encoded data using a predetermined encoding method (decoding method), and outputs the generated Gaussian data to the application unit 323.

[0221] The application unit 323 includes an input interface unit 324, a rendering unit 325, and a presentation unit 326. The input interface unit 324 acquires user operations. For example, based on user operations, the input interface unit 324 generates viewpoint information indicating the viewpoint entered by the user.

[0222] The rendering unit 325 generates a three-dimensional model or a two-dimensional image by rendering the input Gaussian data. For example, the rendering unit 325 generates a 3D model or a 2D image viewed from a viewpoint entered by the user, indicated by the viewpoint information. The rendering unit 325 may also generate a 3D model or a 2D image viewed from a predetermined viewpoint.

[0223] The presentation unit 326 presents (displays) the three-dimensional model or two-dimensional image generated by the rendering unit 325. Whether to present a 3D model or a 2D image may be determined by, for example, the following method. For example, the rendering unit 325 first generates a three-dimensional model and generates point coordinate information, as shown in Figure 34. After viewpoint information is input, the rendering unit 325 generates a 2D image viewed from a specified viewpoint. At that time, the rendering unit 325 may also accept information specifying a region along with the specified viewpoint. In this way, by generating a 3D model first and then generating a 2D image of the specified region, the amount of processing can be reduced compared to generating a 2D image of all regions. Therefore, the rendering speed can be increased. Alternatively, whether to present a 3D display (3D model) or a 2D image may be determined according to the processing power of the device performing the rendering. For example, in the case of a device with low processing power, the device may not generate a 2D image and may present a 3D model.

[0224] Next, we will describe the process of learning and generating Gaussian data based on multiple two-dimensional images. Figure 40 shows a detailed view of the three-dimensional data generation device 380.

[0225] The three-dimensional data generation device 380 includes an SFM processing unit 381, an initialization unit 382, ​​a Gaussian update unit 383, a two-dimensional ground truth data selection unit 384, a rendering unit 385, a loss calculation unit 386, and a Gaussian parameter optimization unit 387.

[0226] The SFM processing unit 381 executes Structure from Motion (SfM) based on the input two-dimensional images and generates point cloud data. The initialization unit 382 may use the generated point cloud data to determine the initial values ​​of the position coordinates of the Gaussian data.

[0227] The initialization unit 382 determines initial values ​​for other data elements of the Gaussian data in addition to initial values ​​for the position coordinates, using a predetermined method. These other data elements include, for example, scale, orientation, SH coefficient, and opacity. This generates the initial Gaussian data.

[0228] The rendering unit 385 renders the generated Gaussian data based on multiple viewpoint information to produce multiple rendered two-dimensional images. The viewpoint information is, for example, information indicating the viewpoint corresponding to the multiple input two-dimensional images, and may be information estimated by SfM or information input from an external source. The two-dimensional ground truth data selection unit 384 selects a two-dimensional image to be used as ground truth data from among the multiple input two-dimensional images and outputs the selected two-dimensional image as ground truth data for two-dimensional images.

[0229] The loss calculation unit 386 generates error information by comparing the rendered two-dimensional images with the ground truth data of the two-dimensional images, and calculates the loss based on the generated error information. The Gaussian parameter optimization unit 387 optimizes the Gaussian data by changing the parameter values ​​of the Gaussian data based on the loss calculated by the loss calculation unit 386. In addition to optimizing the parameters of the Gaussian data, the Gaussian parameter optimization unit 387 may, for example, add, delete, merge, and split the Gaussian data.

[0230] The Gaussian update unit 383 updates the Gaussian data based on the optimization results from the Gaussian parameter optimization unit 387. For example, the Gaussian update unit 383 may reflect the parameter update results of the Gaussian data and generate updated Gaussian data.

[0231] The three-dimensional data generation device 380 generates Gaussian data by repeatedly optimizing parameters using the Gaussian parameter optimization unit 387 and adding or deleting Gaussian data, and outputs the generated Gaussian data.

[0232] (First Embodiment) [Gaussian Data Encoding Using G-PCC] Figure 41 is a block diagram showing an example configuration of an encoding device 330 (three-dimensional data encoding device) according to this embodiment. The encoding device 330 generates a bitstream (encoded data) by encoding the first Gaussian data.

[0233] The encoding device 330 includes a first preprocessing unit 331, a second preprocessing unit 332, and a G-PCC encoding unit 333.

[0234] The first Gaussian data includes, for example, a covariance matrix. The first preprocessing unit 331 converts the first Gaussian data into second Gaussian data composed of components of a ply format. For example, the components of a ply file include three-dimensional coordinates, scale, rotation, SH coefficients of spherical harmonics, and transmittance, as shown in Figure 33.

[0235] In this example, first Gaussian data is input to the encoding device 330, and the encoding device 330 converts the first Gaussian data into second Gaussian data. However, second Gaussian data may also be input to the encoding device 330. In this case, the encoding device 330 does not need to include the first preprocessing unit 331.

[0236] The second preprocessing unit 332 converts the multiple components included in the second Gaussian data into a format that can be encoded using the encoding scheme corresponding to each component, and outputs the converted multiple components to the G-PCC encoding unit 333.

[0237] The G-PCC coding unit 333 encodes the three-dimensional coordinates using a geometry (position information) coding method (such as an octave coding method, a predictive tree coding method, or a TriSoup method) within the G-PCC coding method (Geometry-based PCC). The G-PCC coding unit 333 encodes information associated with the three-dimensional coordinates, such as rotation, scale, SH coefficient, and transmittance, using an attribute coding method (such as an LoD-based coding method (LoD attribute information coding method) or a Transform-based coding method (transformation attribute information coding method)) within the G-PCC coding method.

[0238] The G-PCC coding unit 333 includes a position information coding unit 341, a rotation coding unit 342, a scale coding unit 343, an SH coefficient coding unit 344, a transmittance coding unit 345, a metadata coding unit 346, and a multiplexing unit 347.

[0239] The position information encoding unit 341 encodes three-dimensional coordinates. These three-dimensional coordinates are reference or representative three-dimensional coordinates for the Gaussian data. For example, the three-dimensional coordinates may be the center coordinates of an ellipse that constitutes the Gaussian data, or they may be other coordinates. For example, the three-dimensional coordinates may be the origin coordinates of the grid used to generate the Gaussian data.

[0240] Furthermore, the three-dimensional coordinates are not processed by the first preprocessing unit 331, but are directly input to the second preprocessing unit 332. The second preprocessing unit 332 converts the input three-dimensional coordinates into positive integer information (three-dimensional coordinates). The scale value and offset value used during the conversion are stored in metadata such as SPS and notified to the decoding device.

[0241] The position information encoding unit 341 generates encoded data by encoding the three-dimensional coordinates output from the second preprocessing unit 332 using the G-PCC geometry encoding method. This encoded data is stored in a GDU (Geometry Data Unit) and output. The metadata necessary for decoding the GDU is stored in a GPS (Geometry Parameter Set) and output.

[0242] Furthermore, if the multiple three-dimensional coordinates of the multiple Gaussian data are sparse (density is below a predetermined threshold), the location information coding unit 341 may use a predictive tree coding scheme suitable for coding sparse three-dimensional points. Also, if the multiple three-dimensional coordinates of the multiple Gaussian data are dense (density is above a predetermined threshold), the location information coding unit 341 may use an octave tree coding scheme suitable for coding dense three-dimensional points.

[0243] The first preprocessor 331 converts the covariance matrix contained in the first Gaussian data into rotation values ​​and scale values. The first preprocessor 331 converts the color information contained in the first Gaussian data into spherical harmonics and outputs the SH coefficients in the spherical harmonics.

[0244] The second preprocessor 332 converts the data type of each data contained in the second Gaussian data. Specifically, the second preprocessor 332 converts each data into a positive integer type in order to encode it using the G-PCC coding scheme. For example, if the data is of type float, the second preprocessor 332 converts the data into a positive integer type by performing scaling and offset processing on the data. At this time, the scale value and offset value for each attribute information used in the conversion are stored in metadata such as SPS and notified to the decoding device.

[0245] Furthermore, the second preprocessing unit 332 performs format conversion of each data and mapping processing to map each data to attribute components. For example, if the maximum number of elements (dimensions) per attribute component that can be encoded with the G-PCC attribute coding scheme is 3, the second preprocessing unit 332 converts the input components into attribute components that have three dimensions (three elements) per attribute component. For example, since Scale has three elements, it is output as a single attribute component. Orientation has four elements, so it is converted into two two-dimensional attribute components. The number of elements in the SH coefficient changes for each level. For example, SH coefficients with a level exceeding 3 are converted into multiple attribute components.

[0246] In this process, the second preprocessing unit 332 generates mapping information as metadata for each attribute piece of information, indicating which attribute component each element is mapped to.

[0247] The rotation coding unit 342, scale coding unit 343, SH coefficient coding unit 344, and transmittance coding unit 345 each generate multiple coded data by coding rotation, scale, SH coefficient, and transmittance using the attribute coding method in the G-PCC coding scheme. Each of the generated coded data is stored in an ADU (Attribute Data Unit) and output. The metadata necessary for decoding the ADU is stored in an APS (Attribute Parameter Set) and output. In addition, attribute information is coded using three-dimensional coordinates.

[0248] The metadata encoding unit 346 stores conversion information, including conversion parameters used by the first preprocessing unit 331 or the second preprocessing unit 332, in metadata such as SPS or SEI. The metadata encoding unit 346 may also encode the conversion information before storing it in the metadata. The metadata may also include information on the structure of the Gaussian data, or information indicating how the Gaussian data is transmitted using G-PCC encoding.

[0249] The multiplexing unit 347 stores the encoded data (ADU and GDU) and parameter sets (APS, GPS, and SPS) in the TLV unit and transmits them as a bitstream (encoded data). The bitstream may be multiplexed using a predetermined multiplexing scheme. The bitstream may also be formatted.

[0250] [Decoding Gaussian Data Using G-PCC] Figure 42 is a block diagram showing an example configuration of a decoding device 350 (three-dimensional data decoding device) according to this embodiment. The decoding device 350 generates first Gaussian data or second Gaussian data by decoding a bitstream (encoded data). For example, the decoding device 350 decodes a bitstream generated by the encoding device 330 shown in Figure 41.

[0251] The decoding device 350 includes a G-PCC decoding unit 351, a second post-processing unit 352, and a first post-processing unit 353.

[0252] The G-PCC decoding unit 351 generates three-dimensional coordinates, rotation, scale, SH coefficient, and transmittance by decoding the bitstream using the G-PCC coding scheme. The G-PCC decoding unit 351 comprises a demultiplexing unit 361, a position information decoding unit 362, a rotation decoding unit 363, a scale decoding unit 364, an SH coefficient decoding unit 365, a transmittance decoding unit 366, and a metadata decoding unit 367.

[0253] The demultiplexing unit 361 analyzes multiple TLV units contained in the input bitstream and generates data units of encoded data such as GDU, ADU, SPS, GPS, and APS. The position information decoding unit 362, rotation decoding unit 363, scale decoding unit 364, SH coefficient decoding unit 365, transmittance decoding unit 366, and metadata decoding unit 367 decode these data units using the information in the parameter set, with each attribute component using its respective encoding scheme.

[0254] The position information decoding unit 362 decodes three-dimensional coordinates from the GDU. The rotation decoding unit 363, scale decoding unit 364, SH coefficient decoding unit 365, and transmittance decoding unit 366 decode attribute information (rotation, scale, SH coefficient, and transmittance) from the ADU. The attribute information is decoded using three-dimensional coordinates. The metadata decoding unit 367 obtains conversion information, etc., from metadata such as SPS or SEI. The metadata decoding unit 367 may also decode the conversion information from the metadata. The metadata may also include information on the composition of the Gaussian data, or information indicating the correspondence of how the Gaussian data is transmitted using G-PCC coding.

[0255] The second post-processing unit 352 generates second Gaussian data by performing an inverse transformation on the decoded data. Here, the inverse transformation is the reverse process of the transformation process performed in the second pre-processing unit 332, and is performed based on the transformation information (transformation parameters) contained in the metadata.

[0256] The first post-processing unit 353 generates the first Gaussian data by inversely transforming the second Gaussian data. Here, the inverse transformation is the reverse process of the transformation process performed by the first pre-processing unit 331. At least one of the first Gaussian data and the second Gaussian data is output. Note that the decoding device 350 does not necessarily have to include the first post-processing unit 353.

[0257] [Mapping Process] The second preprocessing unit 332 maps the Gaussian data to the attribute components of the G-PCC. The method of this mapping and the mapping information generated at that time will be explained below.

[0258] The following examples illustrate the mapping of rotation, scale, and SH coefficient, omitting the explanation of transmittance. Similar processing can be applied when other attribute information is added.

[0259] Figure 43 shows an example of mapping multiple attribute components contained in Gaussian data to attribute components of the coding scheme G-PCC. Here, we show an example where the maximum number of dimensions (number of subcomponents) that can be coded in G-PCC is 3.

[0260] In the following, attribute components of Gaussian data may be referred to as Gaussian components, and attribute components of G-PCC may be referred to as G-PCC components. Additionally, attribute components may sometimes be simply referred to as components.

[0261] For example, the Gaussian component of the scale has three subcomponents (B1 to B3) with three dimensions, and is mapped to a single G-PCC component that has three-dimensional subcomponents. In this case, the mapping information indicates that the identifier of the Gaussian component of the scale corresponds to the identifier of the G-PCC component.

[0262] Furthermore, the Gaussian component of rotation has four subcomponents (A1 to A4) with a dimension of 4, and cannot be mapped to a single G-PCC component that has three-dimensional subcomponents. Therefore, the four subcomponents of rotation (A1 to A4) are mapped to two G-PCC components that have two-dimensional subcomponents. In this case, the mapping information indicates that the identifier of the Gaussian component of rotation and the identifier indicating the subcomponent number contained in that attribute component correspond to the identifier of the G-PCC component and the identifier indicating the subcomponent number contained in that G-PCC component.

[0263] Since there are 16 SH coefficients for each level, corresponding to three-dimensional (RGB) elements, the SH coefficients have a total of 48 dimensions (C1 to C48). These 48 dimensions are mapped to 16 G-PCC components, each having three-dimensional subcomponents.

[0264] In Figure 43, "ID" is the attribute component ID, which is an identification number used to identify the attribute component of the encoded data. "ND" is the number of dimensions, which indicates the number of elements (number of subcomponents) per attribute component.

[0265] Furthermore, the method for partitioning and mapping Gaussian data is not limited to the examples above; any combination of mapping method and number of dimensions may be used.

[0266] The second preprocessor unit 332 generates mapping information in the mapping process that indicates the G-PCC component corresponding to the Gaussian component. This mapping information is stored as metadata in the header or SEI.

[0267] The decoding device decodes mapping information from the bitstream. The second post-processing unit 352 reconstructs the Gaussian data by remapping the decoded data for each G-PCC component to Gaussian data based on the mapping information.

[0268] Figure 44 shows an example of mapping information. In Figure 44, the mapping information is stored in SEI (Gaussian Data Attribute Mapping Information SEI). The mapping information includes the number of Gaussian data components. The number of Gaussian data components indicates the number of attribute information included in the Gaussian data. For example, in the example shown in Figure 43, rotation, scale, and SH coefficient are present, and the number of Gaussian data components is 3.

[0269] Furthermore, the mapping information includes, for each Gaussian component, the Gaussian data component ID, the Gaussian data type, and the number of Gaussian data dimensions.

[0270] The Gaussian component ID is an identifier used to uniquely identify a Gaussian component. The Gaussian data type indicates the type of Gaussian component. For example, a value of 0 indicates "rotation," a value of 1 indicates "scale," a value of 2 indicates "SH coefficient," and a value of 3 indicates "transparency." Note that the combinations of values ​​and types shown here are just examples, and the combinations are not limited to these.

[0271] The Gaussian data dimension indicates the number of dimensions (subcomponents) of the Gaussian components. For example, in the example shown in Figure 43, the Gaussian data dimension for rotation is 4, the Gaussian data dimension for scale is 3, and the Gaussian data dimension for SH coefficients is 48.

[0272] Furthermore, the mapping information includes G-PCC component IDs, G-PCC dimension IDs, and transformation information as information for each dimension of the Gaussian component. The G-PCC subcomponent ID corresponds to the "ID" shown in Figure 43. The G-PCC dimension ID is an identifier indicating the subcomponent number. For example, C1 in the SH coefficient shown in Figure 43 has a G-PCC dimension ID of 0, C2 has a G-PCC dimension ID of 1, and C3 has a G-PCC dimension ID of 2.

[0273] Thus, the mapping information shows the ID and dimension ID of the G-PCC component corresponding to each dimension of each component of the Gaussian data. For example, for subcomponent C46 shown in Figure 43, the G-PCC component ID = 18 and the G-PCC dimension ID = 2 are shown.

[0274] In Figure 44, an example of mapping information showing the subcomponents (dimensions) of the G-PCC corresponding to each subcomponent (dimension) of the Gaussian data is shown. However, mapping information showing the subcomponents of the Gaussian data corresponding to each subcomponent of the G-PCC may also be used.

[0275] Furthermore, while Figure 44 shows an example of mapping information indicating the correspondence between subcomponents of Gaussian data and subcomponents of G-PCC, if there are no constraints on the number of subcomponents in the G-PCC component, mapping information indicating the correspondence between the Gaussian component and the G-PCC component may be used. In other words, it is not necessary to show the correspondence between dimensions.

[0276] The conversion information shows the parameters (conversion parameters) used when converting the component data. Figure 45 shows an example of the structure of the conversion information. For example, if a scale and an offset were used in the conversion, the conversion information would include the scale value and the offset value.

[0277] For example, the transformation is performed using the scale value and offset value according to the following formula.

[0278] Converted data = Original data × Scale value + Offset

[0279] The formula used for the conversion may be a predefined formula, or multiple conversion formulas may be defined and the formula used may be switched according to predetermined conditions. Furthermore, information indicating the used conversion formula may be stored in the bitstream.

[0280] In Figure 44, transformation information is provided for each subcomponent (dimension) of the Gaussian data, but transformation information may also be provided for each component of the Gaussian data. In other words, the transformation method may be changed on a component-by-component basis, or it may be changed on a subcomponent basis.

[0281] Note that some of the syntax shown in Figure 44 may be omitted. For example, if the order of listing the rotation, scale, and SH coefficient information, and the number of dimensions are predetermined, the Gaussian data component ID, Gaussian data type, and Gaussian data dimension may be omitted.

[0282] Furthermore, each attribute information is encoded based on the mapping information to the attribute components determined by the second preprocessing unit 332. Figure 46 shows an example of the structure of encoded data when Gaussian data is encoded using G-PCC encoding.

[0283] The SPS contains, for each G-PCC component, a G-PCC attribute component ID (attr_id), a G-PCC attribute type (attribute_type), and the number of dimensions of the G-PCC attribute component (num_dimension). Here, a new G-PCC attribute type may be defined to indicate that each element of the Gaussian data is encoded.

[0284] Figure 47 shows an example of G-PCC attribute types for Gaussian data. As shown in Figure 47, for example, G-PCC attribute types may be defined that represent the rotation, scale, SH coefficient, and transmittance of the Gaussian data, respectively.

[0285] The ADU header is assigned a G-PCC attribute component ID (attr_id). The decryption device can use the G-PCC attribute component ID contained in the ADU header and the information contained in the SPS to identify which G-PCC attribute type each ADU belongs to.

[0286] Furthermore, mapping information (Gaussian data attribute mapping information SEI) may be transmitted as SEI or included in SPS. For example, SPS may include mapping information indicating the Gaussian component corresponding to the G-PCC component.

[0287] Figure 48 is a flowchart of the processing performed by the decoding device according to this embodiment. First, the decoding device acquires multiple TLV units from the bitstream, and acquires multiple data units (GDU, ADU) and multiple parameter sets (SPS, GPS, APS) from the multiple TLV units (S101).

[0288] Next, the decoding device decodes the three-dimensional coordinates using SPS, GPS, and GDU (S102). Next, the decoding device decodes the multiple attribute components described in SPS using APS, ADU, and the three-dimensional coordinates (S103).

[0289] Next, the decoding device obtains mapping information by analyzing the Gaussian data attribute mapping information SEI (S104). Next, the decoding device reconstructs the second Gaussian data from the decoded attribute components and mapping information (S105). Next, the decoding device obtains transformation information and generates the first Gaussian data by inversely transforming the second Gaussian data using the transformation information (S106).

[0290] [Other] For the sake of simplicity, the above explanation uses an example where the Gaussian data is in a single frame. However, this method can also be applied when there are multiple frames of Gaussian data. In this case, multi-frame coding in the G-PCC method may be used, or the interpretation method in the G-PCC method may be used.

[0291] The encoding device may also generate multiple partitioned data by dividing the Gaussian data into multiple regions and encode each of the multiple partitioned data. Alternatively, the encoding device may assign the partitioned data to tiles or slices in a G-PCC and encode them.

[0292] In the above explanation, an example was shown of encoding attribute information (scale, rotation, SH coefficient, transmittance, etc.) of the second Gaussian data. However, by using the method of this embodiment, it is also possible to encode other attribute information.

[0293] Next, we will describe the process of rendering Gaussian data (3DGS) containing multiple 3D Gaussians (ellipsoids), which constitute the 3D data of Gaussian Platting, using the rasterization method. In this embodiment, data containing multiple 2D Gaussians (ellipsoids), obtained by projecting 3DGS onto a 2D plane, is called 2D Gaussian data (2DGS). Rasterization is the process of converting a 3D object into a 2D image, and includes processes such as culling unnecessary 3DGS, projecting 3DGS onto a 2D plane (3D to 2D projection), and blending. In particular, the process of converting three-dimensional data into each pixel of two-dimensional data is called rasterization.

[0294] Figure 49 shows an example of the configuration of the rendering unit 385. The rendering unit 385 comprises a 3DGS calculation unit 391, a 3DGS projection unit 392, a 2DGS selection unit 393, a sorting unit 394, and a blending unit 395. The rendering unit 385 receives Gaussian data (3DGS) and viewpoint information as input, and outputs a two-dimensional image viewed from that viewpoint.

[0295] The 3DGS calculation unit 391 determines the viewpoint based on the viewpoint information and extracts a portion of the 3DGS contained within the viewing frustum of that viewpoint. The 3DGS calculation unit 391 may, for example, generate a portion of the Gaussian data (gs_vis) used for rendering from the Gaussian data (3DGS).

[0296] The 3DGS projection unit 392 projects a portion of the extracted 3DGS (gs_vis) onto a two-dimensional plane to generate a 2DGS. The 3DGS projection unit 392 may also generate information related to the 2DGS (gs2d_info), for example.

[0297] The 2DGS selection unit 393 may duplicate the 2DGS for each tile that intersects it on the two-dimensional plane. For example, the 2DGS selection unit 393 may generate a list of 2DGS for each tile (tile_GS_list) and two-dimensional bounding box information for the 2DGS (2D_bb_info), and output them to a subsequent process.

[0298] The sorting unit 394 sorts the 2DGS based on the depth to the camera. For example, the sorting unit 394 may sort the 2DGS tile by tile and determine the order from front to back or from back to front.

[0299] The blending unit 395 blends the 2DGS in order from the front to the back according to the sorting result to generate a two-dimensional image.

[0300] Figure 50 is a flowchart illustrating an example of the rendering process shown in Figure 49. The rendering unit 385 determines the viewpoint to be used for rendering based on viewpoint information (S401). The rendering unit 385 extracts a portion of the 3DGS contained within the frustum of the determined viewpoint (S402). The rendering unit 385 projects the extracted 3DGS onto a two-dimensional plane to generate 2DGS (S403). The rendering unit 385 duplicates the 2DGS on tiles that intersect with the 2DGS on the two-dimensional plane (S404). The rendering unit 385 sorts the 2DGS based on the depth to the camera (S405). The rendering unit 385 blends the 2DGS from front to back according to the sorting result to generate a two-dimensional image (S406).

[0301] In this way, the rendering unit 385 can generate a two-dimensional image viewed from a given viewpoint based on Gaussian data (3DGS) and viewpoint information.

[0302] Figure 51 shows the 3DGS extraction and projection within the viewing frustum. Figure 52 shows the syntax (example structure) of the 2DGS information gs2d_info.

[0303] Based on Figures 51 and 52, an example of a process to generate two-dimensional Gaussian data (2DGS) containing multiple two-dimensional Gaussians (ellipsoids) obtained by extracting (culling) Gaussian data (3DGS), which is a collection of multiple three-dimensional Gaussians (ellipsoids) contained in Gaussian data, for rasterization and projecting it onto a two-dimensional plane will be described. The rendering unit 385 includes a 3DGS calculation unit 391 and a 3DGS projection unit 392.

[0304] The 3DGS calculation unit 391 receives Gaussian data (3DGS) and viewpoint information (camera parameters, view information) at the determined viewpoint. Here, the viewpoint specifies the position and orientation of the camera in the three-dimensional scene and affects the viewing frustum. The viewpoint is expressed using parameters included in the viewpoint information. Viewpoint information may include parameters such as camera position in 3D space, camera orientation, field of view, projection type (perspective or orthographic), and clipping planes (distance between near and far). Depending on the application, additional parameters such as lens distortion, image center, aspect ratio, focal length, and skew may also be included.

[0305] The 3DGS calculation unit 391 derives a view frustum (visibility cone) based on viewpoint information. As shown in Figure 51, the view frustum is the range of the three-dimensional scene visible from that viewpoint, defined by the field of view from the camera position, the near plane, and the far plane.

[0306] The 3DGS calculation unit 391 determines whether each 3DGS is contained within a viewing frustum based on its position, scale, and rotation. The 3DGS calculation unit 391 extracts 3DGS contained within a viewing frustum as visible Gaussians and outputs the extracted set of 3DGS as a part of the 3DGS (gs_vis) to the next stage. On the other hand, the 3DGS calculation unit 391 determines that 3DGS not contained within a viewing frustum are invisible Gaussians and are not subject to rasterization processing, and does not extract such 3DGS. Here, the number of parts of the 3DGS (gs_vis) may be num_gs_vis. In other words, a portion of the 3DGS (gs_vis) is extracted from multiple 3DGSs and used as input for 3D to 2D projection by the 3DGS projection unit 392.

[0307] The 3DGS projection unit 392 receives a portion of the 3DGS (gs_vis) extracted by the 3DGS calculation unit 391 as input. Based on the camera parameters included in the viewpoint information, the 3DGS projection unit 392 projects the input portion of the 3DGS (gs_vis) onto a two-dimensional plane (2D plane, screen) used for rendering, and generates information about the projected two-dimensional Gaussian data (2DGS). Specifically, the 3DGS projection unit 392 projects the three-dimensional elliptical information such as the position, scale, and rotation of the 3DGS onto the two-dimensional plane, and derives the two-dimensional position coordinates (gs2d_pos), rotation (gs2d_rot), and scale (gs2d_scale) of the projected 2DGS. Furthermore, the 3DGS projection unit 392 may derive the distance or depth between the projected 2DGS and the camera and store it as the depth of the 2DGS (gs_depth). In addition, the 3DGS projection unit 392 may include the color information (information based on the coefficients of spherical harmonics) or transmittance associated with the 2DGS in the 2DGS information and output it to the subsequent 2DGS selection unit 393 (Select 2DGS).

[0308] As shown in Figure 52, the syntax (example structure) of the information about 2DGS (gs2d_info) may include, for example, num_gs_vis, and for num_gs_vis 2DGS, it may include gs2d_pos, gs2d_rot, gs2d_scale, gs_depth, gs_sh, and gs_opacity as elements. Here, gs_sh is, for example, color information based on the coefficients of spherical harmonics, and gs_opacity indicates, for example, the transmittance (opacity) of the 2DGS. An identifier (2dgs_id) may be assigned to each 2DGS. Alternatively, the order in which the 2DGS are listed may be used as the identifier, or any one element or a combination of multiple elements may be used as the identifier.

[0309] As described above, the 3DGS calculation unit 391 can derive a viewing frustum based on viewpoint information and extract a portion of the 3DGS (gs_vis) contained within the viewing frustum. Furthermore, the 3DGS projection unit 392 projects the extracted portion of the 3DGS (gs_vis) onto a two-dimensional plane, generates information related to 2DGS (gs2d_info), and outputs it to the 2DGS selection unit 393 for subsequent processing.

[0310] Figure 53 is a diagram that divides a two-dimensional plane into tiles and identifies the tiles where the projected two-dimensional Gaussian data intersects.

[0311] The 2DGS selection unit 393 of the rendering unit 385 selects tile by tile the 2D Gaussian data that contributes to rendering (blending) based on the information (2DGS information: gs2d_info) of the 2D Gaussian data projected onto the 2D plane.

[0312] Here, two-dimensional Gaussian data refers to data relating to an ellipse obtained by projecting three-dimensional Gaussian data onto a two-dimensional plane (screen), and includes at least two-dimensional position coordinates, two-dimensional scale and two-dimensional rotation, and attribute information such as depth, SH coefficient (gs_sh), and transmittance (gs_opacity) (2DGS). 2DGS information (gs2d_info) may be represented, for example, as the data structure shown in Figure 52.

[0313] A tile is a region obtained by dividing a two-dimensional plane into multiple areas. For example, if the number of columns of a tile is N and the number of rows is M, then the two-dimensional plane is divided into N × M tiles, and the total number of tiles (num_tiles) is N × M. By dividing the two-dimensional plane into multiple tiles, processing can be done in parallel on a tile-by-tile basis. For example, in the example shown in Figure 53, N = 6, M = 6, and num_tiles = 36.

[0314] The 2DGS selection unit 393 identifies the tiles that each 2DGS affects. Specifically, the 2DGS selection unit 393 estimates a region that can contribute to the blending process, including not only the inside of the 2DGS shape but also the area surrounding the 2DGS. The 2DGS selection unit 393 may represent the estimated region using, for example, a two-dimensional bounding box (2D bounding box: gs2d_bb). The size of the two-dimensional bounding box is determined, for example, based on the two-dimensional position coordinates, two-dimensional scale, and two-dimensional rotation of the 2DGS, and may be defined by the top-left and bottom-right coordinates, or by the top-left coordinate, width, and height.

[0315] Next, the 2DGS selection unit 393 determines for each tile whether or not it intersects with the two-dimensional bounding box (Figure 53). If it determines that it intersects, the 2DGS selection unit 393 determines that the tile is affected by the 2DGS and registers the 2DGS in association with that tile. Since one 2DGS can affect multiple tiles, the 2DGS selection unit 393 may register the 2DGS for each of the multiple tiles that it determined to intersect. For example, in the example shown in Figure 53, the 2DGS is registered for eight tiles.

[0316] As a result of the above processing, the 2DGS selection unit 393 outputs a 2DGS list (Tile_GS_list) and two-dimensional bounding box information (gs2d_bb) for each tile.

[0317] Figure 54 shows the syntax for the 2DGS list (Tile_GS_list) for each tile.

[0318] The syntax for the 2DGS list for each tile includes, for each tile, the number of 2DGS registered for that tile (num_gs_in_tile) and the identifier of each 2DGS registered for that tile (2dgs_id). The 2DGS selection unit 393 may store or enumerate the 2DGS list for each tile as syntax.

[0319] For example, a tile-specific 2DGS list may be represented as a 2DGS list (2DGS list) that enumerates the 2DGSs contributing to the tile, and may include syntax that stores the total number of tiles (num_tiles) and the tile identifier (tile_id) for each tile, and for each tile, enumerates the number of 2DGSs registered for that tile (num_gs_in_tile[tile_id]) and the identifier (2dgs_id) of each 2DGS registered for that tile. Each 2dgs_id may be linked to 2DGS information (gs2d_info).

[0320] Note that the 2DGS list for each tile may store 2DGS information (gs2d_info) instead of 2dgs_id. Also, 2dgs_id may be assigned to each 2DGS, or any data, or a combination of multiple data, may be used as an identifier instead of 2dgs_id.

[0321] Figure 55 shows an example of the sequence of identifiers for two-dimensional Gaussian data obtained by projecting three-dimensional Gaussian data onto a two-dimensional plane.

[0322] The sorting unit 394 of the rendering unit 385 (Figure 49) sorts the 2DGS based on the depth to the camera in order to determine the order of the subsequent blending process. The sorting unit 394 receives the tile-specific 2DGS list (Tile_GS_list, which can also be written as tile_GS_list in syntax) and the two-dimensional bounding box information (2D bounding box: gs2d_bb), which are the outputs of the 2DGS selection unit 393. For each tile, the sorting unit 394 sorts the enumeration order of the identifiers (2dgs_id) based on the depth to the camera for each tile containing multiple 2DGS in the tile-specific 2DGS list.

[0323] As shown in Figure 55, in the 2DGS list for each tile before sorting, the multiple identifiers (2dgs_id) registered for each tile (tile_index) may be listed in identifier order (id order). On the other hand, in the 2DGS list for each tile after sorting, the multiple identifiers (2dgs_id) for each tile are listed in depth order (depth sorted order). For example, if the 2dgs_id for tile_index = 0 is "2, 5, 6, 10", after sorting it will be listed as "10, 5, 6, 2". Also, if the 2dgs_id for tile_index = 1 is "5, 6, 12, 13", after sorting it will be listed as "12, 5, 6, 13".

[0324] Figure 56 shows the syntax of a 2DGS list for each tile sorted by depth.

[0325] The syntax for a tile-specific 2DGS list sorted by depth includes the total number of tiles (num_tiles) and the tile identifier (tile_id) for each tile. For each tile, it lists the number of 2DGS registered to that tile (num_gs_in_tile[tile_id]) and the multiple identifiers (2dgs_id) registered to that tile, in order of depth. Note that the data structure is the same before and after sorting; the only thing that changes with sorting is the order in which the identifiers (2dgs_id) for each tile are listed.

[0326] Figure 57 shows an example of determining pixel values ​​by blending multiple 2DGS within a tile.

[0327] The blending process is performed by the blending unit 395 of the rendering unit 385 (Figure 49). The blending unit 395 receives information on two-dimensional Gaussian data (2DGS information: gs2d_info) and a 2DGS list for each tile (tile_GS_list). Here, the 2DGS information (gs2d_info) is data relating to an ellipse obtained by projecting three-dimensional Gaussian data onto a two-dimensional plane (screen), and includes at least two-dimensional position coordinates, two-dimensional scale and rotation, and attribute information such as transparency (opacity) and SH coefficient (SH coef) (2DGS). The 2DGS list for each tile (tile_GS_list) is a list that enumerates the 2DGS identifiers (2dgs_id) registered for each tile in an order sorted by depth.

[0328] The blending unit 395 sequentially blends multiple 2DGSs for each tile in the order listed in the 2DGS list for each tile, which is sorted by depth. For each pixel in the tile, the blending unit 395 determines the color and weight of the pixel corresponding to each 2DGS, and sequentially synthesizes the pixel values ​​of the pixel using the determined color and weight.

[0329] In the example shown in Figure 57, one tile contains two 2DGSs. For example, if the sorted order is "2DGS#1, 2DGS#2", the blending unit 395 first determines the color and weight of the pixel associated with 2DGS#1 and blends them. Next, the blending unit 395 determines the color and weight of the pixel associated with 2DGS#2 and blends them. Here, the color and weight are calculated based on at least the position, scale, rotation, opacity, SH coefficient, and pixel position.

[0330] In pixels where 2DGS#1 and 2DGS#2 overlap, the blending unit 395 determines the color associated with 2DGS#1 and the color associated with 2DGS#2, respectively, and blends them according to the determined order.

[0331] Furthermore, the blending unit 395 can reduce the amount of processing required by using the bounding box of the 2DGS when determining the pixels related to the 2DGS.

[0332] Figure 58 shows the configuration of the three-dimensional data generation device 380A (GSC generator).

[0333] In volumetric rendering rasterization, the Gaussian data list (GS list) needs to be sorted based on the depth from the camera in order to apply blending. However, sorting volume data during the decoded rasterization process is time-consuming, which can lead to delays before display begins and increased rendering time.

[0334] Therefore, the three-dimensional data generation device 380A shown in Figure 58 comprises an SFM processing unit 381, an initialization unit 382, ​​a Gaussian update unit 383, a two-dimensional ground truth data selection unit 384, a rendering unit 385, a loss calculation unit 386, and a Gaussian parameter optimization unit 387. The explanation of each of these components is the same as the explanation in Figure 40.

[0335] Prior to generating and encoding Gaussian data, the three-dimensional data generation device 380A generates sort information corresponding to a specific viewport (viewport, viewport). The three-dimensional data generation device 380A may generate sort information used for blending in the viewport in response to the processing of the rendering unit 385, and output the Gaussian data and sort information to the encoding device 388.

[0336] The encoding device 388 may encode the Gaussian data and sort information input from the three-dimensional data generation device 380A into a bitstream and notify the sort information as the syntax of the SEI. The sort information included in the SEI syntax may include, for example, camera parameter information for each viewpoint and sorted key for sorting the GS list for each viewpoint.

[0337] Figure 59 shows the configuration of the three-dimensional data decoding device 396 (decoder).

[0338] The three-dimensional data decoding device 396 includes a Gaussian decoding unit 397, a three-dimensional Gaussian rasterization unit 398, and a display unit 399.

[0339] The Gaussian decoding unit 397 decodes the reconstructed Gaussian data and the sort information contained in the SEI syntax from the compressed Gaussian bitstream (including the sort information contained in the SEI), and outputs them to the three-dimensional Gaussian rasterization unit 398.

[0340] The three-dimensional Gaussian rasterization unit 398 performs rasterization based on the reconstructed Gaussian data and sort information input from the Gaussian decoding unit 397, and performs blending using the sort information. For example, if sort information corresponding to the viewport has already been obtained, the three-dimensional Gaussian rasterization unit 398 can omit the sorting process after decoding and start blending according to the sort information. As a result, the three-dimensional data decoding device 396 can quickly display the image on the display unit 399.

[0341] In this way, rendering time can be reduced by using sort information. In particular, by notifying the sort information corresponding to the first viewpoint as SEI syntax, the initial rendering delay can be reduced.

[0342] Figure 60 shows the functional blocks of the rendering unit 385A.

[0343] The rendering unit 385A includes a 3DGS calculation unit 391, a 3DGS projection unit 392, a 2DGS selection unit 393, a sorting unit 394, and a blending unit 395.

[0344] The rendering unit 385A receives Gaussian data (3DGS) as input. The rendering unit 385A also receives viewpoint information as input.

[0345] The 3DGS calculation unit 391 extracts a portion of the 3DGS contained within the viewing frustum based on the input Gaussian data (3DGS) and viewpoint information, and generates a portion of the Gaussian data (gs_vis) that represents the extracted portion of the 3DGS.

[0346] The 3DGS projection unit 392 projects a portion of the Gaussian data (gs_vis) onto a two-dimensional plane and generates 2DGS as a projection result. 2DGS is data relating to an ellipse obtained by projecting three-dimensional Gaussian data onto a two-dimensional plane (screen). 2DGS may include at least two-dimensional position coordinates, two-dimensional scale and rotation, and attribute information such as depth, SH coefficient (gs_sh), and transmittance (gs_opacity). The 2DGS projection unit 392 generates 2DGS information (2DGS information: gs2d_info) and outputs it to the 2DGS selection unit 393.

[0347] The 2DGS selection unit 393 selects 2DGS elements that contribute to rendering (blending) for each tile based on the input 2DGS information (gs2d_info), and generates a 2DGS list (tile_GS_list) and two-dimensional bounding box information (2D bounding box information: 2D_bb_info) for each tile.

[0348] Specifically, the 2DGS selection unit 393 estimates a region that can contribute to the blending process for each 2DGS, including not only the inside of the 2DGS shape but also the area surrounding the 2DGS. The 2DGS selection unit 393 may also represent the estimated region as a two-dimensional bounding box (2D bounding box: gs2d_bb).

[0349] The 2DGS selection unit 393 determines for each tile whether or not it intersects with the two-dimensional bounding box. If it determines that it intersects, it registers the 2DGS associated with that tile. Since one 2DGS can affect multiple tiles, the 2DGS selection unit 393 may register the 2DGS for each of the multiple tiles that it determines to intersect.

[0350] The sorting unit 394 sorts the 2DGS for each tile based on the depth to the camera, using the 2DGS information (gs2d_info), the 2DGS list for each tile (tile_GS_list), and the 2D bounding box information (2D_bb_info) input from the 2DGS selection unit 393, and generates sorting information.

[0351] The sorting unit 394 outputs sorting information corresponding to the order of the 2DGS sorted for each tile to the blending unit 395, and also outputs it to the outside of the rendering unit 385A.

[0352] The blending unit 395 blends the 2DGS in the sorted order for each tile based on the sorting information input from the sorting unit 394, and generates a two-dimensional image.

[0353] Figure 61 is a flowchart showing an example of the processing procedure in the rendering unit 385A.

[0354] The rendering unit 385A determines the viewpoint (S401).

[0355] The rendering unit 385A extracts a portion of the 3DGS within the viewing frustum corresponding to the determined viewpoint (S402).

[0356] The rendering unit 385A projects a portion of the extracted 3DGS onto a two-dimensional plane to generate a 2DGS (S403).

[0357] The rendering unit 385A duplicates the 2DGS to intersecting tiles on a two-dimensional plane and generates a 2DGS list (tile_GS_list) and 2D bounding box information (2D_bb_info) for each tile (S404).

[0358] The rendering unit 385A sorts the 2DGS based on the depth to the camera and generates sorting information (S405).

[0359] The rendering unit 385A blends the 2DGS in order from the front to the back to generate a two-dimensional image (S406).

[0360] The rendering unit 385A generates sorting information based on the data calculated by the sorting unit 394 and outputs the generated sorting information (S407).

[0361] Figure 62 shows an example of the syntax for rasterization information SEI (rasterizing_information_SEI).

[0362] The encoding device may include viewpoint information, a 2D bounding box (2d_boundingbox), and a 2DGS list for each tile (Tile_GS_list) (hereinafter also referred to as tile_gs_list) in the rasterized information SEI.

[0363] Here, two-dimensional Gaussian data refers to data relating to an ellipse obtained by projecting three-dimensional Gaussian data onto a two-dimensional plane (screen), and includes at least two-dimensional position coordinates, two-dimensional scale and rotation, and attribute information such as depth, SH coefficient (gs_sh), and transmittance (gs_opacity) (2DGS).

[0364] The syntax shown in Figure 62 is an example where tile_gs_list is a list enumerating the identifiers (2dgs_id) of the 2DGS. In other words, the syntax shown in Figure 62 has the same structure as the data structure and semantics of the 2DGS list for each tile.

[0365] The rasterization information SEI may include, for example, camera parameter information (camera_param_info) and the total number of tiles (num_tiles), as shown in Figure 62. Furthermore, within the loop corresponding to the total number of tiles, it may also include a tile identifier (tile_id), the number of 2DGS registered in that tile (num_gs_in_tile[tile_id]), the 2D bounding box corresponding to that tile (2d_boundingbox[tile_id]), and the identifier of the 2DGS registered in that tile (2dgs_id). In Figure 62, tile_gs_list is represented, for example, as a portion that includes the number of 2DGS registered in the tile (num_gs_in_tile) and the identifier of the 2DGS (2dgs_id) for each tile identifier (tile_id).

[0366] A 2D bounding box is information that indicates the area referenced in the blending process. A 2D bounding box may include, for example, the origin coordinates (origin_x, origin_y), width, and height.

[0367] The total number of tiles (num_tiles) is the number of tiles used in the rasterization process. For example, if a two-dimensional plane is divided into N × M tiles, then num_tiles is N × M.

[0368] The tile identifier (tile_id) is the index of the tile. The tile identifier does not necessarily have to be stored as syntax. For example, if the tile identifier is not stored as syntax, the order of the loop corresponding to the total number of tiles (num_tiles) may be used as the tile index.

[0369] The number of 2DGS registered in the tile (num_gs_in_tile) is the number of 2DGS contained in that tile. The identifier of the 2DGS (2dgs_id) is the index of the 2DGS.

[0370] Figure 63 shows an example of the syntax for camera parameter information.

[0371] Camera parameter information may include at least one of intrinsic parameters and external parameters. Camera parameter information may also include, for example, a flag (extrinsic_flag) indicating the presence or absence of external parameters and a flag (intrinsic_flag) indicating the presence or absence of intrinsic parameters.

[0372] `extrinsic_flag` is a flag for external parameters. `intrinsic_flag` is a flag for internal parameters.

[0373] In the example shown in Figure 63, when extrinsic_flag is up, it includes camera pose information, and when intrinsic_flag is up, it includes camera parameter information.

[0374] Figure 64 shows an example of posture information.

[0375] The posture information may include, for example, position coordinates _x, _y, and _z indicating the position of the viewpoint, and direction vectors _x, _y, and _z indicating the direction of the line of sight.

[0376] The position coordinate _x is the viewpoint position coordinate on the X axis. The position coordinate _y is the viewpoint position coordinate on the Y axis. The position coordinate _z is the viewpoint position coordinate on the Z axis. The direction vector _x is the line of sight direction vector on the X axis. The direction vector _y is the line of sight direction vector on the Y axis. The direction vector _z is the line of sight direction vector on the Z axis.

[0377] Figure 65 shows an example of camera parameter information.

[0378] Camera parameter information may include, for example, the angle of view, image sensor size, focal length, aperture value, shutter speed, ISO sensitivity, depth of field, and frame rate.

[0379] Figure 66 shows an example of the configuration of the rendering unit 385B used for rasterization after decoding.

[0380] The rendering unit 385B includes a 3DGS calculation unit 391, a 3DGS projection unit 392, a 2DGS selection unit 393, a sorting unit 394, and a blending unit 395.

[0381] The 3DGS calculation unit 391 extracts a portion of the 3DGS (gs_vis) contained within the viewing frustum based on the input Gaussian data (3DGS) and viewpoint information.

[0382] The 3DGS projection unit 392 projects the gs_vis extracted by the 3DGS calculation unit 391 onto a two-dimensional plane to generate two-dimensional Gaussian data (hereinafter referred to as two-dimensional Gaussian splatting (2DGS) data (gs2d_info)).

[0383] The 2DGS selection unit 393 replicates the 2DGS to intersecting tiles on a two-dimensional plane based on the input gs2d_info, and generates a list of 2DGS contributing to each tile (tile_GS_list) and two-dimensional bounding box information for the 2DGS (2D_BB_info).

[0384] The sorting unit 394 sorts the 2DGS belonging to each tile based on the depth to the camera and generates the sorted tile_GS_list.

[0385] The blending unit 395 blends the 2DGS for each tile according to the order shown in the sorted tile_GS_list to generate a two-dimensional image.

[0386] The rendering unit 385B can receive rasterized information obtained when the rasterized information SEI is decoded in the decoding device. The rasterized information includes viewpoint information (camera_param_info) and sorting information (sorted tile_GS_list and 2D_BB_info).

[0387] When using rasterization information, the rendering unit 385B can omit processing by the 2DGS selection unit 393 and the sorting unit 394, and perform blending by the blending unit 395 using the input rasterization information (viewpoint information and sorting information).

[0388] On the other hand, if the rendering unit 385B does not use rasterization information, it uses the sorting information calculated by the 2DGS selection unit 393 and the sorting unit 394 to perform blending by the blending unit 395.

[0389] Figure 67 is a flowchart showing an example of the rasterization process after decoding, which is performed by the rendering unit 385B shown in Figure 66.

[0390] The rendering unit 385B determines the viewpoint (S411).

[0391] The rendering unit 385B extracts a portion of the 3DGS contained within the viewing frustum based on the viewpoint information (S412).

[0392] The rendering unit 385B projects the extracted 3DGS onto a two-dimensional plane to generate a 2DGS (S413).

[0393] The viewpoint may be set externally, for example, by an application. Also, if the viewpoint information input from the decoding device is the initial rendering position (default) recommended by the encoding device, the rendering unit 385B can use that viewpoint information as the initial rendering position.

[0394] Subsequently, the rendering unit 385B analyzes the rasterization information and determines whether or not to use the rasterization information (S414). For example, the rendering unit 385B may determine whether or not to use the rasterization information based on whether or not rasterization information exists, or whether or not it includes rasterization information for the viewport to be rasterized.

[0395] If the rendering unit 385B determines that it will not use rasterization information (No in S414), it duplicates the 2DGS onto intersecting tiles on the two-dimensional plane (S415).

[0396] The rendering unit 385B sorts the 2DGS based on the depth to the camera (S416).

[0397] The rendering unit 385B performs blending using the sort information calculated in steps S415 and S416 (S417).

[0398] On the other hand, if the rendering unit 385B determines that it will use rasterization information (Yes in S414), it omits the processing in steps S415 and S416 and performs blending using the input rasterization information (viewpoint information, sorted tile_GS_list and 2D_BB_info) (S418).

[0399] Figure 68 shows an example of 2DGS information (gs2d_info) output by the 3DGS projection unit 392 (Project 3DGS) of the rendering unit 385. As shown in Figure 68, the 2DGS information (gs2d_info) includes 2DGS data that corresponds to the identifier (2dgs_id) of each 2DGS, and includes the two-dimensional position coordinates, two-dimensional scale and two-dimensional rotation of the 2DGS, as well as attribute information such as the SH coefficient (SH) and opacity.

[0400] Figure 69 shows an example of a tile-specific 2DGS list (tile_GS_list) included in the rasterized information (Rasterizing information) decoded by the three-dimensional data decoding device 396. As shown in Figure 69, the tile-specific 2DGS list (tile_GS_list) associates each tile with a list of 2DGS identifiers (2dgs_id) to be blended in that tile, in the order of blending. For example, in the example shown in Figure 69, "10, 5, 6, 2" is listed for tile_index = 0, and "12, 5, 6, 13" is listed for tile_index = 1.

[0401] The 2DGS information (gs2d_info) in Figure 68 becomes the input data for the blending process in step S418 of Figure 67. In contrast, the tile-specific 2DGS list (tile_GS_list) in Figure 69 is a list of 2DGS for each tile input from the three-dimensional data decoding device 396. The blending unit 395 of the rendering unit 385 combines the 2DGS information (gs2d_info) and the tile-specific 2DGS list (tile_GS_list) to associate the order of the 2DGS identifiers (2dgs_id) to be blended with the 2DGS data corresponding to those identifiers (2dgs_id) for each tile. As a result, the blending unit 395 of the rendering unit 385 can refer to the 2DGS data for each tile in the order listed in the tile-specific 2DGS list (tile_GS_list) and perform the blending process.

[0402] Furthermore, if the rendering unit 385 uses the input rasterization information and skips steps S416 and S417 in Figure 67, the 3DGS projection unit 392 (Project 3DGS) does not need to include depth in the output 2DGS information (gs2d_info). In other words, since the rendering unit 385 can obtain the blend order from the 2DGS list for each tile (tile_GS_list), it is no longer necessary to include depth in the 2DGS information (gs2d_info) on the premise that the 3DGS projection unit 392 (Project 3DGS) performs sorting based on depth (step S416).

[0403] The encoding device 388 notifies the rendering unit 385 of a tile-specific 2DGS list (tile_GS_list) as a sorted GS list (sorted GS list) via SEI, thereby eliminating the need for the rendering unit 385 to generate a GS list for each tile and to sort the GS list. This allows the rendering unit 385 to reduce its processing load and improve rendering speed (display speed).

[0404] Furthermore, the encoding device 388 notifies the information of the two-dimensional bounding box (2D bounding box) (2D_BB_info) via SEI, and the rendering unit 385 uses the two-dimensional bounding box (2D bounding box) to select the 2DGS data to be blended, and can use the selected 2DGS data for blending. This allows the rendering unit 385 to further reduce the processing load in the blending process.

[0405] Furthermore, the rendering unit 385 can continue to effectively use the sorting information for a tile even if the user moves the viewport horizontally or vertically by one tile. Specifically, the rendering unit 385 only needs to sort the Gaussian candidates contained in the newly added row or column tile, and can reuse the already notified sorting information for other tiles.

[0406] Furthermore, the rendering unit 385 may be able to reduce the computational load of sorting even when the user changes the viewport by rotating it. Generally, since the object of the user's attention is located near the camera, if the rendering unit 385 can appropriately track the object, a portion of the sorted Gaussian candidates in gs_sorted_keys can be used as pre-sort data for the tiles updated after rotation. For example, the rendering unit 385 may apply insertion sort to the rotated tiles and assist in the sorting process using a portion of the already sorted Gaussian candidates. Alternatively, the rendering unit 385 may limit the amount of rotation performed by the user and allow the continued use of a portion of the already sorted Gaussian candidates. This allows the rendering unit 385 to reduce the computational overhead required for sorting and improve performance.

[0407] Furthermore, tile_id is useful when extending the invention. For example, the encoding device 388 may be configured to transmit only tile candidates that satisfy predetermined conditions in SEI. The predetermined conditions may include, for example, the condition that the number of Gaussians included in the candidate tile (num_gs_in_candidate_tile) is greater than or equal to a minimum number. This allows the encoding device 388 to limit the amount of sorting information encoded in SEI and reduce the size of the SEI message.

[0408] Figure 70 shows an example of the syntax for rasterization information SEI (rasterizing_information_SEI). Rasterization information SEI includes at least camera parameter information (camera_param_info) corresponding to view information (view information) and a tile-specific 2DGS list (tile_gs_list). Camera parameter information (camera_param_info) may include internal and external parameters, similar to the view information described above.

[0409] The tile-specific 2DGS list (tile_gs_list) is described for multiple tiles, corresponding to the total number of tiles in the rasterization process (num_tiles). Specifically, the tile-specific 2DGS list (tile_gs_list) is described repeatedly according to num_tiles, and for each tile, it includes the tile identifier (tile_id) and the number of 2DGS contained in that tile (num_gs_in_tile[tile_id]). The tile identifier (tile_id) is the index of each tile, and the encoding device 388 does not necessarily have to notify the encoding device 388 of this index. In this case, the three-dimensional data decoding device 396 can use the repetition order (for loop order) corresponding to num_tiles as the index for each tile.

[0410] Furthermore, the tile-specific 2DGS list (tile_gs_list) enumerates the depths (2dgs_depth) corresponding to each 2DGS contained in that tile, according to the number of 2DGSs (num_gs_in_tile[tile_id]). Thus, the syntax shown in Figure 70 shows that the tile-specific 2DGS list (tile_gs_list) is a list of 2DGS depths (2dgs_depth) rather than a list of 2DGS identifiers (2dgs_id). The combination of a tile identifier (tile_id) and a 2DGS depth (2dgs_depth) may also be called a GS key (gs_key).

[0411] The rasterized information SEI may or may not include a two-dimensional bounding box (2D bounding box: 2d_boundingbox).

[0412] Figure 71 is a diagram showing the functional blocks of the rendering unit 385C that generates a two-dimensional image downstream of the three-dimensional data decoding device 396. The rendering unit 385C includes a 3DGS calculation unit 391, a 3DGS projection unit 392, a 2DGS selection unit 393, a sorting unit 394, and a blending unit 395.

[0413] The three-dimensional data decoding device 396 decodes the rasterized information SEI and inputs the rasterized information, including viewpoint information (camera_param_info), a sorted 2DGS list for each tile (tile_GS_list), and a two-dimensional bounding box (2D_BB_info), to the rendering unit 385C.

[0414] The 3DGS calculation unit 391 of the rendering unit 385C extracts a portion of the 3DGS within the viewing frustum (gs_vis) based on the input Gaussian data (3DGS) and viewpoint information, and the 3DGS projection unit 392 projects the extracted 3DGS onto a two-dimensional plane to generate 2DGS information (gs2d_info).

[0415] The 2DGS selection unit 393 identifies tiles where 2DGS intersect on a two-dimensional plane and generates a 2DGS list and a two-dimensional bounding box for each tile. The sorting unit 394 sorts the 2DGS based on the depth to the camera when input rasterization information is not used.

[0416] The blending unit 395 performs blending using a sorted 2DGS list and a two-dimensional bounding box for each tile when using input rasterization information, and performs blending using the sorting result from the sorting unit 394 when not using input rasterization information.

[0417] Figure 72 is a flowchart showing the flow of the rasterization process performed by the rendering unit 385C shown in Figure 71 after the three-dimensional data decoding device 396.

[0418] The rendering unit 385C determines the viewpoint (S421).

[0419] The rendering unit 385C extracts a portion of the 3DGS within the viewing frustum based on the determined viewpoint (S422).

[0420] The rendering unit 385C projects the extracted 3DGS onto a two-dimensional plane to generate a 2DGS (S423).

[0421] The rendering unit 385C duplicates the 2DGS onto intersecting tiles on a two-dimensional plane, and may use the two-dimensional bounding box output in this process as the notified two-dimensional bounding box if it has been notified in the rasterization information SEI (S424).

[0422] The rendering unit 385C determines whether or not to use the rasterized information input from the three-dimensional data decoding device 396, based on whether or not such rasterized information exists and whether or not such rasterized information includes sort information corresponding to the viewport to be rendered (S425).

[0423] If the rendering unit 385C determines in step S425 that it will not use the rasterization information (No in S425), it sorts the 2DGS based on the depth to the camera and obtains sort information calculated in the rasterization process (S426).

[0424] If the rendering unit 385C determines in step S425 that it will not use the rasterization information (No in S425), it performs a blending process using the sort information obtained in step S426 (S427).

[0425] If the rendering unit 385C determines in step S425 that it will use the rasterization information (Yes in S425), it skips step S426 and performs blending using the viewpoint information, the sorted 2DGS list for each tile, and the two-dimensional bounding box included in the input rasterization information (S428).

[0426] Figure 73 shows an example of 2DGS information (2DGS_info) output by the 3DGS projection unit 392 (Project 3DGS) of the rendering unit 385. As shown in Figure 73, the 2DGS information (2DGS_info) is associated with a key gs_key (a combination of tile_id and depth) and includes 2DGS data corresponding to the gs_key. The 2DGS data includes, for example, two-dimensional position coordinates, two-dimensional scale and two-dimensional rotation, and attribute information such as SH coefficient and opacity.

[0427] Figure 74 shows an example of a tile-specific 2DGS depth list (2DGS_depth list in Tile) output from the rasterized information SEI decoded by the three-dimensional data decoding device 396. As shown in Figure 74, the tile-specific 2DGS depth list enumerates the 2DGS depths (depth) to be blended in each tile, corresponding to each tile identifier (tile_id), in the order of blending. Therefore, the rendering unit 385 can sequentially determine the combination of tile_id and depth according to the depth enumeration order shown in Figure 74, and treat this combination as gs_key.

[0428] The rendering unit 385 combines the 2DGS information (2DGS_info) shown in Figure 73 with the 2DGS depth list for each tile shown in Figure 74, thereby associating the order of the gs_keys to be blended with the 2DGS data corresponding to those gs_keys for each tile. As a result, the rendering unit 385 can refer to the gs_keys for each tile in the order shown in Figure 74, acquire the 2DGS data shown in Figure 73, and perform the blending process. Step S427 here corresponds to the process in which the blending unit 395 of the rendering unit 385 performs the blending process without using rasterization information.

[0429] The rendering unit 385 can continue to effectively use the sorting information for a tile even if the user moves the viewport horizontally or vertically by one tile. Specifically, the rendering unit 385 only needs to sort the Gaussian candidates included in the newly added row or column tiles, and can reuse the already obtained sorting information for other tiles.

[0430] Furthermore, the rendering unit 385 may be able to reduce the computational load of sorting even when the user changes the viewport by rotating it. Generally, since the object of the user's attention is located near the camera, if the rendering unit 385 can appropriately track the object, it may use a portion of the sorted Gaussian candidates in gs_sorted_keys as preliminary sort data for the tiles updated after rotation. For example, the rendering unit 385 may apply insertion sort to the rotated tiles and assist the sorting process using a portion of the already sorted Gaussian candidates, or it may limit the amount of rotation operation by the user so that a portion of the already sorted Gaussian candidates can be used continuously. This allows the rendering unit 385 to reduce the computational overhead required for sorting and improve performance.

[0431] Furthermore, tile_id is useful when extending the invention. For example, the encoding device 388 may be configured to transmit only tile candidates that satisfy predetermined conditions in rasterized information SEI, and the predetermined conditions may include, for example, the condition that the number of Gaussians included in the candidate tile (num_gs_in_candidate_tile) is greater than or equal to a minimum number. This allows the encoding device 388 to limit the amount of sorting information encoded in the SEI and reduce the size of the SEI message.

[0432] Furthermore, in the method using gs_id, the 2dgs_id assigned by the encoding device 388 and the 2dgs_id assigned by the three-dimensional data decoding device 396 must be the same. For example, if the order of the Gaussian data in the encoding device 388 is the same as the order of the Gaussian data after decoding by the three-dimensional data decoding device 396, the same ID can be assigned in the same order. However, if the Gaussian data is shuffled in the subsequent stage of the encoding device 388, it becomes difficult to maintain the correspondence.

[0433] In contrast, the depth-based method maintains the correspondence between tile_id and depth, allowing blending to be performed using the above correspondence even when Gaussian data is shuffled.

[0434] Figure 75 shows an example of the syntax for sort key type SEI. The encoding device 388 (a device that transmits or generates information) can switch between using depth (gs2d_depth; also written as gs2_depth in the above embodiment) or the 2DGS index (gs2d_index) as the sort key for determining the blend order for each tile by notifying the sort key type SEI in SEI. The three-dimensional data decoding device 396 (a device that receives, renders, or displays information) decodes the sort key type SEI, and the rendering unit 385C determines the blend order for each tile based on the decoded sort key.

[0435] The sort key type SEI includes viewpoint information (camera_param_info) and the total number of tiles (num_tiles). A loop processing over the total number of tiles (num_tiles) notifies the tile identifier (tile_id), the number of Gaussian candidates contained in that tile (num_gs_in_candidate_tile[tile_id]), and the sort key type (sort_key_type) indicating the type of sort key for each tile. Furthermore, a loop processing over the number of Gaussian candidates contained in that tile (num_gs_in_candidate_tile[tile_id]) notifies the sort key for each Gaussian candidate according to the sort key type (sort_key_type).

[0436] Specifically, if the sort key type (sort_key_type) is 0, the encoding device 388 notifies the depth (gs2d_depth) of each Gaussian candidate as the sorted key. Conversely, if the sort key type (sort_key_type) is 1, the encoding device 388 notifies the 2DGS index (gs2d_index) of each Gaussian candidate as the sort key.

[0437] Thus, the sort key type SEI allows switching between a method that notifies the depth (gs2d_depth) as a sorted list of Gaussian candidates for a tile, and a method that notifies the 2DGS index (gs2d_index).

[0438] For example, when the rendering unit 385C uses a 2DGS index (gs2d_index), this method is easily applicable when the Gaussian data (GS data) is not sorted downstream of the three-dimensional data generator 380A (Generator). Furthermore, by not using depth as a sort key, the rendering unit 385C can reduce the number of bits allocated to notify the depth (gs2d_depth), and the processing time for the rasterization process performed downstream of the three-dimensional data decoding device 396 can be reduced.

[0439] In contrast, when the rendering unit 385C uses depth (gs2d_depth), this method can be applied even when the Gaussian data (GS data) is sorted downstream of the three-dimensional data generation device 380A (Generator), and it can reduce the processing time of the rasterization process performed downstream of the three-dimensional data decoding device 396.

[0440] Therefore, the encoding device 388 can utilize the advantages of both of the above depending on the situation by enabling the sort key type (sort_key_type) to be switched according to the sort key type SEI.

[0441] Figure 76 shows an example of SEI syntax for switching between multiple sort keys and notifying the system.

[0442] An encoding device (a device that transmits or generates information) may notify a decoding device (a device that receives, renders, or displays information) of sort keys (sorted_keys) used to determine the blend order as part of the rasterized information SEI. Alternatively, the decoding device may obtain the sort keys (sorted_keys) by decoding the rasterized information SEI.

[0443] The sort_multi_type_SEI() shown in Figure 76 includes viewpoint information (camera_param_info), the total number of tiles (num_tiles), and the tile identifier (tile_id), as well as the number of Gaussian candidates contained in each tile (num_gs_in_candidate_tile[tile_id]) and information indicating the sort type (sort_type).

[0444] The encoding device may, based on the sort_type_flag shown in Figure 76, select and notify the contents of the sorted keys (sorted_keys) for each candidate tile from a plurality of methods. The decoding device may, based on the sort_type_flag, interpret the contents of the sorted keys (sorted_keys) for each candidate tile from a plurality of methods.

[0445] For example, when sort_type_flag is 0, depth (gs2d_depth) is used as a sort key, and when sort_type_flag is 1, opacity (gs2d_opacity) is used as a sort key.

[0446] Further, when sort_type_flag is 2, both depth (gs2d_depth) and opacity (gs2d_opacity) are used as sort keys, and when sort_type_flag has other values, the 2DGS index (gs2d_index) is used as a sort key.

[0447] In this way, by enabling notification of the content of sorted keys (sorted_keys) as two or more elements such as depth (gs2d_depth) and opacity (gs2d_opacity), blend efficiency can be improved by using opacity (gs2d_opacity) alone or in combination with depth (gs2d_depth) in blending processing.

[0448] FIG. 77 shows the syntax of recommended re-sorting method notification.

[0449] sort_multi_type_method_SEI() shown in FIG. 77 is a syntax for enumerating, on a tile (tile) basis, a tile identifier (tile_id) and the number of Gaussian candidates included in a candidate tile (num_gs_in_candidate_tile[tile_id]), and notifying a sort key for the candidate tile and a recommended re-sorting method (sort_method). For example, the sort key is enumerated as gs2d_depth, gs2d_opacity, gs2d_depth and gs2d_opacity, or gs2d_index according to the value of sort_type_flag.

[0450] sort_method indicates a recommended sorting method for re-sorting processing executed when a viewport is changed. For example, when sort_method is 0, Timsort is recommended, and when sort_method is 1, Adaptive Shivers Sort is recommended.

[0451] In general, different sorting methods have different sorting speeds. Therefore, an apparatus that generates sorting information (such as the encoding apparatus 388) notifies the type of re-sorting method as sort_method, and an apparatus that uses the sorting information (such as the three-dimensional data decoding apparatus 396 or the rendering unit 385) selects a re-sorting method according to the notified sort_method. This allows the apparatus that uses the sorting information to execute re-sorting processing at the fastest possible sorting speed.

[0452] FIG. 78 is an example syntax of View_information_SEI including viewpoint information of a plurality of viewports.

[0453] The syntax of View_information_SEI illustrated in FIG. 78 is an example syntax of an SEI (Supplemental Enhancement Information) message configured in a format that can be referred to by a decoding apparatus or a rendering unit (display processing unit) in order to notify viewpoint information corresponding to a plurality of viewports. extrinsic_flag indicates whether or not pose information is included in the syntax; when extrinsic_flag is true, pose information is included in the syntax.

[0454] intrinsic_flag indicates whether or not camera parameter information is included in the syntax; when intrinsic_flag is true, camera parameter information is included in the syntax.

[0455] `rasterization_info_flag` is a flag that indicates whether or not rasterization information is signaled. If `rasterization_info_flag` is 1, rasterization information is included in the syntax; if `rasterization_info_flag` is 0, rasterization information is not included in the syntax.

[0456] If rasterization information is included in the syntax, the syntax includes the total number of tiles (num_tiles), and, associated with each tile identifier (tile_id), the two-dimensional bounding box (2d_boundingbox[tile_id]) and the number of 2DGS contained in the tile (num_gs_in_tile[tile_id]). The syntax may also include, for each tile, the identifiers of the 2DGS contained in that tile (2dgs_id) listed a number of times equal to num_gs_in_tile[tile_id].

[0457] Thus, the syntax includes viewpoint information in a common format for each of the multiple viewports, and may also include rasterization information corresponding to that viewport, as needed.

[0458] Figure 79 shows the syntax for sort_by_depth_fix_intersect_SEI().

[0459] The syntax for sort_by_depth_fix_intersect_SEI() shown in Figure 79 includes viewpoint information (camera_param_info) and the number of tiles (num_tiles), as well as a method for generating a two-dimensional bounding box (bb_gen_method) to estimate the influence of candidate Gaussians. The syntax for sort_by_depth_fix_intersect_SEI() shown in Figure 79 also includes the number of Gaussians contained in the candidate tiles of a given tile (num_gs_in_candidate_tile[tile_id]) associated with each tile identifier (tile_id), and lists the depth (gs2d_depth) as a list of those Gaussians.

[0460] Here, the encoding device may assume that the method for replicating the Gaussian on the intersecting tiles is constant when encoding the minimum information of the camera's viewpoint and the sorting information into the SEI. That is, the encoding device may minimize the amount of viewpoint information and sorting information included in the SEI, assuming that the method used for replicating the Gaussian on the intersecting tiles is constant.

[0461] Furthermore, the rasterizer (for example, the rendering unit 385) may be able to set the tile dimensions (tile_dim_info) and the two-dimensional bounding box generation method (bb_gen_method) in order to replicate the Gaussian to intersecting tiles. For example, the rasterizer may change the tile dimensions by changing the number of tile divisions in the two-dimensional plane (number of rows and columns of tiles), or it may select a different two-dimensional bounding box generation method as a way to estimate the influence of the candidate Gaussian.

[0462] If the rasterizer can set the tile dimensions and the method for generating the two-dimensional bounding box in this way, the encoding device may encode tile dimension information (tile_dim_info) indicating the tile dimensions and the method for generating the two-dimensional bounding box (bb_gen_method) as additional information into the SEI. This allows the decoding processing unit (e.g., rendering unit 385) to continue using the sort information (gs_sorted_keys) included in the SEI as a fixed screen division.

[0463] Furthermore, by using the same two-dimensional bounding box generation method (bb_gen_method) in both the encoding device and the decoding device (e.g., rendering unit 385), the same Gaussian candidate is obtained for each tile. This allows the decoding device (e.g., rendering unit 385) to use the sort information (gs_sorted_keys) notified by the encoding device via SEI as a consistent list of Gaussian candidates corresponding to each tile.

[0464] Figure 80 shows an example of dividing a two-dimensional plane into tiles and generating a two-dimensional bounding box based on the projected two-dimensional Gaussian data.

[0465] As shown in Figure 80, the two-dimensional plane is divided into multiple tiles, each consisting of multiple tile rows and multiple tile columns. The rendering unit 385 generates a two-dimensional bounding box based on the projected two-dimensional Gaussian data and identifies the tiles that intersect the two-dimensional bounding box as intersecting tiles affected by the two-dimensional Gaussian data.

[0466] At this time, when the encoding device notifies the two-dimensional bounding box generation method (bb_gen_method) and tile dimensions (tile_dim_info) described in Figure 79, the rendering unit 385 can generate a two-dimensional bounding box according to the notification and identify Gaussian candidates for each tile. As a result, the rendering unit 385 can use the sort information (gs_sorted_keys) notified by the encoding device in SEI as sort information corresponding to fixed screen division, and ensure consistency of Gaussian candidates used in blending.

[0467] Figure 81 is an example of a scene composed of multiple frames.

[0468] In one example of entertainment applications, a scene is captured as a sequence of multiple frames taken by an array of front-facing cameras.

[0469] In such an environment, user manipulation of the viewpoint (viewport) may be limited. For example, the user may be able to move left and right and up and down relative to the default viewport distance, and may be able to zoom in and zoom out of the viewpoint, but these operations may be limited to a predetermined range.

[0470] Figure 82 shows how the tiles change as the viewport moves and zooms.

[0471] As shown in Figure 82(a), the encoding-side processing unit (e.g., encoding device 388) and the decoding-side processing unit (e.g., rendering unit 385) may encode sorted keys for the tile corresponding to the default viewport and the surrounding tiles of that tile. For example, the encoding-side processing unit may encode the sorted keys for the surrounding tiles of the tile corresponding to the default viewport into SEI.

[0472] Furthermore, as shown in Figure 82(b), for the zoomed-in viewport, sorted keys may also be encoded for the tile corresponding to the zoomed-in viewport and the surrounding tiles of that tile. For example, the encoding processing unit may encode the sorted keys for the surrounding tiles of the tile corresponding to the zoomed-in viewport into SEI.

[0473] As described above, by encoding sorted keys for the default viewport and the surrounding tiles of the zoomed-in viewport, the decoding processing unit (e.g., rendering unit 385) does not need to re-sort the candidate Gaussians for rendering even when the user moves the viewpoint up or down or left or right, or zooms in and out. This allows the decoding processing unit to omit the sorting process when the viewpoint changes, thereby reducing the rendering load.

[0474] Figure 83 shows the syntax of reuse_prev_frame_sorting_SEI().

[0475] The syntax of reuse_prev_frame_sorting_SEI() shown in Figure 83 includes viewpoint information (camera_param_info), as well as the number of tiles (num_tiles) and each tile identifier (tile_id). Furthermore, the syntax of reuse_prev_frame_sorting_SEI() shown in Figure 83 includes a flag (reuse_previous_frame_sorting_flag) indicating whether or not to reuse the sort information (gs_sorted_key) from the previous frame at the same tile position.

[0476] As explained in Figures 81 and 82, in entertainment applications, a series of frames acquired by multiple front-facing cameras is treated as a scene, and the background may be largely static. Even when the background is largely static, the characteristics of candidate Gaussian elements within the same tile can change between frames.

[0477] The `reuse_previous_frame_sorting_flag` indicates whether the processing unit (such as the rendering unit 385) that performs rendering after decoding will use the sorted key (gs_sorted_key) obtained in the previous frame for the same tile position. For example, if `reuse_previous_frame_sorting_flag` is enabled, the processing unit that performs rendering after decoding may use the gs_sorted_key of the previous frame corresponding to the same tile position as the pre-sort input for that tile.

[0478] This allows the processing unit that performs rendering after decoding to reduce the computational overhead required for sorting compared to performing sorting for each tile from the beginning. Furthermore, the encoding device can reduce the amount of sorting information it notifies in the SEI message.

[0479] Furthermore, reuse_previous_frame_sorting_flag may be used as a flag indicating whether or not to update sort information. For example, when reuse_previous_frame_sorting_flag indicates that sort information is not to be updated, notification of sort information for the current frame may be omitted.

[0480] FIG. 84 is a diagram illustrating an example in which a user operates a viewport.

[0481] Here, the scene shown in FIG. 84 may be generated based on a sequence of frames acquired by a configuration including a plurality of front cameras in an application for entertainment use.

[0482] In such an environment, operations that the user can perform on the scene may be limited. For example, the user may move the viewport within a limited range in the left direction, right direction, up direction, or down direction from the default view port distance, and may perform zoom in or zoom out on the viewport. FIG. 84 shows examples of each operation of left panning, right panning, and zooming in.

[0483] FIG. 85 is a diagram illustrating the left-right movement and zooming of the viewport shown in FIG. 84 on a tile-divided two-dimensional plane.

[0484] As shown in FIG. 85, with respect to the default viewport, the user may select a viewport moved to the left (left pan), a viewport moved to the right (right pan), or an enlarged viewport (zoomed in). Through selection of these viewports, the user can change the display target region or display magnification within a limited operation range.

[0485] In the present embodiment, sorting information is an example of rendering information. Therefore, the sorting information described in the embodiment may be read as rendering information.

[0486] Furthermore, in this embodiment, "rasterize" may be read as "render."

[0487] Furthermore, in this embodiment, sort information and rendering information are examples of metadata.

[0488] Some of the syntax for sort information, rendering information, and rasterization information may be included in the SEI, in metadata such as the header of the encoded data, or in the system format. Furthermore, some of this syntax may be included in other metadata.

[0489] Here, the names of the metadata are not limited to sorting information, rendering information, and rasterization information. For example, any metadata containing content equivalent to the sorting information, rendering information, or rasterization information described in the embodiment may be called by other names.

[0490] In this embodiment, the sorting information, rendering information, and rasterization information were described as structures generated during the generation of 3DGS. However, the encoding device may include at least a part of the rendering or rasterization process internally, and the encoding device may generate the sorting information, rendering information, or rasterization information.

[0491] Furthermore, in the embodiment, the decoding device was described as a configuration in which it decodes metadata such as sort information, rendering information, and rasterization information, and outputs the decoded metadata. However, the decoding device may notify the application of the decoded metadata as is. Moreover, the decoding device may use the metadata information decoded by the decoding unit for decoding.

[0492] In this embodiment, sort information is described as being used to skip the selecting and sorting processes in the rendering process. However, sort information may also be used to skip only one of the selecting or sorting processes.

[0493] [Encoding Device] Figure 86 is a diagram showing an example of the configuration of an encoding device. Figure 87 is a flowchart showing an example of an encoding method using the encoding device.

[0494] The encoding device 1700 includes a circuit 1701 and a memory 1702 connected to the circuit 1701.

[0495] Circuit 1701 performs the following operations.

[0496] Circuit 1701 generates one or more three-dimensional Gaussian data based on multiple two-dimensional images (S1701). In the process of generating one or more three-dimensional Gaussian data, circuit 1701 performs a rasterization process corresponding to at least one viewpoint (S1702). Circuit 1701 outputs the rasterization parameters used in the rasterization process (S1703).

[0497] This allows subsequent processing units to reference the rasterization conditions used in the generation of three-dimensional Gaussian data. This facilitates the generation of two-dimensional images based on the same conditions and improves the reproducibility of the generation results.

[0498] For example, the rasterization parameters include at least one viewpoint information indicating at least one viewpoint, generation information for generating one or more two-dimensional Gaussian data obtained by projecting one or more three-dimensional Gaussian data onto a two-dimensional plane, and distance-related information relating to one or more distances between the viewpoint and the one or more two-dimensional Gaussian positions indicated by the one or more two-dimensional Gaussian data.

[0499] This allows for the clear sharing of rasterization processing conditions and two-dimensional Gaussian data generation conditions for each viewpoint as viewpoint information and generation information. This enables the appropriate determination of the processing order or estimation of the processing range according to the viewpoint using distance-related information, thereby stabilizing the generation of two-dimensional images.

[0500] For example, the generated information includes two-dimensional bounding box information associated with one or more two-dimensional Gaussian data points.

[0501] This allows us to understand the affected region on a two-dimensional plane as two-dimensional bounding box information. This reduces the computational cost required for selecting two-dimensional Gaussian data and makes the rasterization process more efficient.

[0502] For example, distance-related information includes the order of one or more two-dimensional Gaussian data points based on a distance of one or more.

[0503] This allows for determining the processing order of distance-based two-dimensional Gaussian data without the need for additional sorting. This reduces the processing load during rasterization and speeds up the generation of two-dimensional images.

[0504] For example, distance-related information includes identification information for one or more two-dimensional Gaussian data points.

[0505] This allows the target two-dimensional Gaussian data to be identified based on the identification information contained in the distance-related information. This preserves the correspondence between distance information and the two-dimensional Gaussian data, making it easier to reference in subsequent processing.

[0506] For example, distance-related information includes distance information contained in one or more two-dimensional Gaussian data points.

[0507] This allows distance information contained in two-dimensional Gaussian data to be directly used as distance-related information. As a result, the processing order or weighting can be determined according to the viewpoint without repeatedly calculating distances.

[0508] For example, circuit 1701 further encodes one or more three-dimensional Gaussian data and rasterization parameters into a bitstream.

[0509] This allows for the integrated transmission or recording of three-dimensional Gaussian data and rasterization parameters within the same bitstream. This suppresses errors in the mapping between three-dimensional Gaussian data and rasterization parameters, ensuring reliable use on the decoding side.

[0510] For example, rasterization parameters are included in the bitstream metadata.

[0511] This allows rasterization parameters to be treated as metadata and referenced independently of the three-dimensional Gaussian data. This enables the decryption process to easily extract the rasterization parameters and apply them to post-decryption processing.

[0512] Figure 88 is a diagram showing an example of the configuration of a decoding device. Figure 89 is a flowchart showing an example of a decoding method using the decoding device.

[0513] The decoding device 1710 includes a circuit 1711 and a memory 1712 connected to the circuit 1711.

[0514] Circuit 1711 performs the following operations.

[0515] Circuit 1711 decodes the rasterization parameters used in the rasterization process that was performed in the process of generating one or more three-dimensional Gaussian data from a bitstream, and which corresponds to at least one viewpoint (S1711). Circuit 1711 outputs the decoded rasterization parameters (S1712).

[0516] This allows the rasterization parameters decoded from the bitstream to be passed to the processing unit after decoding. This enables the generation of a two-dimensional image based on the conditions used on the encoding side, improving the reproducibility of the decoding result.

[0517] For example, the rasterization parameters include at least one viewpoint information indicating at least one viewpoint, generation information for generating one or more two-dimensional Gaussian data obtained by projecting one or more three-dimensional Gaussian data onto a two-dimensional plane, and distance-related information relating to one or more distances between the viewpoint and the one or more two-dimensional Gaussian positions indicated by the one or more two-dimensional Gaussian data.

[0518] This allows the processing entity after decoding to understand the processing conditions based on viewpoint information and the generation information of the two-dimensional Gaussian data. As a result, the processing order can be determined according to the viewpoint using distance-related information, and the computational complexity of the rasterization process can be reduced.

[0519] For example, the generated information includes two-dimensional bounding box information associated with one or more two-dimensional Gaussian data points.

[0520] This allows us to estimate the affected regions based on two-dimensional bounding box information. This reduces the computational cost required for selecting two-dimensional Gaussian data and makes the rasterization process more efficient.

[0521] For example, distance-related information includes the order of one or more three-dimensional Gaussian data points based on a distance of one or more.

[0522] This allows the processing order after decoding to be determined by utilizing the order of three-dimensional Gaussian data based on distance. This reduces the computational cost required for sorting and speeds up the generation of two-dimensional images.

[0523] For example, distance-related information includes identification information for one or more two-dimensional Gaussian data points.

[0524] This allows for the identification of target Gaussian data based on identification information and the maintenance of a correspondence with distance-related information. This simplifies the process by which the decoded processing entity references the Gaussian data and improves processing stability.

[0525] For example, distance-related information includes distance information corresponding to one or more two-dimensional Gaussian data points.

[0526] This allows distance information corresponding to two-dimensional Gaussian data to be referenced after decoding. This enables the determination or weighting of the processing order according to the viewpoint while omitting the calculation of distance.

[0527] For example, rasterization parameters are included in the bitstream metadata.

[0528] This allows for easy extraction of rasterization parameters as metadata from the bitstream. This ensures flexibility in how rasterization parameters are communicated while facilitating their application in post-decryption processing.

[0529] Figure 90 is a diagram showing an example of the configuration of a rasterization device. Figure 91 is a flowchart showing an example of a rasterization method using a rasterization device.

[0530] The rasterization device 1720 includes a circuit 1721 and a memory 1722 connected to the circuit 1721. The rasterization method generates a two-dimensional image based on one or more three-dimensional Gaussian data.

[0531] Circuit 1721 performs the following operations.

[0532] Circuit 1721 obtains rasterization parameters used in the rasterization process corresponding to at least one viewpoint (S1721). Circuit 1721 obtains one or more two-dimensional Gaussian data obtained by projecting one or more three-dimensional Gaussian data onto a two-dimensional plane (S1722). Circuit 1721 controls the blending process of one or more two-dimensional Gaussian data based on the rasterization parameters (S1723).

[0533] This allows us to obtain the rasterization parameters used on the encoding side and control the blending process of the two-dimensional Gaussian data based on those rasterization parameters. This enables us to appropriately apply blending conditions according to the viewpoint and improve the reproducibility of the generated two-dimensional image.

[0534] For example, circuit 1721 further determines whether or not to use rasterization parameters. If rasterization parameters are used, circuit 1721 performs a blending process of one or more two-dimensional Gaussian data based on the rasterization parameters. If rasterization parameters are not used, circuit 1721 performs a blending process of one or more two-dimensional Gaussian data based on the depth information of at least one or more two-dimensional Gaussian data.

[0535] This allows the system to determine whether or not to use rasterization parameters and switch the execution conditions for blending depending on whether or not rasterization parameters are used. As a result, if rasterization parameters are available, processing based on those rasterization parameters is applied, and if rasterization parameters are not available, the generation of a two-dimensional image can continue using processing based on depth information.

[0536] Figure 92 is a block diagram showing an example of a device that generates format data from data. This device includes an encoding unit 1143 and a formatting unit 1144.

[0537] The encoding unit 1143 encodes data consisting of moving images, audio, three-dimensional data (point cloud, mesh, 3D Gaussian splatting, NeRF), neural network models, metadata, SEI, etc., according to a predetermined encoding scheme and generates encoded data. The encoding unit 1143 outputs the generated encoded data as an encoded data bitstream to the formatting unit 1144. In addition to encoded video and audio data, the encoded data may also include metadata such as parameter sets, control information, and SEI. The encoded data may be stored in an encoding unit, which is treated as a processing unit within the encoded data bitstream.

[0538] The formatting unit 1144 formats the input encoded data according to a predetermined system format and outputs it as format data. The formatting unit 1144 multiplexes the encoded data according to, for example, a file format for storage compliant with ISOBMFF or a packet format for transmission compliant with RTP, and can store multiple types of media and related metadata, such as video, audio, three-dimensional data (point cloud, mesh, 3D Gaussian splatting, NeRF), subtitles, application data, file information, reference time information, sensor information, and camera information, in the same format data as needed. The format data may be transmitted externally via input / output means equipped with a communication interface or user interface, or it may be stored in internal memory or storage means.

[0539] Figure 93 is a block diagram showing an example of a device for restoring original data from formatted data.

[0540] The reverse formatting unit 1145 analyzes the input format data according to a predetermined system format, extracts the encoded data stored in the format data, and outputs it as encoded data. The reverse formatting unit 1145 may also extract encoded data according to formats such as ISOBMFF, MPEG-DASH, MMT, AVI, MPEG-2 TS Systems, RTP, glTF, and USD.

[0541] The decoding unit 1146 decodes the encoded data input from the inverse formatting unit 1145 according to a predetermined encoding scheme and outputs data consisting of video, audio, three-dimensional data (point cloud, mesh, 3D Gaussian splatting, NeRF), neural network models, metadata, etc. The output data may be used by applications.

[0542] Figure 94 is a conceptual diagram illustrating an example of how an encoded data bitstream is stored in a system format.

[0543] As shown in this figure, the encoded data bitstream includes not only encoded video and audio data, but also supplementary information such as parameter sets and SEI. The encoded data, parameter sets, and SEI are grouped together as encoding units and treated as NAL units or data units. Examples of these encoding units include the Video NAL unit in video encoding, the TLV unit in G-PCC, and the V3C unit for storing video-based volumetric media, but other types of units and data units may also be used. These encoding units are stored in a system format compliant with ISOBMFF or a transmission format compliant with RTP. Furthermore, the metadata (parameter sets, control information, SEI) of the encoded data described in this embodiment, a part of the metadata, the syntax structure, and the header of the encoded data may be defined as a box structure. These boxes may be contained in the Movie box "moov" or in the Media Data box "mdat". If the metadata is per frame, a new metadata track may be defined and the metadata may be stored in sample entries within the metadata track. Furthermore, the encoded video and three-dimensional data described in this embodiment may also be defined as a box structure and may be stored in an mdat box as a sample or subsample. The header portion may be stored in a moov box and the payload portion in an mdat box. In addition, the SEI described in this embodiment may also be defined as a box structure and may be contained within a moov box or an mdat box. The encoded data may be stored in units, or parts of units may be extracted and stored, or the syntax structure may be modified as needed. Note that the above is an example of storing encoded data in ISOBMFF, and the same approach can be applied when storing encoded data in other formats.

[0544] Figure 95 shows an example of the box structure of ISOBMFF.

[0545] ISOBMFF is an ISO-based media file format specified in ISO / IEC 14496-12, and is a media-independent file format standard for multiplexing and storing various media such as video, audio, and text. In ISOBMFF, a file is composed of multiple boxes, each box consisting of type, length, and data. This embodiment shows an example in which a File type box "ftyp", a Movie box "moov" that stores metadata such as control information, and a Media Data box "mdat" that stores media data such as encoded data are used. The ftyp box indicates the file brand using 4CC and shows compatibility, while the moov box contains a track box that shows media-specific information, including an hdlr box indicating the media type, an stsd box indicating decoding parameters, a tref box indicating reference relationships, and an stco box indicating the data storage location. The mdat box stores samples and subsamples of encoded data corresponding to each media track in the moov box. A sample is a unit of encoded data corresponding to an access unit or a frame at the same time, and a subsample is segmented data corresponding to a part of an access unit or frame. Note that the method of storing each media in ISOBMFF is specified separately; for example, the storage method for AVC video and HEVC video is specified in ISO / IEC 14496-15, and the storage method for three-dimensional data is specified in ISO / IEC 23090-10 and ISO / IEC 23090-18.

[0546] The above describes the encoding device (three-dimensional data encoding device) and decoding device (three-dimensional data decoding device), etc., according to embodiments and modifications of the present disclosure. However, the present disclosure is not limited to these embodiments.

[0547] Furthermore, each processing unit included in the encoding device and decoding device, etc., according to the above embodiment is typically implemented as an LSI, which is an integrated circuit. These may be individually integrated into a single chip, or some or all of them may be integrated into a single chip.

[0548] Furthermore, integrated circuit implementation is not limited to LSIs; it may also be achieved using dedicated circuits or general-purpose processors. Alternatively, an FPGA (Field Programmable Gate Array), which can be programmed after LSI manufacturing, or a reconfigurable processor capable of reconfiguring the connections and settings of circuit cells within the LSI, may be used.

[0549] Furthermore, in each of the above embodiments, each component may be implemented by being composed of dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory.

[0550] Furthermore, this disclosure may be implemented as an encoding method (three-dimensional data encoding method) or a decoding method (three-dimensional data decoding method), etc., performed by an encoding device (three-dimensional data encoding device) and a decoding device (three-dimensional data decoding device), etc.

[0551] Furthermore, this disclosure may be implemented as a program that causes a computer, processor, or device to execute the above-described encoding or decoding method. Alternatively, this disclosure may be implemented as a bitstream generated by the above-described encoding method. Furthermore, this disclosure may be implemented as a recording medium on which the program or the bitstream is recorded. For example, this disclosure may be implemented as a non-temporary, computer-readable recording medium on which the program or the bitstream is recorded.

[0552] Furthermore, the division of functional blocks in the block diagram is just one example; multiple functional blocks can be implemented as a single functional block, a single functional block can be divided into multiple parts, or some functions can be moved to other functional blocks. In addition, the functions of multiple functional blocks with similar functions can be processed in parallel or time-sharing by a single piece of hardware or software.

[0553] Furthermore, the order in which each step in the flowchart is performed is illustrative for the purpose of specifically illustrating this disclosure, and may be in a different order. Also, some of the above steps may be performed simultaneously (in parallel) with other steps.

[0554] Although one or more embodiments of encoding and decoding devices, etc., have been described above based on embodiments, this disclosure is not limited to these embodiments. Without departing from the spirit of this disclosure, various modifications that a person skilled in the art could conceive of are applied to these embodiments, and forms constructed by combining components from different embodiments may also be included within the scope of one or more embodiments.

[0555] This disclosure is applicable to encoding and decoding devices.

[0556] 10 Encoding device 11, 21 Processor 12, 22 Memory 20 Decoding device 101 Three-dimensional data encoding system 102 Three-dimensional data decoding system 103 Sensor terminal 104 External connection unit 111 Three-dimensional data generation system 112 Presentation unit 113 Encoding unit 114 Multiplexing unit 115 Input / output unit 116 Control unit 117 Sensor information acquisition unit 118 Three-dimensional data generation unit 121 Sensor information acquisition unit 122 Input / output unit 123 Demultiplexing unit 124 Decoding unit 125 Presentation unit 126 User interface 127 Control unit 130 First encoding unit 131 Location information encoding unit 132 Attribute information encoding unit 133 Additional information encoding unit 134 Multiplexing unit 140 First decoding unit 141 Demultiplexing unit 142 Location information decoding unit 143 Attribute Information Decoding Unit 144 Additional Information Decoding Unit 150 Second Encoding Unit 151 Additional Information Generation Unit 152 Position Image Generation Unit 153 Attribute Image Generation Unit 154 Video Encoding Unit 155 Additional Information Encoding Unit 156 Multiplexing Unit 160 Second Decoding Unit 161 Demultiplexing Unit 162 Video Decoding Unit 163 Additional Information Decoding Unit 164 Position Information Generation Unit 165 Attribute Information Generation Unit 171 Tree Encoding Unit 172 Prediction Tree Encoding Unit 173 Tree Decoding Unit 174 Prediction Tree Decoding Unit 181 Tree Generation Unit 182 Geometric Information Calculation Unit 183 Encoding Table Selection Unit 184 Entropy Encoding Unit 191 Tree Generation Unit 192 Geometric Information Calculation Unit 193 Encoding Table Selection Unit 194 Entropy Decoding Unit 201 201 LoD attribute information coding unit 202 Transform attribute information coding unit 203 LoD attribute information decoding unit 204 Transform attribute information decoding unit 211 Sorting unit 212 Haar transform unit 213 Quantization unit 214 Inverse quantization unit 215 Inverse Haar transform unit 216 Memory 217 Arithmetic coding unit 221 Arithmetic decoding unit 222 Inverse quantization unit 223 Inverse Haar transform unit 224 Memory 231 Data partitioning unit 232 Coding unit 233 Decoding unit 234 Data merging unit 241 Coding unit 242 TLV storage unit 301 Sensor information input unit 302 Three-dimensional data generation unit 303 Rendering unit 304 Spherical harmonics 305 Rendering unit 311 Coding unit 312 Multiplexing unit 321 Inverse multiplexing unit 322 Decoding Unit323 Application Unit 324 Input Interface Unit 325 Rendering Unit 326 Presentation Unit 330 Encoding Device 331 First Preprocessing Unit 332 Second Preprocessing Unit 333 G-PCC Encoding Unit 335A Rendering Unit 341 Position Information Encoding Unit 342 Rotation Encoding Unit 343 Scale Encoding Unit 344 SH Coefficient Encoding Unit 345 Transparency Encoding Unit 346 Metadata Encoding Unit 347 Multiplexing Unit 350 Decoding Device 351 G-PCC Decoding Unit 352 Second Postprocessing Unit 353 First Postprocessing Unit 361 Demultiplexing Unit 362 Position Information Decoding Unit 363 Rotation Decoding Unit 364 Scale Decoding Unit 365 SH Coefficient Decoding Unit 366 Transparency Decoding Unit 367 Metadata Decoding Unit 371 Second Preprocessing Unit 372 G-PCC Encoding Unit 373 G-PCC decoding unit 374 Second post-processing unit 375 Conversion unit 376 Mapping unit 380 Three-dimensional data generation device 380A Three-dimensional data generation device 381 SFM processing unit 382 Initialization unit 383 Gaussian update unit 384 Two-dimensional ground truth data selection unit 385 Rendering unit 385A Rendering unit 385B Rendering unit 385C Rendering unit 386 Loss calculation unit 387 Gaussian parameter optimization unit 388 Encoding device 391 DGS calculation unit 392 DGS projection unit 393 DGS selection unit 394 Sorting unit 395 Blending unit 396 Three-dimensional data decoding device 397 Gaussian decoding unit 398 Three-dimensional Gaussian rasterization unit 399 Display 400 Encoding device 401 Second pre-processing unit 402 G-PCC coding unit 403 V3C coding unit 404 Metadata coding unit 405 Multiplexing unit 410 Decoding unit 411 Demultiplexing unit 412 G-PCC decoding unit 413 V3C decoding unit 414 Metadata decoding unit 415 Second post-processing unit 1143 Coding unit 1144 Formatting unit 1145 Deformatting unit 1146 Decoding unit 1700 Coding unit 1701 Circuit 1702 Memory 1710 Decoding unit 1711 Circuit 1712 Memory 1720 Rasterization unit 1721 Circuit 1722 Circuit 1722 Memory

Claims

1. A three-dimensional data generation method performed by a three-dimensional data generation device, comprising: generating one or more three-dimensional Gaussian data based on a plurality of two-dimensional images; performing a rasterization process corresponding to at least one viewpoint in the process of generating the one or more three-dimensional Gaussian data; and outputting the rasterization parameters used in the rasterization process.

2. The method for generating three-dimensional data according to claim 1, wherein the rasterization parameter includes at least one of: viewpoint information indicating the at least one viewpoint; generation information for generating one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane; and distance-related information relating to one or more distances between the at least one viewpoint and the one or more two-dimensional Gaussian positions indicated by the one or more two-dimensional Gaussian data.

3. The method for generating three-dimensional data according to claim 2, wherein the generated information includes two-dimensional bounding box information associated with each of the one or more two-dimensional Gaussian data.

4. The method for generating three-dimensional data according to claim 2 or 3, wherein the distance-related information includes the order of the one or more two-dimensional Gaussian data based on the one or more distances.

5. The method for generating three-dimensional data according to claim 2 or 3, wherein the distance-related information includes identification information for each of the one or more two-dimensional Gaussian data.

6. The method for generating three-dimensional data according to claim 2 or 3, wherein the distance-related information includes distance information contained in the one or more two-dimensional Gaussian data.

7. The method for generating three-dimensional data according to claim 2 or 3, further comprising encoding the one or more three-dimensional Gaussian data and the rasterization parameters into a bitstream.

8. The method for generating three-dimensional data according to claim 7, wherein the rasterization parameters are included in the metadata of the bitstream.

9. A three-dimensional data decoding method performed by a three-dimensional data decoding device, comprising decoding rasterization parameters used in a rasterization process performed in a process of generating one or more three-dimensional Gaussian data from a bitstream, corresponding to at least one viewpoint, and outputting the decoded rasterization parameters.

10. The three-dimensional data decoding method according to claim 9, wherein the rasterization parameter includes at least one of viewpoint information indicating the at least one viewpoint, generation information for generating one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane, and distance-related information relating to one or more distances between the at least one viewpoint and the one or more two-dimensional Gaussian positions indicated by the one or more two-dimensional Gaussian data.

11. The method for decoding three-dimensional data according to claim 10, wherein the generated information includes two-dimensional bounding box information associated with each of the one or more two-dimensional Gaussian data.

12. The three-dimensional data decoding method according to claim 10 or 11, wherein the distance-related information includes the order of the one or more three-dimensional Gaussian data based on the one or more distances.

13. The method for decoding three-dimensional data according to claim 10 or 11, wherein the distance-related information includes identification information for each of the one or more two-dimensional Gaussian data.

14. The three-dimensional data decoding method according to claim 10 or 11, wherein the distance-related information includes distance information corresponding to one or more two-dimensional Gaussian data.

15. The three-dimensional data decoding method according to claim 10 or 11, wherein the rasterization parameters are included in the metadata of the bitstream.

16. A rasterization method performed by a rasterization device that generates a two-dimensional image based on one or more three-dimensional Gaussian data, comprising: obtaining rasterization parameters used in the rasterization process corresponding to at least one viewpoint; obtaining one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane; and controlling the blending process of the one or more two-dimensional Gaussian data based on the rasterization parameters.

17. The rasterization method according to claim 16, further comprising determining whether or not to use the rasterization parameters, performing a blending process of the one or more two-dimensional Gaussian data based on the rasterization parameters if the rasterization parameters are used, and performing a blending process of the one or more two-dimensional Gaussian data based on the depth information of at least one or more two-dimensional Gaussian data if the rasterization parameters are not used.

18. A three-dimensional data generation device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, generates one or more three-dimensional Gaussian data based on a plurality of two-dimensional images, performs a rasterization process corresponding to at least one viewpoint in the process of generating the one or more three-dimensional Gaussian data, and outputs the rasterization parameters used in the rasterization process.

19. A three-dimensional data decoding device comprising a circuit and a memory connected to the circuit, wherein the circuit, in operation, generates one or more three-dimensional Gaussian data from a bitstream, decodes rasterization parameters used in rasterization processing corresponding to at least one viewpoint, and outputs the decoded rasterization parameters.

20. A rasterization device that generates a two-dimensional image based on one or more three-dimensional Gaussian data, comprising: a circuit; and a memory connected to the circuit, wherein the circuit, in operation, acquires rasterization parameters used in rasterization processing corresponding to at least one viewpoint; acquires one or more two-dimensional Gaussian data obtained by projecting the one or more three-dimensional Gaussian data onto a two-dimensional plane; and controls blending processing of the one or more two-dimensional Gaussian data based on the rasterization parameters.