Encoding device, decoding device, encoding method, and decoding method

By encoding multiple networks of the 3D data generation model and utilizing metadata compression technology, the problem of low efficiency in 3D data storage and transmission was solved, resulting in a reduction in storage capacity and transmission volume.

CN121925682APending Publication Date: 2026-04-24PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
Filing Date
2024-10-09
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing technologies, the storage capacity and transmission volume of 3D data are large, making it difficult to compress effectively, resulting in low storage and transmission efficiency.

Method used

The encoding device encodes multiple networks that constitute the 3D data generation model, and outputs a bit stream containing first metadata and second metadata. The first metadata and second metadata contain parameters representing the number of network layers and nodes, thereby achieving effective compression of the 3D data.

Benefits of technology

This reduces the storage capacity and transmission volume of 3D data, and improves storage and transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121925682A_ABST
    Figure CN121925682A_ABST
Patent Text Reader

Abstract

An encoding device (1570) is provided with a circuit (1571) and a memory (1572) connected to the circuit (1571), the circuit (1571), during operation, acquires a plurality of networks constituting a three-dimensional data generation model, encodes the plurality of networks, and outputs a bit stream including the plurality of encoded networks, the bit stream including: (1) first metadata that stores parameters common in a sequence; and (2) second metadata storing the frame, the access unit, or a parameter common to the plurality of frames.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to encoding devices, decoding devices, encoding methods, and decoding methods. Background Technology

[0002] Devices and services utilizing 3D data are expected to become increasingly common in a wide range of fields, including computer vision, mapping information, surveillance, infrastructure inspection, and image distribution, which enable autonomous movement of vehicles or robots. 3D data is acquired through various methods, such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.

[0003] As one method of representing 3D data, there is the so-called point cloud method, which uses a group of points in 3D space to represent the shape of a 3D structure. For point clouds, the position and color of the point groups are preserved. It is envisioned that point clouds will become the mainstream method of representing 3D data, but the data volume of point groups is extremely large. Therefore, in the accumulation or transmission of 3D data, similar to 2D moving images (for example, MPEG-4 AVC or HEVC, which are standardized through MPEG), data compression based on encoding is necessary.

[0004] In addition, point cloud compression is partially supported by publicly available libraries that perform point cloud association processing (such as Point Cloud Library).

[0005] In addition, there are known technologies that use three-dimensional map data to retrieve and display facilities located around a vehicle (for example, see Patent Document 1).

[0006] Existing technical documents Patent documents Patent Document 1: International Publication No. 2014 / 020663 Non-patent literature Non-patent document 1: ISO / IEC 15938-17:2022 (Information technology - Multimedia content description interface - Part 17: Compression of neural networks for multimedia content description and analysis (https / / www.iso.org / standard / 78480.html)) Summary of the Invention

[0007] The problem that the invention aims to solve The purpose of this disclosure is to provide encoding devices, etc., that can reduce the storage capacity of three-dimensional data or reduce the amount of three-dimensional data transmitted.

[0008] Methods for solving problems An encoding apparatus according to one aspect of the present disclosure includes: a circuit; and a memory connected to the circuit; wherein the circuit acquires a plurality of networks constituting a three-dimensional data generation model during operation, encodes the plurality of networks, and outputs a bit stream containing the encoded plurality of networks, the bit stream including: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames.

[0009] One type of decoding apparatus disclosed herein includes at least one of the first metadata and the second metadata containing a first parameter representing the number of layers in a first network of the plurality of networks and a second parameter representing the number of nodes in each layer.

[0010] Furthermore, these general or specific technical solutions can be implemented through systems, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or through any combination of systems, methods, integrated circuits, computer programs, and recording media.

[0011] Invention Effects The encoding device disclosed herein can reduce the storage capacity of three-dimensional data or reduce the amount of three-dimensional data transmitted. Attached Figure Description

[0012] Figure 1 This is a diagram illustrating an example of the configuration of a three-dimensional data encoding and decoding system according to Embodiment 1.

[0013] Figure 2 This is a diagram showing the composition of the point group data in Implementation Method 1.

[0014] Figure 3 This is a diagram illustrating an example of the structure of a data file that describes information about point group data in Implementation 1.

[0015] Figure 4 This is a diagram showing the composition of the three-dimensional mesh data in Implementation Method 1.

[0016] Figure 5 This is a diagram illustrating an example of the structure of a data file containing information about three-dimensional mesh data in Implementation 1.

[0017] Figure 6 This is a diagram used to illustrate the three-dimensional model in Implementation Method 1.

[0018] Figure 7This is a diagram representing the types of three-dimensional data in Implementation Method 1.

[0019] Figure 8 This is a diagram used to illustrate the encoding process of three-dimensional data in Implementation Method 1.

[0020] Figure 9 This is a diagram used to illustrate the decoding process of three-dimensional data in Implementation Method 1.

[0021] Figure 10 This is a two-dimensional schematic diagram representing the tiles and slices of the three-dimensional data in Implementation Method 1.

[0022] Figure 11 This is a block diagram illustrating an example of the functional configuration of the server and terminal in Implementation Method 1.

[0023] Figure 12 This is a block diagram illustrating another example of the data generation unit of the server in Implementation Method 1.

[0024] Figure 13 This is a diagram used to illustrate the relationship between the three-dimensional space and the encoded data in Implementation Method 1.

[0025] Figure 14 This is a diagram illustrating an example of the syntax of the encoding method unit in Implementation Method 1.

[0026] Figure 15 This is a diagram illustrating an example of the syntax of the encoded point group in Implementation Method 1.

[0027] Figure 16 This is a diagram illustrating an example of the syntax of the encoding grid in Implementation 1.

[0028] Figure 17 This is a diagram illustrating an example of the syntax for encoding a three-dimensional model in Implementation Method 1.

[0029] Figure 18 This is a diagram illustrating an example of the syntax for three-dimensional data information in Implementation Method 1.

[0030] Figure 19 This is a diagram used to illustrate the data structure of the coded point group in Implementation Method 1.

[0031] Figure 20 This is a diagram used to illustrate the data structure of the coded grid in Implementation 1.

[0032] Figure 21 This is a diagram used to illustrate the data structure of the encoded three-dimensional model in Implementation Method 1.

[0033] Figure 22 This is a diagram illustrating an example of multiple three-dimensional spaces in Implementation Method 1, presented in two dimensions.

[0034] Figure 23 This is a diagram showing an example of a bounding box in Implementation 1.

[0035] Figure 24 This is a diagram illustrating an example of the syntax for three-dimensional spatial information in Implementation 1.

[0036] Figure 25 This is a flowchart illustrating an example of partial decoding in Implementation 1.

[0037] Figure 26 This is a diagram illustrating an example of a three-dimensional spatial region that becomes part of the decoded object in Implementation 1.

[0038] Figure 27 This is a diagram illustrating an example of the data structure of the partially decoded group of encoded points in Implementation 1.

[0039] Figure 28 This is a diagram illustrating an example of the data structure of the partially decoded encoded grid in Implementation 1.

[0040] Figure 29 This is a diagram illustrating an example of the data structure of the partially decoded encoded 3D model in Implementation 1.

[0041] Figure 30 This is a diagram illustrating an example of the configuration of the decoding device in Embodiment 1.

[0042] Figure 31 This is a flowchart illustrating an example of a decoding method of the decoding device in Embodiment 1.

[0043] Figure 32 This is a flowchart illustrating another example of a decoding method using a decoding device.

[0044] Figure 33 This is a diagram illustrating an example of the configuration of an encoding device.

[0045] Figure 34 This is a flowchart illustrating an example of an encoding method using an encoding device.

[0046] Figure 35 This is a diagram illustrating the processing during the learning of the three-dimensional data generation model in Implementation Method 2.

[0047] Figure 36 This diagram illustrates the process of generating still images of a subject from any viewpoint using a three-dimensional data generation model in Implementation 2.

[0048] Figure 37 This is a diagram illustrating the motion image generation method using the three-dimensional data generation model of Example 1 in Embodiment 2.

[0049] Figure 38 This is a diagram illustrating a first example of the configuration of the encoding device in Embodiment 1 of Embodiment 2.

[0050] Figure 39 This is a diagram illustrating a first example of the configuration of the decoding device in Embodiment 1 of Embodiment 2.

[0051] Figure 40 This is a diagram illustrating a second example of the configuration of the encoding device in Embodiment 1 of Embodiment 2.

[0052] Figure 41 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 1 of Embodiment 2.

[0053] Figure 42 This is a diagram illustrating the motion image generation method using the extended three-dimensional data generation model of Embodiment 2 in Implementation 2.

[0054] Figure 43 This is a diagram illustrating a first example of the configuration of the encoding device in Embodiment 2 of Implementation 2.

[0055] Figure 44 This is a diagram illustrating a first example of the configuration of the decoding device in Embodiment 2 of Implementation 2.

[0056] Figure 45 This is a diagram illustrating a second example of the configuration of the encoding device in Embodiment 2 of Implementation 2.

[0057] Figure 46 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 2 of Implementation 2.

[0058] Figure 47 This is a diagram illustrating a motion image generation method using an extended three-dimensional data generation model, which is used to explain a variation of Embodiment 2.

[0059] Figure 48 This is a diagram illustrating a motion image generation method using a three-dimensional data generation model, which is used to explain a variation of Embodiment 2.

[0060] Figure 49 This is a diagram illustrating an example of the configuration of the encoding device in Embodiment 2.

[0061] Figure 50 This is a flowchart illustrating an example of the encoding method of the encoding device in Embodiment 2.

[0062] Figure 51 This is a diagram illustrating an example of the configuration of the decoding device in Embodiment 2.

[0063] Figure 52This is a flowchart illustrating an example of a decoding method of the decoding device in Embodiment 2.

[0064] Figure 53 This is a diagram illustrating an example of the configuration of an encoding device.

[0065] Figure 54 This is a diagram illustrating an example of the configuration of a decoding device.

[0066] Figure 55 This is a block diagram illustrating an example of the structure of the encoding device in Embodiment 3.

[0067] Figure 56 This is a block diagram illustrating an example of the structure of the decoding device in Embodiment 3.

[0068] Figure 57 This is a block diagram illustrating an example of the structure of an encoding apparatus for encoding multiple networks in Embodiment 3.

[0069] Figure 58 This is a diagram showing an example of the encoded data of the first network after learning in Implementation Method 3.

[0070] Figure 59 This is a diagram showing an example of the encoded data of the second network after learning in Implementation Method 3.

[0071] Figure 60 This is a diagram illustrating an example of the detailed structure of the first network learning unit in Embodiment 3.

[0072] Figure 61 This is a flowchart illustrating an example of the learning process of the network in Implementation Method 3.

[0073] Figure 62 This is a block diagram illustrating an example of the structure of a decoding device that decodes multiple networks in Embodiment 3.

[0074] Figure 63 This is a diagram illustrating an example of the syntax of metadata for sequence units in Implementation 3.

[0075] Figure 64 This is a diagram of an example of the syntax for representing metadata of frame units.

[0076] Figure 65 This is a diagram illustrating an example of the syntax of data units in a high-density network in Implementation 3.

[0077] Figure 66 This is a diagram illustrating an example of the syntax of data units in a low-density network in Implementation 3.

[0078] Figure 67This is a diagram illustrating an example of the structure of the data unit of the first network in Implementation 3.

[0079] Figure 68 This is a diagram illustrating an example of the structure of the data unit of the second network in Implementation 3.

[0080] Figure 69 This is a diagram illustrating an example of the syntax of the encoded data of the three-dimensional model of NeRF in Implementation 3.

[0081] Figure 70 This is a diagram illustrating an example of the cell type of NeRF in Implementation 3.

[0082] Figure 71 This is a diagram illustrating an example of the data structure of the encoded data of the three-dimensional model of NeRF in Implementation 3.

[0083] Figure 72 This is a diagram illustrating an example of the SPS syntax of the three-dimensional model of NeRF in Implementation 3.

[0084] Figure 73 This is a diagram illustrating an example of the syntax of the structural information of the three-dimensional model of NeRF in Implementation 3.

[0085] Figure 74 This is a diagram representing an example of component_type in implementation method 3.

[0086] Figure 75 This is a diagram illustrating an example of the component coding type in implementation method 3.

[0087] Figure 76 This is a diagram illustrating the reference relationships of the encoded data of the three-dimensional model of NeRF in Implementation 3.

[0088] Figure 77 This is a diagram illustrating an example of how the data of a frame in Implementation 3 is divided into three three-dimensional spaces.

[0089] Figure 78 This is a diagram illustrating an example of an ID assigned to the segmented data in Implementation Method 3.

[0090] Figure 79 This is a diagram used to illustrate the first example of the encoding method in Implementation Method 3.

[0091] Figure 80 This is a diagram illustrating the first example of the output of the decoding device in Embodiment 3.

[0092] Figure 81 This is a diagram illustrating a second example of the encoding method in Implementation Method 3.

[0093] Figure 82 This is a diagram illustrating a second example of the output of the decoding device in Embodiment 3.

[0094] Figure 83 This is a diagram illustrating a third example of the encoding method in Implementation Method 3.

[0095] Figure 84 This is a diagram illustrating a third example of the output of the decoding device in Embodiment 3.

[0096] Figure 85 This is a diagram illustrating the fourth example of the encoding method in Implementation Method 3.

[0097] Figure 86 This is a diagram illustrating the fourth example of the output of the decoding device in Embodiment 3.

[0098] Figure 87 This is a diagram illustrating the fifth example of the encoding method in Implementation Method 3.

[0099] Figure 88 This is a diagram illustrating the fifth example of the output of the decoding device in Embodiment 3.

[0100] Figure 89 This diagram illustrates the data exchange between the decoding unit and the control unit in Embodiment 3.

[0101] Figure 90 This is a diagram used to illustrate the consistency points in Implementation Method 3.

[0102] Figure 91 This is a diagram illustrating an example of a bitstream containing multiple networks in Implementation 3.

[0103] Figure 92 This is a diagram illustrating an example of the syntax of the layered structure of multiple networks in Implementation 3.

[0104] Figure 93 This is a block diagram illustrating an example of the structure of a modified version of the encoding device in Embodiment 3.

[0105] Figure 94 This is a block diagram illustrating an example of the structure of a modified version of the decoding device in Embodiment 3.

[0106] Figure 95 This is a diagram illustrating an example of the structure of the encoding device in Embodiment 3.

[0107] Figure 96 This is a flowchart illustrating the first example of the encoding method of the encoding device in Embodiment 3.

[0108] Figure 97This is a diagram illustrating an example of the structure of the decoding device in Embodiment 3.

[0109] Figure 98 This is a flowchart illustrating a first example of the decoding method of the decoding device in Embodiment 3.

[0110] Figure 99 This is a flowchart illustrating a second example of the encoding method of the encoding device in Embodiment 3.

[0111] Figure 100 This is a flowchart illustrating a second example of the decoding method of the decoding device in Embodiment 3.

[0112] Figure 101 This is a flowchart illustrating a third example of the encoding method of the encoding device in Embodiment 3.

[0113] Figure 102 This is a flowchart illustrating a third example of the decoding method of the decoding device in Embodiment 3.

[0114] Figure 103 This is a flowchart illustrating a fourth example of the encoding method of the encoding device in Embodiment 3.

[0115] Figure 104 This is a flowchart illustrating a fourth example of the decoding method of the decoding device in Embodiment 3. Detailed Implementation

[0116] The encoding apparatus of the first aspect of this disclosure includes: a circuit; and a memory connected to the circuit; the circuit acquires multiple networks constituting a three-dimensional data generation model during operation, encodes the multiple networks, and outputs a bit stream containing the encoded multiple networks, the bit stream containing: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames.

[0117] Therefore, the encoding device can output a bitstream containing encoded data obtained from multiple networks that will constitute a 3D generative model. This allows for a reduction in the storage capacity or the amount of 3D data transmitted.

[0118] The second type of encoding apparatus disclosed herein, in the first type of encoding apparatus, includes at least one of the first metadata and the second metadata containing a first parameter representing the number of layers of the first network in the plurality of networks and a second parameter representing the number of nodes in each layer.

[0119] In the third-party encoding apparatus disclosed herein, in a second-mode encoding apparatus, the second metadata contains the same specific parameter as the specific parameter contained in the first metadata, and the specific parameter contained in the second metadata is used preferentially over the specific parameter contained in the first metadata.

[0120] In the third-party encoding apparatus disclosed herein, in a second-mode encoding apparatus, the second metadata contains the same specific parameter as the specific parameter contained in the first metadata, and the specific parameter contained in the second metadata is used preferentially over the specific parameter contained in the first metadata.

[0121] In the fourth type of encoding apparatus disclosed herein, in the second type or third type of encoding apparatus, the first metadata includes a flag indicating which of the first or second parameters is included in the first metadata and the second metadata.

[0122] The fifth encoding apparatus of this disclosure, in any of the first to fourth encoding apparatuses, includes an identifier representing the three-dimensional data generation model in the header of the unit contained in the bitstream. Therefore, the decoding apparatus that has obtained the bitstream can identify whether it is a network for the same three-dimensional model based on the identifier contained in the bitstream.

[0123] The sixth encoding apparatus of this disclosure, in the fifth encoding apparatus, divides a frame corresponding to the three-dimensional data generation model into multiple spaces, at least one of the multiple spaces into multiple subspaces, and the identifier contains different values ​​corresponding to the multiple subspaces respectively.

[0124] The seventh encoding apparatus of this disclosure, in any of the first to fourth encoding apparatuses, includes an identifier in the header of each unit in the bitstream that represents the frame number of the frame corresponding to the three-dimensional data generation model. Therefore, the decoding apparatus that has obtained the bitstream can identify whether data units belong to the same frame based on the identifiers contained in the bitstream.

[0125] The eighth encoding apparatus of this disclosure, in any of the first to fourth encoding apparatuses, includes a header in the bitstream containing an identifier indicating which region among the multiple spaces into which the three-dimensional data generation model is segmented. Therefore, the decoding apparatus that has obtained the bitstream can identify whether the data belongs to the same region based on the identifiers contained in the bitstream.

[0126] In the ninth aspect of the encoding apparatus of this disclosure, in any one of the first to eighth aspects of the encoding apparatus, the first metadata includes a parameter representing the number of components constituting the bitstream and an identifier identifying the components.

[0127] The tenth aspect of the decoding apparatus disclosed herein includes: a circuit; and a memory connected to the circuit; the circuit acquires a bitstream during operation, the bitstream comprising a plurality of encoded networks constituting a three-dimensional data generation model, the bitstream comprising: (1) first metadata storing parameters common to the sequence; and (2) second metadata storing parameters common to frames, access units, or multiple frames, the circuit decoding the plurality of networks based on the bitstream during operation. Therefore, the decoding apparatus is capable of appropriately decoding bitstreams that achieve reduced transmission throughput.

[0128] The decoding apparatus of the eleventh aspect of this disclosure, in the decoding apparatus of the tenth aspect, at least one of the first metadata and the second metadata includes a first parameter representing the number of layers of the first network in the plurality of networks and a second parameter representing the number of nodes in each layer.

[0129] The decoding apparatus of the twelfth aspect of this disclosure, in the decoding apparatus of the eleventh aspect, includes a second metadata containing the same specific parameter as a specific parameter contained in the first metadata, wherein the specific parameter contained in the second metadata is used preferentially over the specific parameter contained in the first metadata.

[0130] The decoding apparatus of the thirteenth aspect of this disclosure, in the decoding apparatus of the eleventh or twelfth aspect, includes a flag indicating which of the first metadata or the second parameter is included in the first metadata and the second metadata.

[0131] The decoding apparatus of the fourteenth aspect of this disclosure, in any of the decoding apparatuses of the tenth to thirteenth aspects, includes a header in the bitstream containing an identifier representing the three-dimensional data generation model. Therefore, the decoding apparatus that has obtained the bitstream can identify whether it is a network for the same three-dimensional data generation model based on the identifier contained in the bitstream.

[0132] The decoding apparatus of the fifteenth aspect of this disclosure, in the decoding apparatus of the fourteenth aspect, a frame corresponding to the three-dimensional data generation model is divided into multiple spaces, at least one of the multiple spaces is divided into multiple subspaces, and the identifier contains different values ​​corresponding to the multiple subspaces respectively.

[0133] The decoding apparatus of the sixteenth aspect of this disclosure, in any of the decoding apparatuses of the tenth to thirteenth aspects, includes an identifier in the header of each unit in the bitstream that represents the frame number of the frame corresponding to the three-dimensional data generation model. Therefore, the decoding apparatus that has obtained the bitstream can identify whether data units belong to the same frame based on the identifier contained in the bitstream.

[0134] The seventeenth aspect of the decoding apparatus of this disclosure, in any of the tenth to thirteenth aspects of the decoding apparatus, includes a header in the bitstream containing an identifier indicating which region among the multiple spaces into which the three-dimensional data generation model is segmented. Therefore, the decoding apparatus that has obtained the bitstream can identify whether the data belongs to the same region based on the identifier contained in the bitstream.

[0135] The decoding apparatus of the eighteenth aspect of this disclosure, in any of the decoding apparatuses of the tenth to the seventeenth aspects, wherein the first metadata includes a parameter representing the number of components constituting the bitstream and an identifier identifying the components.

[0136] The nineteenth method of this disclosure is an encoding method executed by an encoding device, which obtains a three-dimensional data generation model, encodes multiple networks constituting the three-dimensional data generation model, and outputs a bit stream containing the encoded multiple networks, wherein the bit stream contains: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames.

[0137] Therefore, it is possible to output a bitstream containing encoded data obtained by encoding multiple networks that constitute a 3D generative model. This allows for a reduction in the storage capacity or the transmission volume of 3D data.

[0138] The twentieth aspect of this disclosure is a decoding method executed by a decoding device, which obtains a bitstream containing multiple encoded networks, the bitstream comprising: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames, and decodes the multiple networks based on the bitstream. Therefore, it is possible to appropriately decode a bitstream that achieves reduced transmission volume.

[0139] Furthermore, these general or specific technical solutions can be implemented through systems, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or through any combination of systems, methods, integrated circuits, computer programs, and recording media.

[0140] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Furthermore, the embodiments described below represent specific examples of this disclosure. The numerical values, shapes, materials, constituent elements, arrangement positions of constituent elements, connection methods, steps, and order of steps shown in the following embodiments are examples and are not intended to limit this disclosure. Additionally, constituent elements in the following embodiments that are not described in the independent claims representing the highest-level concept will be described as arbitrary constituent elements.

[0141] (Implementation Method 1) The configuration of the three-dimensional data encoding and decoding system of this embodiment will be described. Figure 1 This is a diagram illustrating an example of the configuration of the three-dimensional data encoding and decoding system of this embodiment. (See diagram for example.) Figure 1 As shown, the three-dimensional data encoding and decoding system includes a three-dimensional data encoding system 1001, a three-dimensional data decoding system 1002, a sensor terminal 1003, and an external connection unit 1004.

[0142] The 3D data encoding system 1001 generates encoded data or reused data by encoding 3D data. Furthermore, the 3D data encoding system 1001 can be a single device or a system implemented by multiple devices. Additionally, the 3D data encoding device may include a portion of the multiple processing units included in the 3D data encoding system 1001.

[0143] The 3D data encoding system 1001 includes a 3D data generation system 1011, a prompting unit 1012, an encoding unit 1013, a multiplexing unit 1014, an input / output unit 1015, and a control unit 1016. The 3D data generation system 1011 includes a sensor information acquisition unit 1017 and a 3D data generation unit 1018.

[0144] The sensor information acquisition unit 1017 acquires sensor signals from the sensor terminal 1003 and outputs the sensor signals to the three-dimensional data generation unit 1018. The three-dimensional data generation unit 1018 generates three-dimensional data based on the sensor signals and outputs the three-dimensional data to the encoding unit 1013.

[0145] The prompting unit 1012 prompts the user with sensor signals or three-dimensional data. For example, the prompting unit 1012 displays information or images based on sensor signals or three-dimensional data.

[0146] The encoding unit 1013 encodes (compresses) the three-dimensional data and outputs the resulting encoded data, control information obtained during the encoding process, and other additional information to the multiplexing unit 1014. The additional information may include, for example, sensor signals.

[0147] The multiplexing unit 1014 generates multiplexed data by multiplexing the encoded data, control information, and additional information input from the encoding unit 1013. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.

[0148] The input / output unit 1015 (e.g., a communication unit or interface) outputs multiplexed data to the outside. Alternatively, the multiplexed data is stored in an internal memory or other storage unit. The control unit 1016 (or application execution unit) controls each processing unit. That is, the control unit 1016 performs control such as encoding and multiplexing. The control unit 1016 can also perform control such as demultiplexing, decoding, or prompting.

[0149] Alternatively, sensor signals can be input to the encoding unit 1013 or the multiplexing unit 1014. Furthermore, the input / output unit 1015 can directly output 3D data or encoded data to the outside.

[0150] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 1001 is input to the three-dimensional data decoding system 1002 via the external connection unit 1004.

[0151] The 3D data decoding system 1002 generates 3D data by decoding encoded or multiplexed data. Furthermore, the 3D data decoding system 1002 can be a single device or a system comprised of multiple devices. Additionally, the 3D data decoding device may include a portion of the multiple processing units included in the 3D data decoding system 1002.

[0152] The 3D data decoding system 1002 includes a sensor information acquisition unit 1021, an input / output unit 1022, a demultiplexing unit 1023, a decoding unit 1024, a prompting unit 1025, a user interface 1026, and a control unit 1027.

[0153] The sensor information acquisition unit 1021 acquires sensor signals from the sensor terminal 1003.

[0154] The input / output unit 1022 acquires the transmission signal, decodes the multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the demultiplexing unit 1023.

[0155] The demultiplexing unit 1023 obtains encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 1024.

[0156] The decoding unit 1024 reconstructs the point group data by decoding the encoded data.

[0157] The prompting unit 1025 prompts the user with dot group data. For example, the prompting unit 1025 displays information or images based on the dot group data. The user interface 1026 obtains instructions based on the user's operation. The control unit 1027 (or application execution unit) controls each processing unit. That is, the control unit 1027 performs demultiplexing, decoding, and prompting control, etc.

[0158] Furthermore, the input / output unit 1022 can directly acquire point group data or encoded data from external sources. Additionally, the prompting unit 1025 can acquire additional information such as sensor signals and provide prompts based on this additional information. Furthermore, the prompting unit 1025 can also provide prompts based on user instructions obtained from the user interface 1026.

[0159] The sensor terminal 1003 generates the information obtained from the sensor, namely the sensor signal. The sensor terminal 1003 is a terminal equipped with a sensor or camera, such as a moving object like a car, a flying object like an airplane, a portable terminal, or a camera.

[0160] Sensor signals that can be acquired by sensor terminal 1003 include, for example: (1) signals obtained from LIDAR, millimeter-wave radar, or infrared sensors indicating the distance between sensor terminal 1003 and an object or the reflectivity of the object; (2) signals obtained from multiple monocular camera images or stereo camera images indicating the distance between a camera and an object or the reflectivity of the object. Additionally, sensor signals may also include sensor posture, orientation, gyroscope (angular velocity), position (GPS information or altitude), velocity, or acceleration. Furthermore, sensor signals may also include temperature, air pressure, humidity, or magnetism.

[0161] The external connection unit 1004 is realized through communication with integrated circuits (LSI or IC), external storage units, cloud servers via the Internet, or broadcasting.

[0162] Next, the point group data will be explained. Figure 2 It is a diagram showing the composition of point group data. Figure 3 This is a diagram illustrating an example of the structure of a data file that records information about point group data.

[0163] Point cluster data comprises data from multiple points. Each point's data includes location information (3D coordinates) and attribute information related to that location. A cluster of these points is called a point cluster. For example, a point cluster indicates the 3D shape of an object.

[0164] Position information, such as three-dimensional coordinates, is sometimes referred to as geometric information. Additionally, the data for each point can also contain attribute information across multiple attribute categories. Attribute categories could include, for example, color or reflectivity.

[0165] One attribute can be mapped to one location, or multiple attributes with different attribute categories can be mapped to one location. Furthermore, multiple mappings can be established between attribute categories and one location.

[0166] Figure 3 The example data file shown illustrates a one-to-one correspondence between location information and attribute information, illustrating the location and attribute information of N points constituting the point group data.

[0167] Positional information includes, for example, information along the x, y, and z axes. Attribute information includes, for example, RGB color information. Representative data files include .ply files, etc.

[0168] Next, the three-dimensional mesh data will be explained. Figure 4 It is a diagram showing the composition of three-dimensional mesh data. Figure 5 This is a diagram illustrating an example of the structure of a data file containing information about three-dimensional mesh data.

[0169] 3D mesh data is a data format used in Computer Graphics (CG) that indicates the three-dimensional shape of an object through a collection of face information. These face information points to polygons such as triangles or quadrilaterals. 3D mesh data is also known as polygonal data or polygonal meshes.

[0170] The constituent elements are a group of three-dimensional points, vertices of the multiple three-dimensional points forming the group, edges connecting two vertices of the multiple three-dimensional points, and a set of faces enclosed by the multiple edges. A group of three-dimensional points is a collection of points that contains positional information in three-dimensional space and attribute information corresponding to that positional information. Furthermore, a three-dimensional point can also be simply referred to as a point.

[0171] Vertices can also possess attributes such as color, reflectivity, and normal vectors specific to a 3D point. The relationships between vertices forming an edge or face can also be represented by connectivity. Furthermore, vertices can also be represented by position. The face and back faces can be represented by the orientation of the normal vectors relative to 3D points. Additionally, vertices can also possess surface-specific attribute information.

[0172] Grid data files can take the form of object files, for example. Figure 5 In the mesh data file shown, the position information G(1)~G(N) of the N vertices constituting the mesh, and the attribute information A(1)~A(N) of the vertices are represented as vertex information. In the mesh data file, vertex information may also not include attribute information.

[0173] Furthermore, attribute information does not necessarily have to correspond one-to-one with vertices. Figure 5 The mesh data file shows an example of three-dimensional mesh data with M attribute information A2.

[0174] Face information is represented by a combination of vertex indices. n[1,3,4] represents the face of a triangle formed by vertices n=1, n=3, and n=4.

[0175] Additionally, m[2,4,6] indicates that the attribute information in attribute information A2 with m=2, m=4, and m=6 correspond to three vertices respectively. Furthermore, an example of a face consisting of three vertices is shown here, but a face can have any number of vertices greater than or equal to 3, not limited to 3. For example, if the face is a quadrilateral, the number of vertices is 4; if the face is a polygon, the number of vertices is the same as the number of vertices constituting the polygon.

[0176] Furthermore, attribute information A2 can be represented by a file different from the mesh data file, or it can include its pointer information. For example, attribute information can be stored in a two-dimensional attribute map file, and the attribute map filename and the two-dimensional coordinates in the attribute map can also be represented by attribute information A2 from the mesh data file. In this way, attribute information A2 can be contained in the mesh data file or represented by a file different from the mesh data file; regardless of the method used, attribute information for three-dimensional points can be specified.

[0177] Next, the three-dimensional model will be explained. Figure 6 It is a diagram used to illustrate a three-dimensional model.

[0178] A 3D model is a model generated based on 2D or 3D data.

[0179] The 3D model learning unit 1031 generates a network model, i.e., a 3D model, by learning 2D data (2D images) or 3D data (point groups or meshes) and using neural networks to learn 3D shapes and corresponding attribute information.

[0180] The 3D model learning unit 1031 can also generate a 3D model by learning from a 2D image using NeRF (Neural Radiance Fields). Alternatively, the 3D model learning unit 1031 can generate a 3D model after transforming a 2D image into 3D data through photogrammetry using the 2D image. The 3D model can also be generated using 3D data obtained from a sensor (distance sensor).

[0181] Three-dimensional model data consists of the elements that constitute a three-dimensional model, containing information indicating the structure of the network model, feature quantities, etc. For example, three-dimensional model data contains information related to the constituent elements of a neural network. This information includes, for example, multiple layers such as input layer, intermediate layers, and output layer, nodes in each layer, weight coefficients for nodes, and transformation functions for nodes.

[0182] The 3D model encoding unit 1032 can also encode 3D model data and transmit the encoded 3D model data.

[0183] The 3D model decoding unit 1033 receives the transmitted encoded 3D model data and decodes the 3D model based on the encoded 3D model data.

[0184] The rendering reconstruction unit 1034 reconstructs (generates) two-dimensional data (two-dimensional images) or three-dimensional data (point groups or meshes) based on the decoded three-dimensional model. For example, when using a three-dimensional model modeled via NeRF, the rendering reconstruction unit 1034 obtains viewpoint position or line-of-sight vector information, generates rendered two-dimensional data (two-dimensional images) based on the three-dimensional model and the viewpoint position or line-of-sight vector, and outputs the two-dimensional data. The generated two-dimensional data indicates a two-dimensional image of a three-dimensional object observed from the viewpoint position or from the line of sight indicated by the line-of-sight vector. The three-dimensional object is a three-dimensional object that serves as the source of the two-dimensional or three-dimensional data input to the three-dimensional model learning unit 1031.

[0185] Next, the types of three-dimensional data will be explained. Figure 7 This is a diagram illustrating the types of three-dimensional data. For example... Figure 7 As shown, there are static and dynamic objects in 3D data.

[0186] Static objects are 3D data at any given time (a specific moment). Dynamic objects are 3D data that changes over time. Hereinafter, the point group data at a specific moment will be referred to as a PCC frame or frame. Additionally, grid data at any given time will be referred to as a grid frame or frame.

[0187] The object can be three-dimensional data with the area restricted to some extent, such as typical image data, or it can be three-dimensional data with the area unrestricted, such as map information.

[0188] In addition, points of various densities can exist, including sparse point clusters (sparse grid data) and dense point clusters (dense grid data).

[0189] The details of each processing unit are described below. Sensor information is obtained through various methods such as distance sensors like LIDAR or rangefinders, stereo cameras, or combinations of multiple monocular cameras. The 3D data generation unit 1018 generates point group data based on the sensor information obtained by the sensor information acquisition unit 1017. The 3D data generation unit 1018 generates position information (geometric information) as point group data and adds attribute information specific to that position information to the position information.

[0190] The 3D data generation unit 1018 can also process point group data during the generation of position information or the addition of attribute information. For example, the 3D data generation unit 1018 can reduce the amount of data by deleting point groups with overlapping positions. Furthermore, the 3D data generation unit 1018 can transform position information (position shifting, rotation, or normalization, etc.) and process point group data to generate mesh data. Additionally, the 3D data generation unit 1018 can also render attribute information.

[0191] In addition, Figure 1 In this system, the three-dimensional data generation system 1011 is included in the three-dimensional data encoding system 1001, but it can also be set independently outside the three-dimensional data encoding system 1001.

[0192] The encoding unit 1013 generates encoded data by encoding the three-dimensional data based on a predefined encoding method. Regarding encoding methods, there are G-PCC (encoding method using position information), V-PCC (encoding method using video codec), Draco (grid coding method), and V-DMC (grid coding method). The encoding method is not limited to these methods; for example, it can also be a method for encoding dynamic grids, or other methods combining these methods.

[0193] The decoding unit 1024 decodes the three-dimensional data by decoding the encoded data based on a predefined encoding method.

[0194] The multiplexing unit 1014 generates multiplexed data by multiplexing encoded data using existing multiplexing methods. The generated multiplexed data is transmitted or stored. In addition to multiplexing encoded data of 3D data, the multiplexing unit 1014 also multiplexes other media such as images, sounds, subtitles, applications, and files, or reference time information. Furthermore, the multiplexing unit 1014 can also multiplex attribute information associated with sensor information or point group data.

[0195] As multiplexing methods or file formats, there are ISOBMFF, ISOBMFF-based transmission methods such as MPEG-DASH, MMT, MPEG-2 TS Systems, and RTP.

[0196] The demultiplexing unit 1023 extracts encoded data, other media, and time information from the multiplexed data.

[0197] The input / output unit 1015 transmits multiplexed data using a method that matches the transmission medium or storage medium, such as broadcasting or communication. The input / output unit 1015 can communicate with other devices via the Internet, or with storage units such as cloud servers.

[0198] Use HTTP, FTP, TCP, or UDP as the communication protocol. You can also use either the PULL or PUSH communication method.

[0199] Either wired or wireless transmission can be used. For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or coaxial cable can be used. For wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), or millimeter wave can be used.

[0200] In addition, broadcast methods can be used, such as DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3.

[0201] Next, the process of dividing 3D data into more than one 3D data segment will be explained. Figure 8 It is a diagram used to illustrate the encoding and processing of three-dimensional data. Figure 9 This is a diagram used to illustrate the decoding process of three-dimensional data.

[0202] like Figure 8 As shown, the data segmentation unit 1041 segments the three-dimensional data into one or more three-dimensional spaces, generating one or more segmented three-dimensional data (i.e., one or more segmented three-dimensional data). The encoding unit 1042 can also encode one or more segmented three-dimensional data to generate encoded data. The data segmentation unit 1041 and the encoding unit 1042 are components of an encoding device and can be included in one encoding device or in different devices.

[0203] One or more 3D spaces can be individually labeled as tiles or intervals. A 3D space can be, for example, a bounding box. Furthermore, the 3D data contained within each of the segmented 3D spaces can also be represented as slices. A slice is segmented 3D data, including any of the following: a group of points with location information (Geometry) or attribute information (Attribute), a mesh, or a 3D model. Each slice in the multiple slices is encoded by the encoding unit 1042 according to each constituent element and is output as encoded data. The encoded data includes the encoded slices.

[0204] like Figure 9 As shown, in the decoding process, the decoding unit 1051 decodes one or more segmented 3D data (one or more slices) based on the encoded data. The data combining unit 1052 combines the one or more segmented 3D data to restore (generate) 3D data. The decoding unit 1051 and the data combining unit 1052 are components of a decoding device and can be included in that decoding device or in different devices. The one or more segmented 3D data decoded by the decoding unit 1051 may not be combined. The decoding unit 1051 may also decode a portion of the one or more segmented 3D data based on a portion of the encoded data and output the decoded portion of the segmented 3D data. In this case, the decoding device may not have the data combining unit 1052.

[0205] Figure 10 It is a two-dimensional schematic diagram that illustrates the tiles and slices of three-dimensional data.

[0206] When encoding multiple slices, the encoding device can encode using the dependencies between the slices or without using dependencies. When encoding without dependencies, the encoding device can encode each slice independently, reducing processing time by encoding multiple slices in parallel. Similarly, when the decoding device encodes multiple slices without dependencies, it can decode each slice independently, reducing processing time by decoding multiple slices in parallel. Furthermore, the decoding device can reduce processing load by partially decoding only a portion of the multiple slices.

[0207] When encoding using dependency relationships, the encoding device transmits signals to identifiers indicating dependency relationships and encodes sequentially starting from the dependent data. When encoding multiple slices using dependency relationships, the decoding device decodes sequentially starting from the dependent data based on the identifiers.

[0208] In 3D data segmentation, any number of segments and any segmentation method can be used. The shape of an object can also be determined, and multiple 3D points can be segmented for each object. Alternatively, segmentation can be based on the number of 3D points contained in a slice; that is, an upper limit can be determined for the number of 3D points a slice can contain. Furthermore, 3D data can be segmented using map or location information, based on whether it is contained in 3D space (tile information). Multiple tile shapes can also overlap.

[0209] By dividing 3D data into multiple segmented 3D data in this way, adaptive encoding corresponding to the content or object can be performed, and parallel processing can be performed during decoding.

[0210] Next, the method for selecting the 3D data to be prompted or transmitted from multiple 3D data sets will be explained.

[0211] The server stores multiple 3D data sets for the same space. For example, the server stores point cluster data and mesh data for the same space. The server is an example of an encoding device. The terminal, based on its purpose, switches the 3D data obtained from the server and displays the switched 3D data. The terminal can also be a terminal that parses 3D data. In this case, the terminal can also switch the 3D data to be displayed based on purposes such as parsing or displaying, or user operations. The terminal is an example of a decoding device.

[0212] In switching between 3D data, the focus can be on whether to use a group of prompt points or a grid as the 3D data. Alternatively, the focus can be on whether to transmit a group of prompt points or a grid as the 3D data. For example, the terminal can send the user's selection to the server, receive (download) 3D data based on that selection from the server, and then provide prompts for the received 3D data. The 3D data (point group or grid) may or may not be encoded on the server. If the 3D data is encoded, the terminal can receive the encoded 3D data from the server, decode the 3D data based on the received encoded data, and then provide prompts for the decoded 3D data.

[0213] Next, the configuration of server 1070 and terminal 1090 will be explained. Figure 11 This is a block diagram illustrating an example of the functional configuration of a server and a terminal.

[0214] Server 1070 includes a data generation unit 1071, a synchronization unit 1075, a point group coding unit 1076, a grid coding unit 1077, a model coding unit 1078, a multiplexing unit 1079, and a data extraction unit 1080.

[0215] The data generation unit 1071 generates three-dimensional data based on at least one of two-dimensional data and three-dimensional data. The generated three-dimensional data includes at least two of point group data, mesh data, and three-dimensional model data. The data generation unit 1071 has a point group generation unit 1072, a mesh generation unit 1073, and a model generation unit 1074. The data generation unit 1071 only needs to have at least two of the point group generation unit 1072, mesh generation unit 1073, and model generation unit 1074. The point group generation unit 1072 generates point group data based on at least one of two-dimensional data and three-dimensional data. The mesh generation unit 1073 generates mesh data based on at least one of two-dimensional data and three-dimensional data. The model generation unit 1074 generates three-dimensional model data by performing machine learning based on at least one of two-dimensional data and three-dimensional data.

[0216] The two-dimensional data input to the data generation unit 1071 can also be two-dimensional images acquired by a camera. The three-dimensional data input to the data generation unit 1071 can, for example, be point data of spaces such as a building site, factory, or office acquired by a sensor such as LiDAR. The data generation unit 1071 can also generate color information as attribute information for each point contained in the point data of the three-dimensional data using the two-dimensional image of the two-dimensional data. The three-dimensional data generated by the data generation unit 1071 can also be divided into arbitrary spaces. Point data, mesh data, and three-dimensional model data can also be divided into arbitrary spaces respectively.

[0217] The synchronization unit 1075 acquires the spatial position or time (reproduction time, decoding time, acquisition time, etc.) of the point group data, mesh data, and 3D model data generated by the data generation unit 1071. The time of each data point is the reproduction time, decoding time, acquisition time, etc. Alternatively, the synchronization unit 1075 may not acquire the synchronization of the point group data, mesh data, and 3D model data, but instead generate synchronization information for acquiring synchronization. Furthermore, the synchronization unit 1075 may perform processing to acquire the synchronization of at least two types of 3D data from the point group data, mesh data, and 3D model data generated by the data generation unit 1071, or generate synchronization information (synchronization signal) for acquiring synchronization; alternatively, it may not perform the processing for acquiring the synchronization of all three types of 3D data (synchronization processing).

[0218] The point group encoding unit 1076 encodes the point group data that has been synchronized by the synchronization unit 1075. Alternatively, the point group encoding unit 1076 may not encode the point group data. The point group data may be pre-encoded or encoded upon request from the terminal 1090.

[0219] The grid encoding unit 1077 encodes the grid data that has been synchronized by the synchronization unit 1075.

[0220] The model encoding unit 1078 encodes the three-dimensional model data that has been synchronized by the synchronization unit 1075.

[0221] The multiplexing unit 1079 uses a prescribed format or a prescribed multiplexing method to multiplex the encoded point group data (coded point group), the encoded mesh data (coded mesh data), the encoded 3D model data, and the synchronization information. Alternatively, multiplexing based on the multiplexing unit 1079 may not be performed. In this case, the server 1070 may not have the multiplexing unit 1079.

[0222] The data extraction unit 1080 extracts a portion of the multiplexed 3D data corresponding to the request from the terminal 1090 and sends the extracted portion of 3D data to the terminal 1090. Alternatively, data extraction based on the data extraction unit 1080 may not be performed. In this case, the server 1070 may not have the data extraction unit 1080. Without data extraction based on the data extraction unit 1080, the server 1070 can send the 3D data multiplexed by the multiplexing unit 1079 to the terminal 1090. Furthermore, even without multiplexing based on the multiplexing unit 1079, the server 1070 can send encoded point group data (encoded point group), encoded mesh data (encoded mesh), encoded 3D model data (encoded 3D model), and synchronization information to the terminal 1090, or send a bitstream containing encoded point group data (encoded point group), encoded mesh data (encoded mesh), encoded 3D model data (encoded 3D model), and synchronization information to the terminal 1090.

[0223] The terminal 1090 includes a control unit 1091, a decoding unit 1092, and a prompting unit 1093.

[0224] The control unit 1091 sends a request for a portion of the 3D data to the server 1070. The control unit 1091 can also process user operations to determine a portion of the 3D data.

[0225] The decoding unit 1092 decodes a portion of the 3D data based on the bit stream (encoded data) obtained from the server 1070.

[0226] The prompting unit 1093 renders a portion of the decoded 3D data to provide a prompt.

[0227] Figure 11 The data generation unit 1071 can also be... Figure 12 The data generation unit 1110 shown is used to achieve this. Figure 12 This is a block diagram illustrating another example of the data generation unit of a server.

[0228] The data generation unit 1110 includes a point group generation unit 1111, a grid generation unit 1112, and a model generation unit 1113.

[0229] The point group generation unit 1111 has the same function as the point group generation unit 1072. The point group generation unit 1111 acquires point group data from the point group sensor 1101 and a two-dimensional image from the camera 1102, and generates point group data based on the point group data and the two-dimensional image. The point group data generated by the point group generation unit 1111 includes position information of each point and attribute information (such as color information) extracted from the two-dimensional image, wherein the attribute information corresponds to each point indicated by the position information.

[0230] The mesh generation unit 1112 generates mesh data based on the point group data generated by the point group generation unit 1111.

[0231] The model generation unit 1113 has the same function as the model generation unit 1074. The model generation unit 1113 acquires point group data from the point group sensor 1101 and two-dimensional images from the camera 1102, and performs machine learning based on the point group data and the two-dimensional images to generate three-dimensional model data.

[0232] like Figure 11 As explained, point cluster data, grid data, and 3D model data can also be generated independently. For example... Figure 12 As explained, grid data can also be generated from point cluster data. Furthermore, point cluster data can also be generated from grid data.

[0233] Mesh can be generated from point groups, and point groups can also be generated from mesh.

[0234] Furthermore, point cluster data, mesh data, and 3D model data can be generated by the server 1070, or by sensors or a terminal 1090 equipped with sensors. Sensors include, for example, a point cluster sensor 1101 and a camera 1102.

[0235] Next, the relationship between three-dimensional space and coded data will be explained. Figure 13 It is a diagram used to illustrate the relationship between three-dimensional space and coded data.

[0236] As mentioned above, three-dimensional data includes, for example, point group data, grid data, and any of the three-dimensional models.

[0237] like Figure 13As shown, when 3D data is divided into three 3D data segments in three 3D spaces (tiles or intervals), the encoding device encodes each of the three segments separately and adds a header to perform data unitization. The header contains an identifier (Space_ID) of the space to which the encoded data of that data unit belongs, and an identifier (DataUnit_ID) of the data unit.

[0238] The data unit is further given a header containing the identifier of the data unit or the length information of the data unit, and the encoding method unit is generated through unitization.

[0239] Next, the syntax of the encoding method unit will be explained. Figure 14 This is a diagram illustrating an example of the syntax of an encoding scheme unit. Figure 15 This is a diagram illustrating an example of the syntax for encoding point groups. Figure 16 This is a diagram illustrating an example of the syntax of a coding grid. Figure 17 This is a diagram illustrating an example of the syntax for encoding a three-dimensional model.

[0240] The `unit_type` directive indicates the type of data unit stored in the encoding mode unit. Thus, the type of data unit stored in the encoding mode unit is specified.

[0241] length indicates the length of a data unit.

[0242] The data() function indicates the body of a data unit.

[0243] exist Figure 15 In this context, when `unit_type` is 0, it indicates that the data unit is the location information (geometric information) of a group of encoded points. When `unit_type` is 1, it indicates that the data unit is the attribute information of a group of encoded points. When `unit_type` is 2, it indicates that the data unit is the metadata of a group of encoded points.

[0244] exist Figure 16 In this context, when `unit_type` is 0, it indicates that the data unit is encoding the location information (geometric information) of the mesh. When `unit_type` is 1, it indicates that the data unit is encoding the attribute information of the mesh. When `unit_type` is 2, it indicates that the data unit is encoding the metadata of the mesh.

[0245] exist Figure 17 In the context of `unit_type` being 0, the data unit indicates that it encodes element 1 of the 3D model. With `unit_type` being 1, the data unit indicates that it encodes element 2 of the 3D model. With `unit_type` being 2, the data unit indicates that it encodes metadata of the 3D model.

[0246] in addition, Figures 15-17 The syntax shown is an example and is not limited to the above configuration. These syntaxes can be constructed using a portion of the syntax, or using types (categories) not mentioned above, and the order of the syntactic components can be rearranged. For example, in the syntax of the encoding mode unit, it can also be as follows: Figure 14 The configuration of the encoding mode unit, which is common to multiple encoding modes, is shown in that way. Figures 15-17 The unit_type, length, and data() shown are shown.

[0247] Additionally, a header can be assigned to the encoding unit to indicate its category. The categories of encoding units include, for example, `point_cloud_codec_unit` for point cloud data, `mesh_codec_unit` for mesh data, and `model_codec_unit` for 3D model data. This allows for the comprehensive processing of multiple encoding methods.

[0248] Figure 18 This is a diagram illustrating an example of the syntax for three-dimensional data information.

[0249] In terms of syntax, when multiple encoding methods are saved in a single format, the number of 3D data contained in that format (number_of_3Dformat) and the type of 3D data (format_type) can also be indicated, and data in each format can be saved. Therefore, it is possible to comprehensively process multiple encoding methods or 3D data, and to recognize multiple encoding methods or 3D data.

[0250] 3Ddata_info indicates the format structure information for storing multiple 3D data sets.

[0251] number_of_3Dformat indicates the number of 3D formats used.

[0252] `format_type` indicates the category of the format of the saved 3D data. For example, a number for `format_type` and the corresponding format can be determined as follows: `format_type` of 0 indicates that the saved 3D data is in the format of point cloud. `format_type` of 1 indicates that the saved 3D data is in the format of mesh. `format_type` of 2 indicates that the saved 3D data is in the format of G-PCC (g-pcc). `format_type` of 3 indicates that the saved 3D data is in the format of V-DMC (v-dmc). `format_type` of 4 indicates that the saved 3D data is in the format of 3D model.

[0253] Next, the data structure of the encoded data of multiple three-dimensional data will be explained according to each type of three-dimensional data. Figure 19 It is a diagram used to illustrate the data structure of a group of coded points. Figure 20 It is a diagram used to illustrate the data structure of the coded grid. Figure 21 It is a diagram used to illustrate the data structure of a coded 3D model.

[0254] For each type of three-dimensional data, the encoding device divides the three-dimensional data into multiple three-dimensional data according to each of the multiple spatial regions, and encodes the multiple three-dimensional data (i.e. multiple segmented three-dimensional data) separately to generate encoded data.

[0255] Each encoded data is assigned a header, and at least one of the data_unit_id and space_id is stored.

[0256] Here, `data_unit_id` is an identifier that identifies a data unit within the encoded data and is unique within the encoded data. Additionally, `space_id` indicates identification information for a spatial region. If either `data_unit_id` or `space_id` is common across multiple 3D datasets, it indicates the same value across all three-dimensional datasets.

[0257] exist Figures 19-21 In the example, data units with data_unit_id=0 in the encoded point group, data units with data_unit_id=3 in the encoded mesh, and data units with data_unit_id=0 in the encoded 3D model are all assigned space_id=1. This means that the 3D data is contained in the common 3D space indicated by Space_ID#1.

[0258] Data, including headers, can be contained in bitstream structures such as data units or encoding methods, or stored in file formats specified by ISOBMFF, such as various BOXes.

[0259] Next, the three-dimensional spatial information will be explained. Figure 22 It is a diagram that represents an example of multiple three-dimensional spaces in a two-dimensional way. Figure 23 This is a diagram showing an example of a bounding box. Figure 24 This is a diagram illustrating an example of syntax for three-dimensional spatial information.

[0260] In the syntax of three-dimensional spatial information, 3Dspace_info indicates information about the segmented three-dimensional space. 3Dspace_info can be used for partial decoding.

[0261] number_of_space indicates the number of three-dimensional spaces after partitioning.

[0262] space_id indicates the identifier of the segmented three-dimensional space.

[0263] Three-dimensional spatial information includes bounding box information as a specification Figure 23 Information about the bounding box shown.

[0264] The bounding box information includes bounding_box_xyz and bounding_box_whd.

[0265] `bounding_box_xyz` indicates the coordinates of the reference point of the bounding box. Figure 23 In the example, the coordinates of x, y, and z (x0, y0, z0) are used to represent it.

[0266] `bounding_box_whd` indicates the size of the bounding box. Figure 23 In the example, it can be represented by width w, height h, and depth d (w0, h0, d0).

[0267] Additionally, three-dimensional spatial information may include an identifier for each data unit of the encoded data. Furthermore, three-dimensional spatial information may also omit this identifier; that is, the identifier may not be transmitted as a signal.

[0268] pointcloud_id indicates the identifier of the data unit of the coded point group in the space corresponding to space_id.

[0269] mesh_id indicates the identifier of the data cell of the coded grid of the space corresponding to space_id.

[0270] model_id indicates the identifier of the data unit of the coded 3D model corresponding to space_id.

[0271] Furthermore, if the data unit shows a data_unit_id instead of a space_id, the identifier of each encoded data unit can be stored in the information indicating each space in the three-dimensional spatial information. This allows for the establishment of a correspondence between the three-dimensional spatial information and the segmented three-dimensional encoded data.

[0272] Alternatively, if the space_id is shown in the data unit, the three-dimensional spatial information can be mapped to an identifier for each data unit of the encoded data using the space_id. In this case, it is also possible not to store the identifier for each data unit of the encoded data.

[0273] Alternatively, the 3D spatial information of the point group data and the grid data can be commonalized by making the segmentation method, the origin of each segmented space, and the size of the bounding box the same in both the grid data and the point group data. Alternatively, the same 3D spatial information can be used in both the point group data and the grid data. This allows for the commonalization of 3D spatial information across different types of 3D data, enabling the use of the same 3D spatial information. By making 3D spatial information commonalized, switching between different types of 3D data (e.g., switching prompts or transmissions) becomes easier. Furthermore, in formats that integrate multiple 3D data sets, it is possible to utilize a single 3D spatial information across all 3D data sets instead of setting 3D spatial information for each set, thus reducing the amount of 3D spatial information.

[0274] In addition to point group data and grid data, it can also synchronize the three-dimensional spatial information of the three-dimensional model with other types of three-dimensional data, and can also make the three-dimensional spatial information of other types of three-dimensional data common.

[0275] Next, the relationship between the data structure of 3D data and partial decoding will be explained. Figure 25 This is a flowchart illustrating an example of partial decoding. Figure 26 This is a diagram illustrating an example of a three-dimensional spatial region of an object that is partially decoded. Figure 27 This is a diagram illustrating an example of the data structure of a partially decoded group of coded points. Figure 28 This is a diagram illustrating an example of a data structure for a partially decoded encoded grid. Figure 29 This is a diagram illustrating an example of the data structure of a partially decoded encoded 3D model.

[0276] In partial decoding, firstly, the decoding device determines the three-dimensional spatial region of the object to be partially decoded (S1001).

[0277] Next, the decoding device uses three-dimensional spatial information (3Dspace_info) to determine the region that overlaps with the three-dimensional spatial region of the object based on the bounding box information of multiple three-dimensional spatial regions, and obtains the space_id corresponding to the determined region (S1002).

[0278] Next, the decoding device obtains a data unit with the obtained space_id from the encoded data and decodes it (S1003). Thus, the decoding device performs partial decoding, decoding only a portion of the three-dimensional data. In partial decoding, the decoding device does not decode the entire three-dimensional data, but only a portion of it.

[0279] For example, such as Figure 26 As shown, when the three-dimensional space region of the object being partially decoded is shown in thick lines, the space_id of the obtained three-dimensional space is determined to be #2 based on the three-dimensional space information.

[0280] Then, as Figures 27-29 As shown, the encoded data of various three-dimensional data were used to establish corresponding data units with Space_id=#2 and then decoded.

[0281] In addition, the decoding device can also obtain the data unit ID from the three-dimensional spatial information instead of obtaining the space_id, obtain the data unit with the obtained data unit ID, and perform partial decoding.

[0282] In the above embodiments, point group data, mesh data, and 3D model data are exemplified as 3D data representing 3D objects, but the methods are not limited to these. For example, a 3D object may also be represented by multiple groups, each containing line-of-sight information indicating a line of sight and a 2D image obtained when viewing the 3D object from that line of sight. That is, data containing these multiple groups can also be processed as a type of 3D data. In addition, 3D data may also be data in other formats such as Gaussian splatting data.

[0283] Figure 30 This is a diagram illustrating an example of the configuration of a decoding device. Figure 31 This is a flowchart illustrating an example of a decoding method performed by a decoding device.

[0284] The decoding device 1130 includes a circuit 1131 and a memory 1132 connected to the circuit 1131.

[0285] Circuit 1131 performs the following actions.

[0286] Circuit 1131 acquires encoded data (S1021), the encoded data including: encoding method information (format), indicating one encoding method including first data representing a three-dimensional object and second data representing the three-dimensional object; and identification information, indicating the three-dimensional space containing the three-dimensional object. Next, based on the encoded data, circuit 1131 decodes the first data and second data corresponding to the three-dimensional space (S1022). Next, circuit 1131 renders the first data to generate first prompt data for prompting (S1023). Next, circuit 1131 renders the second data to generate second prompt data for prompting (S1024). Next, circuit 1131 switches from the generated second prompt data to the first prompt data for prompting (S1025). Furthermore, the first prompt data and the second prompt data are, for example, two-dimensional data or three-dimensional data generated by the rendering reconstruction unit 1034.

[0287] Therefore, based on the first and second data corresponding to the three-dimensional space, first and second prompt data are generated, and prompts are given by switching from the second prompt data to the first prompt data. This allows for prompting in a way that does not produce spatial deviation during the switching between the two data representing the three-dimensional object. Thus, the first and second prompt data can be appropriately used to provide prompts.

[0288] For example, the first data is point group data representing the three-dimensional object.

[0289] Therefore, since the prompt is given by switching from the second prompt data to the first prompt data based on the point group data, the prompt can be given in a way that does not produce spatial deviation in the switching between the two data representing the 3D object.

[0290] For example, the second data is mesh data representing the three-dimensional object.

[0291] Therefore, by switching from the second cue data based on grid data to the first cue data for cues, it is possible to switch cues in a way that does not produce spatial deviation when switching between the two data representing the 3D object.

[0292] For example, the second data is three-dimensional model data representing the three-dimensional object. The three-dimensional model data indicates a machine learning model obtained by performing machine learning on multiple sets of views and two-dimensional images.

[0293] Therefore, by switching from the second cue data based on the 3D model data to the first cue data for prompting, it is possible to switch the cue data in a way that does not produce spatial deviation when switching between the two data representing the 3D object.

[0294] For example, the second data is a two-dimensional image obtained when the three-dimensional object is viewed from a specified line of sight.

[0295] Therefore, by switching from second cue data based on a two-dimensional image to first cue data for cues, it is possible to switch between the two data representing a three-dimensional object in a way that does not produce spatial deviation.

[0296] For example, the circuit also receives a switching request for prompt data from the user. In the prompt, the circuit switches from the second prompt data to the first prompt data according to the switching request.

[0297] Therefore, it is possible to switch at a time specified by the user.

[0298] For example, the circuit also receives an operation from the user to change the style of the prompt. In the prompt, the circuit changes the style of the prompt according to the operation, switching from the second prompt data to the first prompt data based on the change.

[0299] Therefore, it can switch at timed intervals corresponding to user actions.

[0300] For example, in the acquisition process, the circuit acquires the encoded data from the encoding device via a communication network. In the prompting process, the circuit switches from the second prompting data to the first prompting data based on the bandwidth of the communication network.

[0301] Therefore, it can switch according to the bandwidth of the communication network. For example, when the bandwidth of the communication network changes from less than the specified bandwidth to more than the specified bandwidth, it can switch from the second prompt data to the first prompt data to make a prompt.

[0302] For example, in the prompt, the circuit switches from the second prompt data to the first prompt data based on the capabilities of the circuit that is available.

[0303] Therefore, it is possible to switch according to the capability of the available circuit. For example, when the capability of the available circuit changes from less than the specified capability to more than the specified capability, it is possible to switch from the second prompt data to the first prompt data to provide a prompt.

[0304] For example, the encoded data includes synchronization information for synchronizing the coordinate system of the first data with the coordinate system of the second data. The circuit, in the prompt, provides prompts for both the first and second prompt data based on the synchronization information.

[0305] Therefore, it is possible to switch from the second prompt data to the first prompt data based on matching the coordinate systems of the first and second prompt data. Thus, it is possible to provide prompts in a way that minimizes spatial deviation when switching between the two data representing a 3D object.

[0306] For example, the circuit further determines whether to synchronize the coordinate system of the first data with the coordinate system of the second data. If the circuit determines that it is necessary to synchronize the coordinate system of the first data with the coordinate system of the second data, the circuit provides a prompt based on the synchronization information for both the first and second prompt data in the prompt.

[0307] Therefore, synchronous processing can be performed when needed and skipped when not needed. This could potentially reduce the processing load.

[0308] For example, the first data and the second data have a common structure in the first data and the second data, respectively.

[0309] Therefore, it is possible to reduce the amount of encoded data. Therefore, it is possible to reduce communication capacity.

[0310] For example, the encoded data includes spatial information for determining the three-dimensional space containing the three-dimensional object. The circuit also obtains an object region indicating a portion of the three-dimensional space. Based on the spatial information, the circuit determines first overlapping data, which is a portion of the first data and overlaps with the object region. In the decoding, the circuit decodes the determined first overlapping data.

[0311] Therefore, for example, the amount of data acquired can be reduced by acquiring only the first overlapping data. This reduces communication capacity. Furthermore, for example, only the first overlapping data can be decoded. This reduces processing load.

[0312] Alternatively, circuit 1131 can also be like... Figure 32 The decoding method shown in the flowchart is followed. Figure 32 This is a flowchart illustrating another example of a decoding method performed by a decoding device.

[0313] Circuit 1131 decodes the encoding information representing the three-dimensional object and indicating a second encoding method different from the first encoding method of the first data (S1031). Circuit 1131 decodes the second data indicating the second encoding method by the encoding information (S1032). The second data is used to generate second prompt data for prompting.

[0314] Therefore, by decoding the second data of the second encoding method indicated by the encoding method information obtained through decoding, second data for generating appropriate prompt data can be obtained.

[0315] Figure 33 This is a diagram illustrating an example of the configuration of an encoding device. Figure 34 This is a flowchart illustrating an example of an encoding method performed by an encoding device.

[0316] The encoding device 1140 includes a circuit 1141 and a memory 1142 connected to the circuit 1141.

[0317] Circuit 1141 performs the following actions.

[0318] Circuit 1141 generates encoding information representing the three-dimensional object and indicating a second encoding method different from the first encoding method of the first data (S1041). Circuit 1141 generates second data indicating the second encoding method of the encoding information (S1042). Circuit 1141 generates a bitstream containing the encoding information and the second data (S1043). The second data is used to generate second prompt data for prompting.

[0319] Therefore, since a bitstream containing encoding method information and second data is generated, the decoding device that obtained the bitstream can obtain second data for generating appropriate second prompt data.

[0320] (Implementation Method 2) A method for generating still images of a subject (three-dimensional object) observed from any viewpoint in still space using a learning-based model, i.e., a three-dimensional data generation model, is explained.

[0321] Figure 35 This is a diagram used to illustrate the processing during the learning of the three-dimensional generative model in Implementation Method 2. Figure 36 This diagram illustrates the process of generating a still image of a subject from any viewpoint using a three-dimensional generative model in Implementation 2.

[0322] Information processing devices acquire 3D data generation models through learning, thereby enabling the generation of still images observed from any viewpoint in static space. For example, there are 3D data generation models generated using methods such as Neural Radiance Fields (NeRF).

[0323] During learning, for example, the information processing device acquires learning data, which includes an image of viewpoint A (correct value) obtained from any viewpoint A, and viewpoint information (camera pose, etc.) of viewpoint A when the image was obtained. The viewpoint information may include viewpoint A and the direction of the line of sight from viewpoint A. The information processing device, for example, uses an evaluation function 1402 to optimize the parameters of the network included in the 3D data generation model, such that the difference between the generated image of viewpoint A output from the 3D data generation model 1401 by inputting the viewpoint information from the aforementioned learning data and the image of viewpoint A as the input image corresponding to viewpoint A is minimized. By performing this learning process using multiple learning data corresponding to multiple different viewpoints, the information processing device can obtain a more accurate 3D data generation model. Learning processing is performed on the learning data corresponding to each of the multiple viewpoints. That is, the same processing as the learning processing for viewpoint A is performed on each viewpoint.

[0324] During generation, if the information processing device inputs viewpoint information, for example, viewpoint B, into the learned 3D data generation model 1403, it outputs a generated image of viewpoint B. If it inputs viewpoint information of viewpoint Z, which is different from viewpoint B, it outputs a generated image of viewpoint Z. The viewpoint information of viewpoint B may include viewpoint B and the viewing direction from viewpoint B. The viewpoint information of viewpoint Z may include viewpoint Z and the viewing direction from viewpoint Z.

[0325] Thus, by learning and obtaining a 3D data generation model 1403, it is possible to generate still images observed from any viewpoint in static space. However, it is not possible to directly generate moving images.

[0326] In addition, Figure 36 The example shown is a 3D data generation model that generates an image of a given viewpoint if that viewpoint information is input. However, this is not a limitation, and the data output from the 3D data generation model can be in any form. For example, the 3D data generation model could also be a network model that outputs learned 3D data of the object space as point cluster data or mesh data. Thus, users can stereoscopically view 3D data such as point cluster data or mesh data in the object space, and can also use the point cluster data or mesh data to measure the dimensions of objects in the object space as 3D data output.

[0327] [Example 1] Figure 37This diagram illustrates the motion image generation method using the three-dimensional data generation model of Embodiment 1 in Embodiment 2. Furthermore, this embodiment describes an example of the configuration of an apparatus and a method for encoding or decoding the three-dimensional data generation models NNt0~NNt5 generated corresponding to times t0~t5. However, it is not limited to this; it can also be applied to an apparatus and method for encoding or decoding the three-dimensional data generation models at any time within any period.

[0328] This embodiment illustrates a method for generating a moving image of an object (subject) observed from any viewpoint using a three-dimensional data generation model. In this method, for example... Figure 37 As shown, by obtaining 3D data generation models corresponding to each time point, still images of an object observed from any viewpoint at each time point can be generated. By arranging the generated still images in chronological order, motion images can be generated. More specifically, in generating motion images for times t0 to t5, multiple 3D data generation models NNt0 to NNt5 corresponding to times t0 to t5 are generated through learning. The viewpoint information (camera pose, etc.) of the viewpoint A from which the motion images are to be generated is input into the generated 3D data generation models NNt0 to NNt5 corresponding to times t0 to t5. Thus, the generated images of viewpoint A at times t0 to t5 are output by the 3D data generation models NNt0 to NNt5. By concatenating these images in time, motion images of the object observed from viewpoint A at times t0 to t5 can be generated.

[0329] However, in this case, maintaining multiple 3D data generation models corresponding to multiple time points requires either a large storage capacity for storing the data of these multiple 3D data generation models on a storage device or a large network bandwidth for transmitting the data of these multiple 3D data generation models over a network. Therefore, the data size can also be reduced by using, for example, NNC (Neural Network Coding) in the MPEG (Moving Picture Experts Group) standard to encode the data of the multiple 3D data generation models corresponding to multiple time points. In this disclosure, a method for more effectively compressing this data is described.

[0330] NNC is shown in Non-Patent Document 1.

[0331] Figure 38 This is a diagram illustrating a first example of the configuration of the encoding device in Embodiment 1 of Embodiment 2.

[0332] The encoding device 1420 includes a three-dimensional data generation model acquisition unit 1421, a buffer unit 1422, and a network model encoding unit 1423.

[0333] The 3D data generation model acquisition unit 1421 acquires learning data from times t0 to t5, and uses this learning data to generate 3D data generation models NNt0 to NNt5 for times t0 to t5 through learning. The learning data includes multiple viewpoint images obtained by photographing an object from one or more viewpoint positions along one or more viewing directions at each time t0 to t5, and one or more viewpoint information indicating one or more viewpoint positions and one or more viewing directions corresponding to the multiple viewpoint images. The one or more viewpoint information may also be the camera position and pose when each of the multiple viewpoint images was photographed. Furthermore, the learning data is not limited to this and may also include information obtained from other sensors. For example, the learning data may also include point cluster data and depth images obtained using LiDAR or TOF sensors at each time. This improves the accuracy of the 3D data generation model obtained through learning.

[0334] The buffer unit 1422 stores the three-dimensional data generation model at time t generated by the three-dimensional data generation model acquisition unit 1421. The buffer unit 1422 is implemented by a storage device such as a memory. The three-dimensional data generation model at time t stored in the buffer unit 1422 can also be used as an initial model when the three-dimensional data generation model acquisition unit 1421 acquires (generates) the three-dimensional data generation model after time t through learning. As a result, the learning time can be shortened and the accuracy of the three-dimensional data generation model after time t can be improved.

[0335] Furthermore, the buffer unit 1422 can also store multiple 3D data generation models corresponding to multiple time points. Therefore, for example, an initial model can be generated based on the multiple 3D data generation models stored in the buffer unit 1422, for example, through averaging or other processing. The 3D data generation model acquisition unit 1421 learns 3D data generation models after time t using this initial model, and can obtain a high-precision 3D data generation model. Furthermore, if the 3D data generation model acquisition unit 1421 does not refer to past 3D data generation models during learning, the encoding device 1420 may not need to include the buffer unit 1422. This reduces the amount of storage used as the buffer unit 1422.

[0336] The network model encoding unit 1423 encodes the three-dimensional data generation models NNt0~NNt5 obtained by the three-dimensional data generation model acquisition unit 1421 and outputs a bit stream.

[0337] Furthermore, as a network model encoding method, the data size can be reduced, for example, by using NNC data encoding from the MPEG standard. That is, the network model encoding unit 1423 uses NNC to encode the three-dimensional data generation models NNt0 to NNt5 and appends the encoding result to the bitstream. In other words, the network model encoding unit 1423 generates encoded data as the encoding result and generates a bitstream containing the encoded data.

[0338] Specifically, the network model encoding unit 1423 first encodes the 3D data generation model NNt0 at time t0 using NNC, and appends the encoding result to the bitstream. Next, the network model encoding unit 1423 encodes the 3D data generation model NNt1 at time t1 using NNC, and appends the encoding result to the bitstream. In this way, the network model encoding unit 1423 can also reduce the amount of encoding by sequentially encoding the 3D data generation model at each time using NNC and appending each encoding result to the bitstream.

[0339] Furthermore, at this time, the network model encoding unit 1423 can also append time information, which indicates the time corresponding to the encoded 3D data generation model, as metadata to the bitstream. Thus, by decoding and referring to the metadata contained in the bitstream, the decoding device can determine the time corresponding to the decoded 3D data generation model and can appropriately generate motion images of the object from any viewpoint.

[0340] In addition, metadata is not limited to time information; it can also include information related to the acquisition (generation) of learning data, or information required by the decoding device to generate motion images.

[0341] For example, the network model encoding unit 1423 may also attach information related to the camera's frame rate when acquiring (generating) learning data as metadata. Thus, the decoding device can decode the frame rate of the generated motion image from the bitstream and appropriately set that frame rate.

[0342] Alternatively, the network model encoding unit 1423 may append the frame number corresponding to each time moment as metadata to the bitstream instead of the time information, and use other parameters to associate each frame number with the time information. For example, the network model encoding unit 1423 may append the time information and frame rate of the first frame as metadata, and the decoding device may calculate the time information of each frame based on this metadata, thereby reducing the amount of encoding corresponding to the time information of each frame.

[0343] Furthermore, the network model encoding unit 1423 can also append viewpoint information from viewpoint images used during learning to the bitstream. Thus, the decoding device can, for example, generate high-quality motion images by preferentially selecting viewpoints close to the viewpoint positions corresponding to the images used during learning. This is because the closer the viewpoint position or time is to the time of learning, the more likely the 3D data generation model is to generate higher-quality viewpoint images.

[0344] Figure 39 This is a diagram illustrating a first example of the configuration of the decoding device in Embodiment 1 of Embodiment 2.

[0345] The decoding device 1425 includes a network model decoding unit 1426 and a rendering unit 1427.

[0346] The network model decoding unit 1426 acquires the bit stream and, based on the acquired bit stream, decodes the three-dimensional data generation model NNt0~NNt5 and metadata such as time information for times t0~t5.

[0347] The rendering unit 1427 uses the 3D data generation models NNt0~NNt5 decoded by the network model decoding unit 1426 and metadata such as time information to generate motion images of viewpoint A based on viewpoint information specified by the user or system. Specifically, the rendering unit 1427 inputs the viewpoint information of viewpoint A into the 3D data generation model NNt0 at time t0 to generate image IMGt0 of viewpoint A at time t0. Next, it inputs the viewpoint information of viewpoint A into the 3D data generation model NNt1 at time t1 to generate image IMGt1 of viewpoint A at time t1. The rendering unit 1427 applies the generation processing of these images at each time to each time t2~t5 to generate images IMGt2~IMGt5 of viewpoint A at times t2~t5. Furthermore, the rendering unit 1427 uses images IMGt0~IMGt5 and metadata such as time information to generate motion images of objects observed from viewpoint A at times t0~t5. The moving image may include, for example, images IMGt0~IMGt5 and cue time information for calculating the cue times of images IMGt0~IMGt5 based on times t0~t5.

[0348] Furthermore, the viewpoint information can change according to time. For example, viewpoint information of viewpoint A can be input into the 3D data generation model NNt0~NNt3 at times t0~t3, and viewpoint information of viewpoint B can be input into the 3D data generation model NNt4~NNt5 at times t4~t5. Thus, the rendering unit 1427 generates multiple images of the object observed from viewpoint A at times t0~t3, and generates multiple images of the object observed from viewpoint B at times t4~t5. In other words, the rendering unit 1427 can generate motion images of the observed object, where the viewpoint switches from viewpoint A to viewpoint B at time t4.

[0349] Furthermore, the rendering unit 1427 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time. Thus, the user can switch between generating moving images and generating still images depending on the application.

[0350] Furthermore, the rendering unit 1427 is not limited to generating moving or still images based on a 3D data generation model. For example, the rendering unit 1427 can also generate point group data or mesh data based on a 3D data generation model, and output the generated point group data or mesh data as dynamic point group data or dynamic mesh data. Thus, users can use dynamic 3D data of audiovisual objects such as HMDs (Head Mount Displays), and can also use the dynamic 3D data to measure the amount of motion of objects.

[0351] Figure 40 This is a diagram illustrating a second example of the configuration of the encoding device in Embodiment 1 of Embodiment 2.

[0352] The encoding device 1430 includes a three-dimensional data generation model acquisition unit 1431, a buffer unit 1432, a difference calculation unit 1433, and a network model encoding unit 1434.

[0353] The three-dimensional data generation model acquisition unit 1431 is the same as the three-dimensional data generation model acquisition unit 1421 of the encoding device 1420.

[0354] The buffer unit 1432 is the same as the buffer unit 1422 of the encoding device 1420, but it differs from the buffer unit 1422 in that it inputs the three-dimensional data generation model stored in the memory or the like as a reference three-dimensional data generation model into the difference calculation unit 1433.

[0355] The difference calculation unit 1433 calculates difference information, which represents the difference between the three-dimensional data generation models NNt0~NNt5 generated by the three-dimensional data generation model acquisition unit 1431 at times t0~t5 and the three-dimensional data generation models (hereinafter referred to as reference three-dimensional data generation models) generated by the three-dimensional data generation model acquisition unit 1431 before each time step. Here, the difference information may include the difference in the weight parameters of the nodes of each network model, etc. For example, the difference calculation unit 1433 obtains the three-dimensional data generation model NNt5 at time t5 from the three-dimensional data generation model acquisition unit 1431 and obtains the three-dimensional data generation model NNt4 at time t4 from the buffer unit 1432 as a reference three-dimensional data generation model.

[0356] Alternatively, the difference calculation unit 1433 can use the three-dimensional data generation model NNt5 and the three-dimensional data generation model NNt4, for example, to calculate the difference (change) between the weight parameters of the nodes in the network model of the three-dimensional data generation model NNt5 and the weight parameters of the nodes in the network model of the three-dimensional data generation model NNt4, and input the difference information representing the difference to the network model encoding unit 1434. Thus, the difference information is encoded by the network model encoding unit 1434. That is, the encoding device 1430 can also perform predictive encoding by encoding the difference between the predicted value and the information related to the network model in the three-dimensional data generation model NNt5 based on the prediction of the three-dimensional data generation model NNt4, thereby reducing the amount of data. Through such predictive encoding, for example, in cases where the changes in the three-dimensional data generation model are small over time, such as when the object hardly moves, the value of the encoded difference becomes smaller, thus improving encoding efficiency. For example, the encoding device 1430 can also be set to RNNt0 = 0 and RNNtn = NNt(n-1) (n is an integer value from 1 to 5), using the previous three-dimensional data generation model as a reference three-dimensional data generation model, and reducing the number of bits through predictive coding.

[0357] Furthermore, in the second example, the encoding device 1430 performs predictive encoding on information related to the network model in the three-dimensional data generation model NNt5 based on information related to the network model in the three-dimensional data generation model NNt4, but is not limited to this. For example, the encoding device 1430 may select a reference three-dimensional data generation model for prediction from one or more three-dimensional data generation models stored in the buffer 1432, and perform predictive encoding using the selected three-dimensional data generation model. In this case, the encoding device 1430 may append information representing the selected three-dimensional data generation model (reference three-dimensional data generation model information) to the bitstream in order to pass the selected three-dimensional data generation model to the decoding device. Thus, the encoding device 1430 can select the optimal reference three-dimensional data generation model from the viewpoint of encoding efficiency, thereby improving encoding efficiency. Furthermore, by decoding the reference three-dimensional data generation model information, the decoding device can appropriately decode the bitstream, which has improved encoding efficiency.

[0358] Furthermore, when the encoding device 1430 performs predictive coding with reference to two or more three-dimensional data generation models stored in the buffer 1432, it can also append information representing the two or more reference three-dimensional data generation models to the bitstream. Thus, the encoding device 1430 can use two or more reference three-dimensional data generation models to improve the coding efficiency of predictive coding. Moreover, the decoding device can appropriately decode the bitstream with improved coding efficiency.

[0359] Furthermore, when the reference 3D data generation model is not stored in the buffer 1432, for example, when encoding the initial 3D data generation model (initial frame) in data order, the encoding device 1430 may encode the 3D data generation model of the processing object without calculating the difference from the predicted value (hereinafter referred to as intra-frame prediction), or it may encode by calculating the difference from the predicted value set to 0. Additionally, when the encoding device 1430 sets a certain time t as a random access point, it can encode the 3D data generation model corresponding to time t through intra-frame prediction, or it may encode by calculating the difference from the predicted value set to 0. Therefore, the decoding device can start decoding the 3D data generation model from the initial 3D data generation model (initial frame) or the random access point in data order, improving the functionality during playback.

[0360] Furthermore, a set of multiple 3D data generation models (multiple frames) can be defined (hereinafter referred to as GOF (Group of Frame)). The first frame of the GOF can also be encoded through intra-frame prediction. Thus, the decoding device can randomly access the first frame of the GOF. In addition, by decoding the first frame of the GOF, functionality such as fast-forward playback can be improved.

[0361] Furthermore, the encoding device 1430 may also append permission information indicating whether inter-GOF prediction referencing is permitted to the bitstream. For example, if the bitstream contains permission information indicating that inter-GOF prediction referencing is prohibited, the decoding device can determine that multiple GOFs can be decoded in parallel. Additionally, for example, by permitting inter-GOF prediction referencing, encoding efficiency can be improved.

[0362] The network model encoding unit 1434 is the same as the network model encoding unit 1423 of the encoding device 1420, but it differs in that it encodes the difference information d0~d5 of the three-dimensional data generation model NNt0~NNt5 input from the difference calculation unit 1433 and outputs a bit stream.

[0363] Furthermore, the encoding device 1430 includes a difference calculation unit 1433 and a network model encoding unit 1434 separately, but it is not limited to this. For example, it may be configured such that the difference calculation unit 1433 is included within the network model encoding unit 1434. That is, the network model encoding unit 1434 may also perform the processing of the difference calculation unit 1433.

[0364] Furthermore, the encoding device 1430 may also append prediction coding information to the bitstream, indicating whether the 3D data generation model was encoded using intra-frame prediction or using a reference 3D data generation model for prediction coding (hereinafter referred to as inter-frame prediction). Thus, by decoding the prediction coding information, the decoding device can appropriately determine whether intra-frame prediction or inter-frame prediction should be used to decode the 3D data generation model.

[0365] Figure 41 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 1 of Embodiment 2.

[0366] The decoding device 1435 includes a network model decoding unit 1436, an addition unit 1437, a buffer unit 1438, and a rendering unit 1439.

[0367] The network model decoding unit 1436 acquires the bit stream and, based on the acquired bit stream, decodes the difference information d0~d5 and other metadata such as time information of the three-dimensional data generation model NNt0~NNt5 at times t0~t5.

[0368] The addition unit 1437 adds the difference information d0~d5 of the three-dimensional data generation model corresponding to times t0~t5, which is decoded by the network model decoding unit 1436, and the reference three-dimensional data generation model RNNt0~RNNt5 obtained from the buffer unit 1438 at the corresponding times to calculate the three-dimensional data generation model NNt0~NNt5. In this way, the decoding device 1435 can also be set to RNNt0=0, RNNtn=NNt(n-1) (n is a value of 1~5), and use the three-dimensional data generation model of the previous time as the reference three-dimensional data generation model for prediction decoding.

[0369] Furthermore, in the second example, the decoding device 1435 separately describes the addition unit 1437 and the network model decoding unit 1436, but it is not limited to this. For example, it could also be a structure in which the addition unit 1437 is included within the network model decoding unit 1436. That is, the network model decoding unit 1436 can also perform the processing of the addition unit 1437.

[0370] Furthermore, if the buffer 1438 does not store a reference 3D data generation model, for example, when decoding the initial 3D data generation model (the first frame) in data order, the decoding device 1435 may perform decoding without adding the difference information to the reference 3D data generation model via the addition unit 1437 and without prediction (hereinafter referred to as intra-frame prediction), or it may add the prediction value set to 0 to the difference information for decoding. Additionally, if the decoding device 1435 sets a certain time t as a random access point, it can decode the 3D data generation model corresponding to time t using intra-frame prediction, or it may add the prediction value set to 0 to the difference information for decoding. Furthermore, if the bitstream contains prediction encoding information indicating that the 3D data generation model to be decoded has been encoded using intra-frame prediction, it can decode the 3D data generation model using intra-frame prediction, or it may add the prediction value set to 0 to the difference information for decoding. Therefore, the decoding device 1435 can begin decoding the 3D data generation model from the 3D data generation model that starts with the data sequence (starting frame), random access points, or 3D data generation models that have been encoded by intra-frame prediction, thereby improving the functionality during reproduction.

[0371] In addition, the decoding device 1435 in the second example performs predictive decoding on information related to the network model in the three-dimensional data generation model NNt5 based on information related to the network model in the three-dimensional data generation model NNt4, but is not limited to this. For example, the decoding device 1435 may select a reference three-dimensional data generation model for prediction from one or more three-dimensional data generation models stored in the buffer 1438, and perform predictive decoding using the selected three-dimensional data generation model. In this case, the decoding device 1435 may also decode the information representing the selected three-dimensional data generation model (reference three-dimensional data generation model information) from the bitstream. Thus, the decoding device 1435 decodes the reference three-dimensional data generation model information from the bitstream generated by the encoding device 1430, which selects the reference three-dimensional data generation model that is optimal from the viewpoint of encoding efficiency, thereby enabling appropriate decoding of the bitstream that improves encoding efficiency.

[0372] Furthermore, when performing predictive decoding with reference to two or more three-dimensional data generation models stored in the buffer 1438, the decoding device 1435 can also decode information representing two or more reference three-dimensional data generation models from the bitstream. Thus, the decoding device 1435 can appropriately decode the bitstream, which improves the coding efficiency of predictive coding, using two or more reference three-dimensional data generation models.

[0373] The rendering unit 1439 is the same as the rendering unit 1427 of the decoding device 1425. The rendering unit 1439 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time.

[0374] [Example 2] Figure 42 This diagram illustrates the motion image generation method using the extended three-dimensional data generation model of Embodiment 2 in Embodiment 2. Furthermore, this embodiment describes an example of the configuration and method of an apparatus for encoding or decoding extended three-dimensional data generation models NNt0-2 and NNt3-5, generated corresponding to periods t0-t2 and t3-t5 respectively, but it is not limited to this; it can also be applied to an apparatus and method for encoding or decoding extended three-dimensional data generation models in any period.

[0375] This embodiment illustrates a method for generating a moving image of an object (subject) observed from any viewpoint using a three-dimensional data generation model. In this method, for example, as... Figure 42In this way, by obtaining a three-dimensional data generation model (hereinafter referred to as the extended three-dimensional data generation model) capable of generating images from any viewpoint within a certain time range (period), it is possible to generate still images of objects observed from any viewpoint at any time within each period. By arranging the generated still images in chronological order, moving images can be generated. Similar to the three-dimensional data generation model in Example 1, the extended three-dimensional data generation model is, for example, a three-dimensional data generation model generated by methods such as NeRF.

[0376] More specifically, when generating motion images from time t0 to t5, an extended 3D data generation model NNt0-2 capable of representing the period from t0 to t2 and an extended 3D data generation model NNt3-5 capable of representing the period from t3 to t5 are generated through learning. The viewpoint information (camera pose, etc.) of the viewpoint A from which the motion images are to be generated is input into the generated extended 3D data generation models NNt0-2 and NNt3-5. Thus, the generated images of viewpoint A from time t0 to t5 are output by the extended 3D data generation models NNt0-2 and NNt3-5. By concatenating these images temporally, motion images from time t0 to t5, in which the object is observed from viewpoint A, can be generated.

[0377] However, in this case, maintaining the extended 3D data generation model corresponding to each period (time period) requires either a large storage capacity to store the extended 3D data generation model data in a storage device or a large network bandwidth to transmit the data of multiple 3D data generation models over a network. Therefore, the data size can also be reduced by using, for example, NNC (Neural Network Coding) in the MPEG (Moving Picture Experts Group) standard to encode the extended 3D data generation model corresponding to each period. In this disclosure, a method for more effectively compressing this data is described.

[0378] Furthermore, based on the above configuration, the information processing device can generate any viewpoint image at any time within the period t0-t5. For example, when acquiring the extended 3D data generation model NNt0-2, the information processing device generates the extended 3D data generation model NNt0-2 by using multi-viewpoint images captured at times t0, t1, and t2 as learning data, and by machine learning based on the camera poses corresponding to the multi-viewpoints. Moreover, when generating a motion image of viewpoint A, the information processing device can generate not only viewpoint images A at times t0, t1, and t2, but also images of any viewpoint at times t0.5 and t1.5 between times t0, t1, and t2. Time t0.5 is the time between time t0 and time t1, and time t1.5 is the time between time t1 and time t2.

[0379] Therefore, not only the time corresponding to the image during learning, the information processing device can also generate an image of any viewpoint corresponding to a time offset from the time corresponding to the image during learning, thus enabling the generation of motion images of viewpoint A at a high frame rate.

[0380] Furthermore, as learning data for the extended 3D data generation model NNt0-2, the information processing device can learn not only the learning data at times t0, t1, and t2, but also, for example, the learning data at time t3. Thus, it is possible to generate viewpoint images from any viewpoint after time t2 with high precision, such as an image from any viewpoint at time t2.5.

[0381] Furthermore, as learning data for the extended 3D data generation model NNt3-5, the information processing device can learn not only the learning data corresponding to times t3, t4, and t5, but also, for example, supplement the learning data corresponding to times t2 and t6. Thus, the information processing device can generate images from any viewpoint before time t3 or from any viewpoint after time t5 with high precision. Moreover, as a switching point of the extended 3D data generation model, for example, in the above example, when generating a viewpoint image at time 2.5 between time t2 and t3, which is the switching point between the extended 3D data generation model NNt0-2 and the extended 3D data generation model NNt3-5, the information processing device can also generate viewpoint images at time t2.5 using both the extended 3D data generation model NNt0-2 and the extended 3D data generation model NNt3-5, and generate the average image of the two generated viewpoint images at time t2.5 as the viewpoint image at time t2.5. Thus, a high-precision viewpoint image at time t2.5 can be generated.

[0382] In this way, by specifying the time and viewpoint information within the period corresponding to the extended three-dimensional data generation model, the information processing device can generate an image of the object observed from the specified viewpoint at the specified time.

[0383] Figure 43 This is a diagram illustrating a first example of the configuration of the encoding device in Embodiment 2 of Implementation 2.

[0384] The encoding device 1450 includes an extended three-dimensional data generation model acquisition unit 1451, a buffer unit 1452, and a network model encoding unit 1453.

[0385] The extended 3D data generation model acquisition unit 1451 acquires learning data for each period t0~t2 and t3~t5, from time t0 to t5. Using the acquired learning data for each period, it generates an extended 3D data generation model NNt0-2 for period t0~t2 and an extended 3D data generation model NNt3-5 for period t3~t5 through learning. The learning data includes multiple viewpoint images of an object captured from one or more viewpoint positions along one or more viewing directions at each time t0~t5, and one or more viewpoint information representing one or more viewpoint positions and one or more viewing directions corresponding to the multiple viewpoint images. The one or more viewpoint information may be the position and pose of the camera when capturing each of the multiple viewpoint images. Furthermore, the learning data is not limited to this and may also include information obtained from other sensors. For example, the learning data may also include point cluster data and depth images acquired at each time using a LiDAR or TOF sensor. As a result, the accuracy of the extended 3D data generation model obtained through learning can be improved.

[0386] The buffer unit 1452 stores the extended three-dimensional data generation model for the period tm-n, from time tm (m is an integer) to time tn (n is an integer greater than m), generated by the extended three-dimensional data generation model acquisition unit 1451. The buffer unit 1452 is implemented using a storage device such as a memory. The extended three-dimensional data generation model for the period tm-n stored in the buffer unit 1452 can also be used as an initial model when the extended three-dimensional data generation model acquisition unit 1451 acquires (generates) the extended three-dimensional data generation model for the period after tm-n through learning. As a result, the learning time can be shortened and the accuracy of the extended three-dimensional data generation model for the period after tm-n can be improved.

[0387] Furthermore, the buffer unit 1452 can also store multiple extended 3D data generation models corresponding to multiple periods. Thus, for example, an initial model can be generated based on the multiple extended 3D data generation models stored in the buffer unit 1452, for example, through averaging or other processing. The extended 3D data generation model acquisition unit 1451 learns extended 3D data generation models for periods after period tm-n by using this initial model, and can obtain a high-precision extended 3D data generation model. Furthermore, if the extended 3D data generation model acquisition unit 1451 does not refer to extended 3D data generation models of past periods during learning, the encoding device 1450 may not need to include the buffer unit 1452. This reduces the amount of storage used as the buffer unit 1452.

[0388] The network model encoding unit 1453 encodes the extended three-dimensional data generation model NNt0-2 and the extended three-dimensional data generation model NNt3-5 obtained by the extended three-dimensional data generation model acquisition unit 1451 and outputs a bit stream.

[0389] Furthermore, as a network model encoding method, the data size can be reduced, for example, by using NNC data encoding from the MPEG standard. That is, the network model encoding unit 1453 uses NNC to encode the extended 3D data generation model NNt0-2 and the extended 3D data generation model NNt3-5, and appends the encoding result to the bitstream. In other words, the network model encoding unit 1453 generates encoded data as the encoding result and generates a bitstream containing the encoded data.

[0390] Specifically, the network model encoding unit 1453 first uses NNC to encode the extended 3D data generation model NNt0-2 for periods t0 to t2, and appends the encoding result to the bitstream. Next, the network model encoding unit 1453 uses NNC to encode the extended 3D data generation model NNt3-5 for periods t3 to t5, and appends the encoding result to the bitstream. In this way, the network model encoding unit 1453 can sequentially encode the extended 3D data generation model for each period using NNC and append each encoding result to the bitstream, thereby reducing the amount of encoding.

[0391] Furthermore, at this time, the network model encoding unit 1453 can also append time information, representing the period to which the encoded extended 3D data generation model corresponds, as metadata to the bitstream. Thus, by decoding and referring to the metadata contained in the bitstream, the decoding device can determine which period the decoded extended 3D data generation model corresponds to, and can appropriately generate motion images of the object from any viewpoint.

[0392] Furthermore, the network model encoding unit 1453 can also generate information as time information indicating which period of viewpoint image the extended 3D data generation model can generate, and append the generated time information as metadata to the bitstream. Thus, in the decoding device, by decoding this metadata, the period during which the extended 3D data generation model can generate viewpoint images can be determined, and motion images can be generated appropriately.

[0393] In addition, metadata is not limited to time information; it can also include information related to the acquisition (generation) of learning data, or information required by the decoding device to generate motion images.

[0394] For example, the network model encoding unit 1453 may also attach information related to the camera's frame rate when acquiring (generating) the learning data as metadata. Thus, the decoding device can decode the frame rate of the generated motion image from the bitstream and appropriately set that frame rate.

[0395] Alternatively, the network model encoding unit 1453 may append the frame number corresponding to each period as metadata to the bitstream instead of the time information, and use other parameters to associate each frame number with the time information. For example, the network model encoding unit 1453 may append the time information and frame rate of the first frame as metadata, and the decoding device may calculate the time information of each frame based on this metadata, thereby reducing the amount of encoding corresponding to the time information of each frame.

[0396] Furthermore, the network model encoding unit 1453 can also append viewpoint information from viewpoint images used in the learning process, or time information indicating the time when the viewpoint image was captured, to the bitstream. Thus, for example, the decoding device can generate high-quality motion images by preferentially selecting viewpoints close to the viewpoint positions corresponding to images used in the learning process, or times close to the times corresponding to images used in the learning process. This is because the closer the viewpoint position or time is to the time of learning, the more likely the extended 3D data generation model is to generate higher-quality viewpoint images.

[0397] Figure 44 This is a diagram illustrating a first example of the configuration of the decoding device in Embodiment 2 of Implementation 2.

[0398] The decoding device 1455 includes a network model decoding unit 1456 and a rendering unit 1457.

[0399] The network model decoding unit 1456 acquires the bit stream and, based on the acquired bit stream, decodes metadata such as the extended three-dimensional data generation model NNt0-2 during period t0~t2 and the extended three-dimensional data generation model NNt3-5 during period t3~t5, as well as the time information corresponding to these extended three-dimensional data generation models NNt0-2 and NNt3-5.

[0400] The rendering unit 1457 uses metadata such as the extended 3D data generation models NNt0-2 and NNt3-5 decoded by the network model decoding unit 1456 and time information to generate motion images of viewpoint A based on viewpoint information specified by the user or system. Specifically, the rendering unit 1457 inputs the viewpoint information of viewpoint A and the time within the period t0~t2 into the extended 3D data generation model NNt0-2, and generates images IMGt0 of viewpoint A at time t0, IMGt1 of viewpoint A at time t1, and IMGt2 of viewpoint A at time t2. The rendering unit 1457 applies the image generation processing of the period t0~t2 to the extended 3D data generation model NNt3-5 for the period t3~t5 to generate images IMGt3~IMGt5 of viewpoint A at times t3~t5. Furthermore, the rendering unit 1457 uses metadata such as images IMGt0~IMGt5 and time information to generate motion images of the object observed from viewpoint A at times t0~t5. The motion image may, for example, include images IMGt0~IMGt5 and prompt time information for calculating prompt times based on the images IMGt0~IMGt5 at times t0~t5.

[0401] Furthermore, the viewpoint information can change according to time. For example, viewpoint information of viewpoint A can be input into the extended 3D data generation model NNt0-2 during the period t0~t2, and viewpoint information of viewpoint B can be input into the extended 3D data generation model NNt3-5 during the period t3~t5. As a result, the rendering unit 1457 generates multiple images of the object observed from viewpoint A during time t0~t2, and generates multiple images of the object observed from viewpoint B during time t3~t5. That is, the rendering unit 1457 can generate motion images of the object observed, with the viewpoint switching from viewpoint A to viewpoint B at time t3.

[0402] Furthermore, the rendering unit 1457 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time. Thus, the user can switch between generating moving images and generating still images depending on the application.

[0403] Furthermore, the rendering unit 1457 is not limited to generating moving or still images based on the extended 3D data generation model. For example, the rendering unit 1457 can also generate point group data or mesh data that the extended 3D data generation model can represent, and output the generated point group data or mesh data as dynamic point group data or dynamic mesh data. Thus, users can use dynamic 3D data of audiovisual objects such as HMDs (Head Mount Displays), and can also use the dynamic 3D data to measure the amount of motion of objects.

[0404] Figure 45 This is a diagram illustrating a second example of the configuration of the encoding device in Embodiment 2 of Implementation 2.

[0405] The encoding device 1460 includes an extended three-dimensional data generation model acquisition unit 1461, a buffer unit 1462, a difference calculation unit 1463, and a network model encoding unit 1464.

[0406] The extended three-dimensional data generation model acquisition unit 1461 is the same as the extended three-dimensional data generation model acquisition unit 1451 of the encoding device 1450.

[0407] The buffer unit 1462 is the same as the buffer unit 1452 of the encoding device 1450, but it differs from the buffer unit 1452 in that it inputs the extended three-dimensional data generation model stored in the memory or the like as a reference extended three-dimensional data generation model into the difference calculation unit 1463.

[0408] The difference calculation unit 1463 calculates difference information, which represents the difference between the extended 3D data generation model NNt0-2 generated by the extended 3D data generation model acquisition unit 1461 for periods t0 to t2 and the extended 3D data generation model NNt3-5 for periods t3 to t5, respectively, and the extended 3D data generation model generated by the extended 3D data generation model acquisition unit 1461 before each period (hereinafter referred to as the reference extended 3D data generation model). Here, the difference information may include the difference in the weight parameters of the nodes of each network model, etc. For example, the difference calculation unit 1463 obtains the extended 3D data generation model NNt3-5 for periods t3 to t5 from the extended 3D data generation model acquisition unit 1461, and obtains the extended 3D data generation model NNt0-2 for periods t0 to t2 from the buffer unit 1462 as the reference extended 3D data generation model.

[0409] The difference calculation unit 1463 can also use extended 3D data generation models NNt3-5 and NNt0-2, for example, to calculate the difference (change) between the weight parameters of the nodes in the network model of extended 3D data generation model NNt3-5 and the weight parameters of the nodes in the network model of extended 3D data generation model NNt0-2, and input the difference information representing the difference to the network model encoding unit 1464. Thus, the difference information is encoded by the network model encoding unit 1464. That is, the encoding device 1460 can also reduce the amount of data by predictive encoding, which encodes the difference between the predicted and predicted values, based on the information related to the network model in extended 3D data generation model NNt0-2 predicted by extended 3D data generation model NNt3-5. Through such predictive encoding, for example, in cases where the changes in the extended 3D data generation model are small over time, such as when the object hardly moves, the value of the encoded difference becomes smaller, thus improving encoding efficiency. For example, the encoding device 1460 can also be set to RNNt0-2 = 0 and RNNt3-5 = NNt0-2, and the extended three-dimensional data generation model of the previous time period can be used as a reference extended three-dimensional data generation model to reduce the number of bits through predictive coding.

[0410] Furthermore, in the second example, the encoding device 1460 performs predictive encoding on information related to the network model in the extended 3D data generation model NNt3-5 based on information related to the network model in the extended 3D data generation model NNt0-2, but is not limited to this. For example, the encoding device 1460 may select a reference extended 3D data generation model for prediction from one or more extended 3D data generation models stored in the buffer 1462, and perform predictive encoding using the selected extended 3D data generation model. In this case, in order to pass the selected extended 3D data generation model to the decoding device, the encoding device 1460 may also append information representing the selected extended 3D data generation model (reference extended 3D data generation model information) to the bitstream. Thus, the encoding device 1460 can select the optimal reference extended 3D data generation model from the viewpoint of encoding efficiency, thereby improving encoding efficiency. Furthermore, by decoding the reference extended 3D data generation model information, the decoding device can appropriately decode the bitstream with improved encoding efficiency.

[0411] Furthermore, when the encoding device 1460 performs predictive coding with reference to two or more extended three-dimensional data generation models stored in the buffer 1462, it can also append information representing the two or more reference extended three-dimensional data generation models to the bitstream. Thus, the encoding device 1460 can use two or more reference extended three-dimensional data generation models to improve the coding efficiency of predictive coding. Furthermore, the decoding device can appropriately decode the bitstream with improved coding efficiency.

[0412] Furthermore, when the reference extended 3D data generation model is not stored in the buffer 1462, for example, when encoding the initial extended 3D data generation model (initial frame) in data order, the encoding device 1460 can encode the extended 3D data generation model of the processing object without calculating the difference from the predicted value (hereinafter referred to as intra-frame prediction), or it can calculate the difference from the predicted value set to 0 for encoding. Additionally, when a certain period tm-n is set as a random access point, the encoding device 1460 can encode the extended 3D data generation model corresponding to period tm-n through intra-frame prediction, or it can calculate the difference from the predicted value set to 0 for encoding. Therefore, the decoding device can decode the extended 3D data generation model starting from the initial extended 3D data generation model (initial frame) or the random access point in data order, improving functionality during playback.

[0413] Furthermore, a set of multiple extended 3D data generation models (multiple frames) can be defined (hereinafter referred to as GOF (Group of Frame)). The first frame of the GOF can also be encoded through intra-frame prediction. Thus, the decoding device can randomly access the first frame of the GOF. In addition, by decoding the first frame of the GOF, functionality such as fast-forward playback can be improved.

[0414] Furthermore, the encoding device 1460 may also append permission information indicating whether inter-GOF prediction referencing is permitted to the bitstream. For example, if the bitstream contains permission information indicating that inter-GOF prediction referencing is prohibited, the decoding device can determine that multiple GOFs can be decoded in parallel. Additionally, for example, by permitting inter-GOF prediction referencing, encoding efficiency can be improved.

[0415] The network model encoding unit 1464 is the same as the network model encoding unit 1453 of the encoding device 1450, but it differs in that it encodes the difference information d0-2 and d3-5 of the extended three-dimensional data generation models NNt0-2 and NNt3-5 input from the difference calculation unit 1463 and outputs a bit stream.

[0416] Furthermore, the encoding device 1460 includes a difference calculation unit 1463 and a network model encoding unit 1464 separately, but it is not limited to this. For example, it may be configured to include the difference calculation unit 1463 within the network model encoding unit 1464. That is, the network model encoding unit 1464 may also perform the processing of the difference calculation unit 1463.

[0417] Furthermore, the encoding device 1460 may also append prediction coding information to the bitstream, indicating whether the extended 3D data generation model was encoded using intra-frame prediction or using a reference extended 3D data generation model for prediction coding (hereinafter referred to as inter-frame prediction). Thus, by decoding the prediction coding information, the decoding device can appropriately determine whether intra-frame prediction or inter-frame prediction should be used to decode the extended 3D data generation model.

[0418] Figure 46 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 2 of Implementation 2.

[0419] The decoding device 1465 includes a network model decoding unit 1466, an addition unit 1467, a buffer unit 1468, and a rendering unit 1469.

[0420] The network model decoding unit 1466 acquires the bit stream and, based on the acquired bit stream, decodes the difference information d0-2, d3-5, and time information of the extended three-dimensional data generation model NNt0-2 and the extended three-dimensional data generation model NNt3-5 during the period t0~t2.

[0421] The addition unit 1467 adds the difference information d0-2, d3-5 of the extended 3D data generation models NNt0-2 and NNt3-5 corresponding to periods t0~t2 and t3~t5, decoded by the network model decoding unit 1466, and the reference extended 3D data generation models RNNt0-2 and RNNt3-5 obtained from the buffer unit 1468 during the corresponding periods to calculate the extended 3D data generation models NNt0-2 and NNt3-5. Thus, the decoding device 1465 can also be set to RNNt0-2 = 0 and RNNt3-5 = NNt0-2, using the extended 3D data generation model from the previous time period as a reference extended 3D data generation model for prediction decoding.

[0422] Furthermore, the decoding device 1465 in the second example separately describes the addition unit 1467 and the network model decoding unit 1466, but it is not limited to this. For example, it could be configured to include the addition unit 1467 within the network model decoding unit 1466. That is, the network model decoding unit 1466 can also perform the processing of the addition unit 1467.

[0423] Furthermore, if the buffer 1468 does not store a reference extended 3D data generation model, for example, when decoding the initial extended 3D data generation model (the first frame) in data order, the decoding device 1465 may perform decoding without adding the difference information to the reference extended 3D data generation model by the addition unit 1467 and without prediction (hereinafter referred to as intra-frame prediction), or it may add the prediction value set to 0 to the difference information to perform decoding. Additionally, if the decoding device 1465 sets a certain period tm-n as a random access point, it can decode the extended 3D data generation model corresponding to period tm-n through intra-frame prediction, or it may add the prediction value set to 0 to the difference information to perform decoding. Furthermore, if the bitstream contains prediction encoding information indicating that the extended 3D data generation model of the decoding target has been encoded through intra-frame prediction, it can decode the extended 3D data generation model through intra-frame prediction, or it may add the prediction value set to 0 to the difference information to perform decoding. Therefore, the decoding device 1465 can start decoding the extended three-dimensional data generation model from the extended three-dimensional data generation model (starting frame) that begins in data order, random access points, or extended three-dimensional data generation models that have been encoded by intra-frame prediction, thereby improving the functionality during reproduction.

[0424] In addition, the decoding device 1465 in the second example performs predictive decoding on information related to the network model in the extended 3D data generation model NNt3-5 based on information related to the network model in the extended 3D data generation model NNt0-2, but is not limited to this. For example, the decoding device 1465 may also select a reference extended 3D data generation model for prediction from one or more extended 3D data generation models stored in the buffer 1468, and perform predictive decoding using the selected extended 3D data generation model. In this case, the decoding device 1465 may also decode information representing the selected extended 3D data generation model (reference extended 3D data generation model information) from the bitstream. Thus, the decoding device 1465 decodes the reference extended 3D data generation model information from the bitstream generated by the encoding device 1460, which selects the reference extended 3D data generation model that is optimal from the viewpoint of encoding efficiency, thereby enabling appropriate decoding of the bitstream with improved encoding efficiency.

[0425] Furthermore, when performing predictive decoding with reference to two or more extended three-dimensional data generation models stored in the buffer 1468, the decoding device 1465 can also decode information representing two or more reference extended three-dimensional data generation models from the bitstream. Thus, the decoding device 1465 can appropriately decode the bitstream that improves the coding efficiency of predictive coding using two or more reference extended three-dimensional data generation models.

[0426] The rendering unit 1469 is the same as the rendering unit 1427 of the decoding device 1425. The rendering unit 1469 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time.

[0427] [Variation Example] Additionally, the encoding device 1460 may include information related to the number of images that the extended 3D data generation model can generate (i.e., the upper limit of the number of images) in the metadata attached to the bitstream of the extended 3D data generation model during the period tm-n. Thus, the decoding device 1465 can know the number of images that the decoded extended 3D data generation model can generate, for example, by appropriately setting the frame rate of the motion picture to be generated, and also, for example, by calculating the number of delayed frames until the motion picture is displayed.

[0428] Additionally, the encoding device 1460 can also append information indicating the time unit (i.e., the smallest time unit) to which the extended 3D data generation model can generate viewpoint images as time information appended to the bitstream. For example, as time information, the encoding device 1460 can append information such as whether the viewpoint image can be generated up to a time unit of 1 msec or 1 μmsec to the bitstream. Thus, the decoding device 1465 can determine the time unit to which the viewpoint information is generated and can accordingly generate high frame rate motion images or 3D data.

[0429] Furthermore, the encoding device 1460 can also append information about the extended 3D data generation model to the metadata of the bitstream attached to the extended 3D data generation model during the period tm-n. For example, by appending the timing information or viewpoint information of the image used for learning as metadata to the bitstream, the encoding device 1460 can, by decoding the metadata, know the timing or viewpoint information at which the extended 3D data generation model can generate viewpoint images with high quality, thereby enabling the production of high-quality motion pictures.

[0430] Furthermore, the width of the viewpoint image generated by the extended 3D data generation model can also be... Figure 47 The width can be switched dynamically as shown. Specifically, the width can also be switched according to the subject. Figure 47 This is a diagram illustrating a motion image generation method using an extended three-dimensional data generation model, which is used to explain a variation of Embodiment 2.

[0431] For example, in scenes with many stationary objects in the subject (scenes where the number of stationary objects in multiple subjects is a first number or more, or scenes where the volume (area) occupied by the stationary objects in multiple subjects is a first amount or more), the encoding device 1460 can generate an extended 3D data generation model with a long duration capable of generating viewpoint images with high image quality by expanding the width of the learning data used during learning (i.e., extending the duration). For example, in scenes with many moving objects in the subject (scenes where the number of moving objects in multiple subjects is a first number or more, or scenes where the volume (area) occupied by the moving objects in multiple subjects is a first amount or more), the encoding device 1460 can generate an extended 3D data generation model with a narrow duration capable of generating viewpoint images with high image quality, even for moving objects, by narrowing the width of the learning data used during learning.

[0432] Alternatively, the encoding device 1460 can use learning data for a certain period tm-n (e.g., a Group of Frames (GOF) representing a set of frames within period tm-n in a learning image) to generate an extended 3D data generation model NNtm-n for that period tm-n. In this case, the encoding device 1460 buffers the learning image frames for period tm-n to generate the extended 3D data generation model and performs compressed transmission, thus generating a transmission delay of the GOF size. The encoding device 1460 can also append information related to this transmission delay, such as the number of GOF frames and the number of delayed frames, to the bitstream. Therefore, the decoding device 1465 can obtain the delay information by decoding the bitstream and can appropriately reproduce the motion image or 3D data taking the delay into account.

[0433] Furthermore, the above embodiments illustrate an example of generating a still image of an arbitrary viewpoint at a given time or period using a 3D data generation model or an extended 3D data generation model, but are not necessarily limited to this. For example, a 3D data generation model or an extended 3D data generation model, such as... Figure 48 As shown, it generates (outputs) 3D data such as point cluster data or mesh data at a specific moment within a certain period. This allows users to perform dimensional measurements of objects or obtain 3D data with higher audiovisual detail. Figure 48 This is a diagram used to illustrate a motion image generation method based on a modified example of embodiment 2 involving a three-dimensional data generation model.

[0434] Furthermore, encoding devices 1420 and 1460 can also include recommended output formats corresponding to the use case in the metadata of the bitstream, indicating the output format of images, point group data, grid data, etc. Thus, the user can select the recommended output format based on the use case.

[0435] Furthermore, encoding devices 1420 and 1460 can also attach more than one viewpoint information to the metadata of the bitstream appended to the 3D data generation model or the extended 3D data generation model. For example, encoding devices 1420 and 1460 can consider including recommended viewpoint information for audiovisual objects or user viewpoint information when acquiring learning data in the metadata. Thus, decoding devices 1425 and 1465 can generate motion graphics or 3D data using viewpoint information selected from more than one viewpoint information appended to the bitstream based on user intent, etc.

[0436] Alternatively, a default viewpoint can be predetermined based on more than one viewpoint. Alternatively, if no user-specified viewpoint is provided, the decoding devices 1425 and 1465 can use the predetermined default viewpoint to generate motion images or 3D data. Thus, the decoding devices 1425 and 1465 can automatically generate motion images or 3D data even without user specification.

[0437] As an example of using this implementation method, there are the following usage methods.

[0438] First, the encoding devices 1420 and 1460 use cameras or sensors to acquire data of a dynamic object that they want to send to a distance, and use the data of the dynamic object as learning data to generate a three-dimensional data generation model or an extended three-dimensional data generation model of the dynamic object.

[0439] Next, the encoding devices 1420 and 1460 encode the three-dimensional data generation model or the extended three-dimensional data generation model using the encoding method described in this embodiment, and transmit the bit stream containing the encoding result to a remote location.

[0440] Then, decoding devices 1425 and 1465 decode the bitstream received from a distance, and use the decoded 3D data of the dynamic object to generate a 3D model or an extended 3D data generation model to generate motion images or 3D data from any viewpoint. The generated 3D data can then be used for appreciation or measurement purposes. In this way, this embodiment can also be applied to all use cases of remotely sharing information in a certain space.

[0441] Furthermore, when there are more than one object in a certain space that needs to be sent to a distant location, the 3D data generation modeling, encoding and transmission, decoding, and rendering processes described in this embodiment can be applied to each object separately. For example, dynamic objects in the foreground and static objects in the background existing in a certain space can be generated and modeled in 3D data and encoded and transmitted separately. As a result, the optimal 3D data generation modeling or encoding method can be applied to each object, thereby improving encoding efficiency.

[0442] Furthermore, it is not necessarily limited to this; multiple objects can also be treated as a single object, and the 3D data generation, modeling, encoding, transmission, decoding, and rendering processes described in this embodiment can be applied to each object separately. This allows for the transmission of multiple objects to a remote location while minimizing processing overhead.

[0443] Figure 49 This is a diagram illustrating an example of the configuration of the encoding device in Embodiment 2. Figure 50 This is a flowchart illustrating an example of the encoding method of the encoding device in Embodiment 2.

[0444] The encoding device 1470 includes circuitry 1471 and memory 1472. The encoding device 1470 is a device that implements the encoding devices 1420 and 1460.

[0445] Circuit 1471 performs the following actions.

[0446] Circuit 1471 acquires a first three-dimensional data generation model (e.g., three-dimensional data generation model NNt0) corresponding to a first time point (e.g., time t0) and a second three-dimensional data generation model (e.g., three-dimensional data generation model NNt1) corresponding to a second time point (e.g., time t1) (S1401). Circuit 1471 generates a bitstream by encoding the acquired first and second three-dimensional data generation models (S1402). The first and second three-dimensional data generation models output two-dimensional images of the subject as viewed from the viewpoint and the viewing direction, respectively, when input with viewpoint information including the viewpoint and the viewing direction.

[0447] Therefore, it is possible to generate a bitstream containing a first three-dimensional data generation model that generates a two-dimensional image corresponding to a first moment based on arbitrary viewpoint information and a second three-dimensional data generation model that generates a two-dimensional image corresponding to a second moment. Thus, it is possible to generate a bitstream that compresses the data of the motion image obtained from an arbitrary viewpoint. Therefore, it is possible to reduce the storage capacity used to store the data of the motion image obtained from an arbitrary viewpoint, or the network bandwidth used to transmit the data.

[0448] For example, the first 3D data generation model and the second 3D data generation model are learning models that use neural networks.

[0449] For example, the bitstream includes first time information representing the first time and second time information representing the second time.

[0450] For example, the bitstream includes a first frame number corresponding to the first time point and a second frame number corresponding to the second time point.

[0451] For example, the bitstream contains frame rate information related to the frame rate of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model. The multiple learning images are 2D images obtained by capturing images at multiple different timings.

[0452] For example, the bitstream contains viewpoint information, which includes the viewpoint and line-of-sight direction of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model.

[0453] For example, the plurality of learning images are two-dimensional images obtained by photographing the subject from mutually different viewpoints and viewing directions. The viewpoint information includes the mutually different viewpoints and viewing directions.

[0454] For example, circuit 1471 calculates difference information representing the difference between the first three-dimensional data generation model and the second three-dimensional data generation model in the encoding of the second three-dimensional data generation model. The bitstream contains the difference information.

[0455] For example, the difference includes the difference between the weight parameters corresponding to the nodes contained in the first 3D data generation model and the second 3D data generation model.

[0456] For example, the bitstream includes reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.

[0457] For example, the first time point corresponds to a random access point. The first 3D data generation model is encoded either by intra-frame prediction or by inter-frame prediction with a prediction value of 0.

[0458] For example, the first 3D data generation model and the second 3D data generation model are contained in one of a plurality of sets. The first 3D data generation model is the first in the data order among the plurality of 3D data generation models contained in the set.

[0459] For example, the bitstream contains licensing information indicating whether the 3D data generation model is permitted to reference other 3D data generation models in the encoding of each of the plurality of 3D data generation models.

[0460] For example, the first three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt0-2) corresponds to a first period (e.g., period t0~t2) that includes the first time point (e.g., time t0). The second three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt3-5) corresponds to a second period (e.g., period t3~t5) that includes the second time point (e.g., time t3).

[0461] For example, the multiple first learning images used in the generation of the first three-dimensional data generation model are two-dimensional images obtained by taking pictures at multiple different timings during the first period.

[0462] For example, when the first three-dimensional data generation model is input with a time contained in the first period, it outputs a two-dimensional image of the subject at the input time.

[0463] For example, the bitstream contains information indicating the upper limit of the number of images that the first 3D data generation model can generate.

[0464] For example, the bitstream contains first information associated with the plurality of first learning images. This first information includes multiple viewpoints and multiple line-of-sight directions corresponding to the plurality of first learning images, as well as multiple different timings.

[0465] For example, the first period or the second period is dynamically determined based on the subject.

[0466] For example, circuit 1471 saves the generated first three-dimensional data generation model in memory 1472. Based on the first three-dimensional data generation model saved in memory 1472, circuit 1471 generates the second three-dimensional data generation model.

[0467] For example, circuit 1471 stores the generated first 3D data generation model and the second 3D data generation model in memory 1472. Based on the first 3D data generation model and the second 3D data generation model stored in memory 1472, circuit 1471 generates an initial model. Based on the initial model, circuit 1471 generates a third 3D data generation model (e.g., 3D data generation model NNt2) corresponding to a third time point (e.g., time t2).

[0468] Figure 51 This is a diagram illustrating an example of the configuration of the decoding device in Embodiment 2. Figure 52 This is a flowchart illustrating an example of a decoding method of the decoding device in Embodiment 2.

[0469] The decoding device 1480 includes circuitry 1481 and memory 1482. The decoding device 1480 is a device that implements the decoding devices 1425 and 1465.

[0470] Circuit 1481 performs the following actions.

[0471] Circuit 1481 acquires the bitstream (S1411). Circuit 1481 decodes from the bitstream a first three-dimensional data generation model (e.g., three-dimensional data generation model NNt0) corresponding to a first time moment (e.g., time t0) and a second three-dimensional data generation model (e.g., three-dimensional data generation model NNt1) corresponding to a second time moment (e.g., time t1) (S1412). When the first three-dimensional data generation model and the second three-dimensional data generation model are input with viewpoint information including the viewpoint and the viewing direction, they output a two-dimensional image of the subject as viewed from the viewpoint and the viewing direction, respectively.

[0472] Therefore, based on the compressed bitstream of motion image data obtained from any viewpoint, it is possible to decode a first three-dimensional data generation model that generates a two-dimensional image corresponding to a first moment based on arbitrary viewpoint information, and a second three-dimensional data generation model that generates a two-dimensional image corresponding to a second moment. Thus, it is possible to appropriately decode the bitstream that reduces the storage capacity for storing the motion image data obtained from any viewpoint or the network bandwidth for transmitting the data.

[0473] For example, the first 3D data generation model and the second 3D data generation model are learning models that use neural networks.

[0474] For example, the bitstream includes first time information representing the first time and second time information representing the second time.

[0475] For example, the bitstream includes a first frame number corresponding to the first time point and a second frame number corresponding to the second time point.

[0476] For example, the bitstream contains frame rate information related to the frame rate of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model. The multiple learning images are 2D images obtained by capturing images at multiple different timings.

[0477] For example, the bitstream contains viewpoint information, which includes the viewpoint and line-of-sight direction of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model.

[0478] For example, the plurality of learning images are two-dimensional images obtained by photographing the subject from mutually different viewpoints and viewing directions. The viewpoint information includes the mutually different viewpoints and viewing directions.

[0479] For example, the bitstream contains difference information representing the difference between the first 3D data generation model and the second 3D data generation model.

[0480] For example, the difference includes the difference between the weight parameters corresponding to the nodes contained in the first 3D data generation model and the second 3D data generation model.

[0481] For example, the bitstream includes reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.

[0482] For example, the first time point corresponds to a random access point. The first 3D data generation model is encoded either by intra-frame prediction or by inter-frame prediction with a prediction value of 0.

[0483] For example, the first 3D data generation model and the second 3D data generation model are contained in one of a plurality of sets. The first 3D data generation model is the first in the data order among the plurality of 3D data generation models contained in the set.

[0484] For example, the bitstream contains licensing information indicating whether the 3D data generation model is permitted to reference other 3D data generation models in the encoding of each of the plurality of 3D data generation models.

[0485] For example, the first three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt0-2) corresponds to a first period (e.g., period t0~t2) that includes the first time point (e.g., time t0). The second three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt3-5) corresponds to a second period (e.g., period t3~t5) that includes the second time point (e.g., time t3).

[0486] For example, the multiple first learning images used in the generation of the first three-dimensional data generation model are two-dimensional images obtained by taking pictures at multiple different timings during the first period.

[0487] For example, when the first three-dimensional data generation model is input with a time period included in the first period, it outputs a two-dimensional image of the subject at the input time.

[0488] For example, the bitstream contains information indicating the upper limit of the number of images that the first 3D data generation model can generate.

[0489] For example, the bitstream contains first information associated with the plurality of first learning images. This first information includes multiple viewpoints and multiple line-of-sight directions corresponding to the plurality of first learning images, as well as multiple different timings.

[0490] For example, the first period or the second period is dynamically determined based on the subject.

[0491] For example, circuit 1471 saves the generated first three-dimensional data generation model in memory 1472. Based on the first three-dimensional data generation model saved in memory 1472, circuit 1471 generates the second three-dimensional data generation model.

[0492] For example, circuit 1471 stores the generated first 3D data generation model and the second 3D data generation model in memory 1472. Based on the first 3D data generation model and the second 3D data generation model stored in memory 1472, circuit 1471 generates an initial model. Based on the initial model, circuit 1471 generates a third 3D data generation model (e.g., 3D data generation model NNt2) corresponding to a third time point (e.g., time t2).

[0493] (other) In one embodiment, a method for generating motion images from a predetermined viewpoint is disclosed. The generation of motion images is achieved, for example, by a device including a memory and circuitry connected to the memory. In one example of this device, a three-dimensional data generation model (Neural Network) generated through learning is stored in the memory, and the circuitry retrieves the stored three-dimensional data generation model (Neural Network) and generates motion images based on the three-dimensional data generation model. Alternatively, the three-dimensional data generation model or an extended three-dimensional data generation model may not be stored in memory. For example, encoding devices 1420 and 1460 may retrieve specified information from a URL on a specified network and obtain the three-dimensional data generation model based on that specified information.

[0494] Figure 53 This is a diagram illustrating an example of the configuration of an encoding device.

[0495] The encoding device 1490 includes a processor 1491 and a memory 1492.

[0496] Processor 1491 is a circuit that performs information processing and is capable of accessing memory 1492. For example, processor 1491 is a dedicated or general-purpose electronic circuit that encodes a 3D data generation model. Processor 1491 can also be a processor like a CPU. Alternatively, processor 1491 can be an assembly of multiple electronic circuits. Furthermore, for example, processor 1491 can also function as multiple components of the aforementioned encoding device, excluding the component for storing information.

[0497] Memory 1492 is a dedicated or general-purpose memory that stores information used by processor 1491 to encode the 3D data generation model. Memory 1492 can be an electronic circuit or connected to processor 1491. Alternatively, memory 1492 can be contained within processor 1491. Alternatively, memory 1492 can be an assembly of multiple electronic circuits. Alternatively, memory 1492 can be a disk or optical disk, or it can be a storage device or recording medium. Alternatively, memory 1492 can be non-volatile memory or volatile memory.

[0498] For example, memory 1492 may store the encoded 3D data generation model, or it may store the stream corresponding to the encoded 3D data generation model. Additionally, memory 1492 may also store the program used by processor 1491 to encode the 3D data generation model.

[0499] Furthermore, in the encoding device 1490, it is possible to omit all of the aforementioned components of the encoding device, and it is also possible to omit all of the aforementioned processes. A portion of the components may be included in other devices, and a portion of the aforementioned processes may be performed by other devices.

[0500] Figure 54 This is a diagram illustrating an example of the configuration of a decoding device.

[0501] The decoding device 1495 includes a processor 1496 and a memory 1497.

[0502] Processor 1496 is a circuit that performs information processing and is capable of accessing memory 1497. For example, processor 1496 is a dedicated or general-purpose electronic circuit for decoding streams. Processor 1496 can also be a processor like a CPU. Alternatively, processor 1496 can be an assembly of multiple electronic circuits. Furthermore, for example, processor 1496 can also function as multiple components of the aforementioned decoding device, excluding the component for storing information.

[0503] Memory 1497 is a dedicated or general-purpose memory that stores information used by processor 1496 to decode the stream. Memory 1497 can be an electronic circuit or connected to processor 1496. Alternatively, memory 1497 can be contained within processor 1496. Alternatively, memory 1497 can be an assembly of multiple electronic circuits. Alternatively, memory 1497 can be a magnetic disk or optical disk, or it can be a storage device or recording medium. Alternatively, memory 1497 can be non-volatile memory or volatile memory.

[0504] For example, the memory 1497 can store a 3D data generation model or a stream. Furthermore, the memory 1497 can also store a program for the processor 1496 to decode the stream.

[0505] Furthermore, in the decoding device 1495, it is possible to omit all of the aforementioned components of the decoding device, and it is also possible to omit all of the aforementioned processes. A portion of the components may be included in other devices, and a portion of the aforementioned processes may be performed by other devices.

[0506] (Implementation Method 3) In Implementation 3, a method for encoding and transmitting a three-dimensional model (a learning model for generating three-dimensional data) is described.

[0507] For example, point group data, such as point clusters and meshes, includes line information connecting 3D points or points, surface information, attribute information corresponding to points, and attribute information corresponding to surfaces. Therefore, when the resolution of points or meshes increases, or the area of ​​points or meshes increases, the amount of data in point group data increases proportionally to the increase in resolution or area.

[0508] When the area of ​​points or grids is large, even if the constituent elements of point group data such as points or grids are encoded, the amount of encoded data will still be large due to the large amount of data itself.

[0509] In contrast, 3D models, as learning models used to generate 3D data, show minimal increase in data volume even as the area of ​​points or grids increases. A 3D model is a network model obtained by learning from 2D data (2D images) or 3D data (point groups or grids) using neural networks or similar techniques to learn the 3D shape and corresponding attribute information. Furthermore, since a 3D model is a network model used to generate 3D data, it can also be called a 3D generative model.

[0510] Therefore, in order to reduce the storage capacity or transmission volume of 3D data, a method for encoding and transmitting 3D models is needed.

[0511] In this embodiment, the three-dimensional model learning unit (three-dimensional model acquisition unit), three-dimensional model encoding unit (network model encoding unit), and three-dimensional model decoding unit (network model decoding unit) described in the previous embodiments will be specifically explained as a method for modeling a three-dimensional model using NeRF (Neural Radiance Fields), as well as the encoding method of the three-dimensional model, the decoding method of the encoded three-dimensional model, and the method for decoding two-dimensional images or three-dimensional data from the three-dimensional model.

[0512] A 3D model generated using basic NeRF can also consist of multiple networks. Furthermore, the networks referred to here are learned models obtained through learning using neural networks. Multiple networks can, for example, include networks learned using sparse sampling points and networks learned using dense sampling points. Thus, multiple networks are networks with different numbers of input sampling points or different sampling point densities. Sampling points can, for example, be 3D points representing 3D locations.

[0513] Alternatively, the multiple networks may include, for example, a network for outputting geometric information such as the object's density, probability of existence, and geometric coordinates, and a network for outputting information (attribute information) associated with geometry, such as color information, reflectivity, normal vector, color coordinates, timestamp, and object ID, based on the geometric information. Multiple networks may include two or more networks. Multiple networks may be three or more networks with different sampling points, or may have two or more networks for outputting geometric information, or two or more networks for outputting attribute information.

[0514] Multiple networks can be encoded using multiple network coding units. Multiple network coding units can encode multiple networks, for example, using existing network coding units such as NNC (Neural Network Coding) in the MPEG standard.

[0515] Figure 55 This is a block diagram illustrating an example of the structure of the encoding device in Embodiment 3.

[0516] The encoding device 1500 includes a NeRF 3D generative model learning unit 1501 and a NeRF generative model encoding unit 1502.

[0517] The NeRF 3D generative model learning unit 1501 generates NeRF model data by acquiring and learning from multiple 2D images (or 3D data, etc.), and then outputs the generated NeRF model data. NeRF model data is an example of a 3D model.

[0518] The NeRF generative model encoding unit 1502 encodes the NeRF model data output from the NeRF 3D generative model learning unit 1501 and outputs the encoded NeRF model data, i.e., the NeRF model encoded data. In the NeRF model encoded data, in addition to the encoded data obtained by encoding the constituent elements of the NeRF model, metadata related to the encoding (additional information, control information, encoding parameters, etc.) is also encoded and output in a prescribed bitstream format. The constituent elements of the NeRF model include, for example, multiple layers constituting a neural network learning model, multiple nodes contained in each of the multiple layers, and edges connecting the nodes between layers. The NeRF model encoded data can also be used in conjunction with the data from Embodiment 1. Figure 1 The encoded data shown is processed in the same way. For example, NeRF model encoded data can be multiplexed into a specified format and stored in a file, or the file can be transmitted. Furthermore, the multiplexed data can also be transmitted using specified communication methods.

[0519] Figure 56 This is a block diagram illustrating an example of the structure of the decoding device in Embodiment 3.

[0520] The decoding device 1510 includes a NeRF generation model decoding unit 1511 and a rendering data reconstruction unit 1512.

[0521] The NeRF generation model decoding unit 1511 acquires the received or reproduced and demultiplexed NeRF model encoded data, decodes the metadata related to the encoding from the acquired NeRF model encoded data, and uses the metadata to decode the NeRF model data.

[0522] The rendering data reconstruction unit 1512 reproduces and outputs 3D data from the NeRF model data. Alternatively, the rendering data reconstruction unit 1512, based on the NeRF model data and the viewpoint position and view vector information, outputs a 2D image viewed from the direction of the viewpoint indicated by the view vector information. That is, the rendering data reconstruction unit 1512 can also generate and output a 2D image from any viewpoint specified by the user, for example, based on the NeRF model data.

[0523] Next, taking NeRF, one of the methods for 3D modeling, as an example, we will use... Figure 57 This paper describes a method for generating a 3D model from multiple 2D images and encoding the generated network model. Furthermore, the method described here is an example and is not limited to the NeRF approach described here; it can also be applied to other NeRF approaches or 3D modeling methods.

[0524] Figure 57 This is a block diagram illustrating an example of the structure of an encoding apparatus for encoding multiple networks in Embodiment 3.

[0525] The encoding device 1520 includes a 3D generative model learning unit 1521 and a 3D generative model encoding unit 1525. The encoding device 1520 may also include a bitstream data constructing unit 1529.

[0526] First, the specific structure of the 3D generative model learning unit 1521 will be described. Specifically, the 3D generative model learning unit 1521 has a first network learning unit 1522, a sampling point determination unit 1523, and a second network learning unit 1524.

[0527] The first network learning unit 1522 learns a 3D model using multiple input 2D images and the viewpoint information (camera pose) of each of the input 2D images. That is, the first network learning unit 1522 learns 2D images corresponding to each viewpoint information based on multiple 2D images and the viewpoint information of each 2D image, thereby generating a 3D model (first network). Furthermore, the viewpoint information includes the viewpoint at which the 2D image was captured and the line-of-sight vector (line-of-sight direction) from that viewpoint. The first network learning unit 1522 can also be input with sampling points (first sampling points) during learning. First sampling points can be, for example, a set of points with large (coarse) intervals between them. The coordinates of each point included in the first sampling points can be predetermined coordinates or coordinates calculated using a prescribed method. The first network learning unit 1522 outputs the learned 3D model (first network). The 3D model (first network) generated by the first network learning unit 1522 is a network that outputs density information for the sampling points. Additionally, the first network learning unit 1522 can also output density information for the sampling points obtained during learning.

[0528] Here, density information represents the density of the object at the sampling point. For example, density information is set higher (i.e., higher than a predetermined value) when the object is a person or a table, lower (i.e., lower than a predetermined value) when the object is a light-transmitting object like glass, and set to a value close to 0 when there is no object. Therefore, density information can also be called information indicating the presence or absence of an object, or information indicating the probability of an object's presence. Additionally, density information can also be called geometric information.

[0529] The sampling point determination unit 1523 determines the density of an object at the coordinates shown by the coarsely sampled first sampling point based on the density information for the first sampling point output from the first network learning unit 1522, and determines a second sampling point for learning in the second network learning unit 1524. Alternatively, in determining the second sampling point, the sampling point determination unit 1523 may, for example, determine the presence of an object if the density of the sampling point is greater than a predetermined density, and decide to sample more meticulously the surrounding area of ​​the determined object (the space where the object is determined to exist and the space around it). Furthermore, the sampling point determination unit 1523 may, for example, determine that there is no object in the space where the sampling point is located if the density of the sampling point is less than a predetermined density, and decide to sample more coarsely in that space, or decide not to sample. The sampling point determination unit 1523 may, for example, use a PDF sampler.

[0530] The meaning of the sampling points output by the sampling point determination unit 1523 varies depending on the density determination method. For example, sampling points output when an object is determined to exist can also be referred to as geometric information representing the coordinates of the object. Furthermore, based on the density of the sampling points, it is possible to distinguish between objects with high transmittance, such as glass, and hard materials, and to perform processing to remove sampling points from the extracted object's sampling points that are determined to be objects with high transmittance, hard materials, etc., within the object (space). In this way, sampling points that become the objects to be extracted can be determined based on the density of the sampling points. Therefore, by extracting sampling points of objects or materials with specific densities that satisfy specific conditions, the extracted sampling points can be determined as geometric information.

[0531] The sampling point determination unit 1523 can use a predetermined method or parameters, or a method or parameters selected from multiple methods or parameters, in determining the sampling points. In this case, information representing the predetermined method or parameters can be encoded as learning metadata and stored in the bitstream. Therefore, information representing the predetermined method or parameters can be notified to the decoding device as learning metadata included in the bitstream.

[0532] The structure of the second network learning unit 1524 is the same as that of the first network learning unit 1522. The second network learning unit 1524 uses multiple input two-dimensional images, viewpoint information (camera pose) of each of the input two-dimensional images, and second sampling points output from the sampling point determination unit 1523 to learn a three-dimensional model. That is, the second network learning unit 1524 learns two-dimensional images corresponding to each viewpoint information based on multiple two-dimensional images, viewpoint information of each of the two-dimensional images, and second sampling points, thereby generating a three-dimensional model (the second network).

[0533] Here, the second network generated by the second network learning unit 1524 is a network capable of outputting color information and density information. Furthermore, the second network learning unit 1524 can also output the density information and color information obtained during learning for the sampled points, and use this density information and color information for other processing.

[0534] Next, the specific structure of the 3D generation model encoding unit 1525 will be explained.

[0535] The 3D generative model encoding unit 1525 includes a first network encoding unit 1526, a second network encoding unit 1527, and a metadata encoding unit 1528.

[0536] The first network encoding unit 1526 encodes the first network that has been learned by the first network learning unit 1522. The first network encoding unit 1526 outputs the encoded data obtained by encoding the first network.

[0537] The second network encoding unit 1527 encodes the second network that has been learned by the second network learning unit 1524. The second network encoding unit 1527 outputs the encoded data obtained by encoding the second network.

[0538] The metadata encoding unit 1528 encodes the metadata generated by the sampling point determination unit 1523. The metadata encoding unit 1528 outputs the encoded data obtained by encoding the metadata.

[0539] Thus, in the 3D generative model encoding unit 1525, encoded data obtained by encoding the first network, the second network, and metadata respectively is generated and the generated encoded data is output.

[0540] Alternatively, the 3D generative model encoding unit 1525 can also use existing network encoding units such as NNC (Neural Network Coding) in the MPEG standard for encoding. The learned first and second networks include: multiple layers such as the network's input layer, intermediate layers, and output layer; nodes in each layer; weight coefficients for each node; and transformation functions for each node. The learned first and second networks can also have networks for outputting density information (density network), color networks for outputting color information, and reflectivity networks for outputting reflectivity information, respectively. Alternatively, the learned first and second networks can also have attribute networks for outputting attribute information such as color information or reflectivity information, respectively.

[0541] Figure 58 This is a diagram showing an example of the encoded data of the first network after learning in Implementation Method 3.

[0542] If the first network after training is a network used to generate sampling points, it may at least include a density network for outputting density information. The first network after training may also include a color network for outputting color information, or an attribute network (reflectance network) for outputting other attribute information (such as reflectance information).

[0543] Figure 59 This is a diagram showing an example of the encoded data of the second network after learning in Implementation Method 3.

[0544] The learned second network can be a network used to output color information or other attribute information for the sampled points. The learned second network can include a density network for outputting density information and a color network for outputting color information. Additionally, the learned second network can also include an attribute network (reflectance network) for outputting other attribute information (such as reflectance information).

[0545] Furthermore, if no attribute information needs to be output, the first or second network after learning may not include a color network for outputting color information, or an attribute network (reflectance network) for outputting other attribute information (such as reflectance information).

[0546] Next, the detailed structure of the first network learning department will be explained. Figure 60 This diagram illustrates an example of the detailed structure of the first network learning unit in Embodiment 3. Furthermore, the structure (learning method) of the first network learning unit 1522 described here is an example of a means of learning a network; other means may also be used.

[0547] The first network learning unit 1522 includes a first network unit 1522a, a comparison unit 1522b, and a rendering unit 1522c.

[0548] The first network unit 1522a estimates density and color information for the multiple sampling points based on the two-dimensional image of each input viewpoint information, the viewpoint information (camera pose) of each input two-dimensional image, and multiple sampling points corresponding to each input viewpoint information, and outputs the estimated density and color information. Specifically, the first network unit 1522a inputs the two-dimensional image of each viewpoint information, the viewpoint information of each two-dimensional image, and multiple sampling points corresponding to each viewpoint information to the first network before learning, and outputs the density and color information output from the first network.

[0549] The rendering unit 1522c performs rendering processing based on the density and color information of each sampling point output by the first network unit 1522a, generates a two-dimensional image of each viewpoint, and outputs the generated two-dimensional image of each viewpoint. The rendering unit 1522c can use a volume renderer.

[0550] The comparison unit 1522b compares the two-dimensional image output by the rendering unit 1522c with the two-dimensional image for learning input to the first network unit 1522a according to the viewpoint information, and feeds back the difference obtained from the comparison result to the first network unit 1522a.

[0551] The first network unit 1522a adjusts the parameters of the first network to reduce the feedback difference.

[0552] The first network learning unit 1522 learns and optimizes the parameters of the first network by repeatedly performing the above processing on the information from each viewpoint, and outputs the learned first network. In addition, the first network unit 1522 outputs the density information of each of the multiple sampling points.

[0553] The first network learning unit 1522 can also output the color information of each of the multiple sampling points obtained during learning. The color information can also be used for other processing.

[0554] Next, the learning process of the network will be explained. Figure 61 This is a flowchart illustrating an example of the learning process of the network in Implementation Method 3.

[0555] The first network learning unit 1522 inputs a two-dimensional image of each viewpoint information and the viewpoint information of each two-dimensional image into the first network before learning (S1501).

[0556] The first network learning unit 1522 learns the first network as a three-dimensional model according to the information of each viewpoint (S1502).

[0557] The first network learning unit 1522 performs the processing of steps S1511 to S1514 in step S1502.

[0558] The first network learning unit 1522 inputs multiple sampling points along the viewing direction into the first network of the learning object according to each pixel of the two-dimensional image to obtain three-dimensional information (density information and color information) (S1511).

[0559] The first network learning unit 1522 generates the colors of the pixels of the two-dimensional image based on the three-dimensional information, and generates the two-dimensional image (S1512). That is, the first network learning unit 1522 renders the two-dimensional image.

[0560] The first network learning unit 1522 extracts the difference between the input two-dimensional image and the rendered two-dimensional image (S1513).

[0561] The first network learning unit 1522 learns the first network based on the difference (S1514). Specifically, the first network learning unit 1522 adjusts the parameters of the first network to make the difference smaller.

[0562] The first network learning unit 1522 determines whether the learning has ended for all viewpoint information (S1503).

[0563] If the first network learning unit 1522 determines that learning has ended for all viewpoint information (S1503: Yes), it proceeds to step S1504. If it determines that learning has not ended for all viewpoint information (S1503: No), it executes step S1502 based on the next viewpoint information.

[0564] The first network learning unit 1522 outputs the completed network (first network) (S1504).

[0565] The first network learning unit 1522 can also output the three-dimensional information obtained during the learning process (S1505). The three-dimensional information may be, for example, the density information or color information of each point of the sampling points.

[0566] Furthermore, the learning process described above for the first network has also been applied to the learning process for the second network.

[0567] Next, the decoding device 1530 for decoding multiple networks will be described. Figure 62 This is a block diagram illustrating an example of the structure of a decoding apparatus for decoding multiple networks in Embodiment 3.

[0568] The decoding device 1530 includes a bitstream data segmentation unit 1531, a three-dimensional generation model decoding unit 1532, and a reconstruction unit 1536.

[0569] The bitstream data segmentation unit 1531 segments the input bitstream into encoded data of a first network, a second network, and metadata.

[0570] Next, the specific structure of the 3D generative model decoding unit 1532 will be described. The 3D generative model decoding unit 1532 includes a first network decoding unit 1533, a second network decoding unit 1534, and a metadata decoding unit 1535.

[0571] The first network decoding unit 1533 decodes the learned first network based on the encoded data of the first network. The first network decoding unit 1533 outputs the decoded learned first network.

[0572] The second network decoding unit 1534 decodes the learned second network based on the encoded data of the second network. The second network decoding unit 1534 outputs the decoded learned second network.

[0573] The metadata decoding unit 1535 decodes the metadata based on the encoded metadata. The metadata decoding unit 1535 outputs the decoded metadata.

[0574] Next, the specific structure of the reconstruction unit 1536 will be described. The reconstruction unit 1536 includes a density estimation unit 1537, a sampling point determination unit 1538, an attribute information estimation unit 1539, and a rendering unit 1540.

[0575] The density estimation unit 1537 uses the learned first network and the first sampling point to estimate the density information for the first sampling point and outputs the estimated density information.

[0576] The sampling point determination unit 1538 determines a second sampling point based on density information. The sampling point determination unit 1538 determines the second sampling point using parameters contained in the metadata, in the same way as the encoding device 1520. The sampling point determination unit 1538 outputs the determined second sampling point.

[0577] The attribute information estimation unit 1539 uses the learned second network and the second sampling point to estimate the density information and color information corresponding to the second sampling point. The attribute information estimation unit 1539 outputs the estimated density information and color information. The attribute information estimation unit 1539 can also estimate the density information and color information corresponding to the sampling point of any input viewpoint, and output the estimated density information and color information, even when sampling points of any input viewpoint are input.

[0578] The rendering unit 1540 performs rendering processing based on the density and color information of each of the second sampling points, generates a two-dimensional image of each viewpoint, and outputs the generated two-dimensional image of each viewpoint.

[0579] In addition, the reconstruction unit 1536 can also directly output the second sampling point and the attribute information (density information and color information) corresponding to the second sampling point estimated by the attribute information estimation unit 1539.

[0580] Next, we will explain the data structure for storing the encoded data, i.e., the encoded data.

[0581] According to this structure, the decoding device 1530 can identify each category of data from the NeRF encoded bitstream, thus enabling data segmentation and decoding by each category. Furthermore, the decoding device 1530 facilitates data-by-data operation (i.e., processing), enabling functions such as parallel decoding, random access, partial decoding, and scalable decoding.

[0582] Furthermore, in a 3D model composed of multiple networks, by assigning the same identification ID to the data that constitute the same 3D model, the decoder can identify the networks of the same 3D model.

[0583] Figure 63 This is a diagram illustrating an example of the syntax of metadata for sequence units in Implementation 3. Figure 64 This is a diagram of an example of the syntax for representing metadata of frame units.

[0584] A network learned via NeRF includes multiple layers such as the input layer, intermediate layers, and output layer; nodes in each layer; weights for each node; and transformation functions for each node. Parameters common to a sequence are stored in metadata common to that sequence, such as the Sequence Parameter Set (SPS). Furthermore, parameters common to each frame, each access unit, or multiple frame units are stored in metadata common to the frame or to multiple frame units.

[0585] For example, when structural information representing the network structure, such as information about multiple layers of the network, is fixed in the sequence, it can be stored in the SPS or in the frame unit metadata. The network structural information can also be stored in both the SPS and the frame unit metadata; in this case, the frame unit metadata can be used preferentially. Furthermore, the network structural information can be stored in either the SPS or the frame unit metadata. In this case, a flag indicating which of the SPS and metadata the network structural information is stored in is stored in the higher-order SPS bits, and the decoding device 1530 can determine which of the SPS and metadata the network structural information is stored in based on this flag.

[0586] Next, the data structure of the network's coding layer will be explained.

[0587] Figure 65 This is a diagram illustrating an example of the syntax of data units in a high-density network in Implementation 3. Figure 66 This diagram illustrates an example of the syntax of data units in a low-density network in Implementation Method 3. Furthermore, high density indicates that the interval between sampling points is larger than a predetermined value, resulting in coarse sampling. Low density indicates that the interval between sampling points is smaller than a predetermined value, resulting in fine sampling.

[0588] Figure 67 This is a diagram illustrating an example of the structure of the data unit of the first network in Implementation 3. Figure 68 This diagram illustrates an example of the structure of the data units in the second network of Embodiment 3. Furthermore, the first network is a high-density network, while the second network is a low-density network.

[0589] The network parameters, such as node weights, obtained through learning, can be stored as network data units within the network. Network data can include network structural information. This structural information includes, for example, information used to determine multiple layers, including the network's input, intermediate, and output layers, the nodes within each layer, the weights for each node, and the transformation functions for each node.

[0590] In addition, data related to the output of density information (density network for outputting density information), data related to the output of color information (color network for outputting color information), and data related to other attribute information (attribute network for outputting other attribute information (e.g., reflectance information) in the network data can also be grouped and saved.

[0591] For example, in the case where the first network, after learning, has a density network for outputting the density of the sampling points, the network data and the density network are encoded and stored in the payload of the network's data unit.

[0592] Additionally, for example, if the learned second network has a density network for outputting the density of the sampling points and a color network for outputting color information, the network data, density network, and color network are encoded and stored in the payload of the network's data unit.

[0593] In network encoding, existing network encoding units such as NNC (Neural Network Coding) from the MPEG standard can also be used. In this case, NNC data units can be used as network data units. Encoded data defined by NNC can be stored in data structures defined by NNC. Furthermore, network data units can also be divided into multiple data formats according to the network structure.

[0594] Next, the data structure of the encoded data of the NeRF 3D model will be explained. Figure 69 This is a diagram illustrating an example of the syntax of the encoded data of the three-dimensional model of NeRF in Implementation 3. Figure 70 This is a diagram illustrating an example of the cell type of NeRF in Implementation 3. Figure 71This is a diagram illustrating an example of the data structure of the encoded data of the three-dimensional model of NeRF in Implementation 3.

[0595] The encoded data of the NeRF 3D model is stored in, for example, the model_codec_unit within the codec_unit() function, which comprehensively handles the encoding method, as described in the data structure of the encoded data in the above-described embodiment, and then transmitted. Alternatively, the encoded data of the NeRF 3D model may be transmitted directly as is, without being stored in the codec_unit. The encoded data of the NeRF 3D model is referred to herein as the NeRF 3D model Unit.

[0596] The encoded data of a NeRF 3D model can be structured with a header and NeRF 3D model data, capable of storing Fine NW Data Units, Coarse NW Data Units, NeRF Metadata, SPS, FPS, NPS, SEI, etc. The nerf_unit_type in the header indicates the category of data stored in the NeRF 3D model data. Therefore, the encoding device can generate data that the decoding device 1530 can recognize as the constituent elements of the encoded data of the NeRF model.

[0597] Here, when the data is a Fine NW Data Unit or a Coarse NW Data Unit, a 3D sub-model ID that identifies the 3D model can also be assigned. For example, when the data are Fine NW Data Units or Coarse NW Data Units in the same 3D model, an ID indicating that they are data units in the same 3D model can also be assigned.

[0598] Furthermore, the encoded data of the NeRF 3D model may also include a 3D model frame ID representing the frame number of the 3D model. Additionally, data units of frames within the same time frame may be assigned the same frame ID in the encoded data of the NeRF 3D model. Furthermore, when the 3D model data is region-segmented data, the encoded data of the NeRF 3D model may also include a space ID indicating which region it belongs to. That is, data within the same region are assigned the same space ID. Moreover, this space ID is the same as the space ID described in the above embodiments.

[0599] Furthermore, the 3D model sub-model ID can be the same as the data unit ID described in the above embodiments. These identifiers allow related data to be mapped to each other, enabling the decoding device 1530 to identify the corresponding data.

[0600] Next, the structural information of the three-dimensional model of NeRF (first network or second network) will be explained. Figure 72 This is a diagram illustrating an example of the SPS syntax of the three-dimensional model of NeRF in Implementation 3. Figure 73 This is a diagram illustrating an example of the syntax of the structural information of the three-dimensional model of NeRF in Implementation 3. Figure 74 This is a diagram representing an example of component_type in implementation method 3. Figure 75 This is a diagram illustrating an example of the component coding type in implementation method 3.

[0601] The SPS (Sequence Parameter Set) stores the structural information of the NeRF that constitutes the sequence corresponding to the SPS. Therefore, the decoding device 1530 can obtain information about the constituent elements or components of the bitstream storing the SPS, and can begin decoding based on this information.

[0602] SPS includes: `number_of_component`, representing the number of components constituting the bitstream; and `component_type`, which serves as an identifier for each component based on its number. For example, `component_type` indicates whether the encoded data is geometric information or the density of geometric information. Furthermore, `component_type` can also represent... Figure 74 The types shown.

[0603] In addition, SPS includes a component coding type indicating which encoding method the component was encoded with. The encoding method indicated by the component coding type is, for example, MPEG G-PCC for point group compression, VVC for video codecs, NNC for network compression, etc. Furthermore, the component coding type can also represent... Figure 75 The encoding method is illustrated.

[0604] In this embodiment, in the example of a bitstream consisting of two networks, the number of components is 2, as shown below.

[0605] component0:component_type=1 or 5, component coding type=4 component1:component_type=2 or 6, component coding type=4 In addition, the above indicates whether it is the case where component_type=1 for component0 and component_type=2 for component1, or the case where component_type=5 for component0 and component_type=6 for component1.

[0606] Thus, the structural information of the three-dimensional model of NeRF is communicated to the decoding device 1530, which can then begin decoding based on the structural information of the three-dimensional model of NeRF.

[0607] in addition, Figure 75 The component_type shown in 0~6 is one example; it may not represent all components or only a portion of them.

[0608] Next, the reference relationships of the encoded data of the NeRF 3D model will be explained. Figure 76 This is a diagram illustrating the reference relationships of the encoded data of the three-dimensional model of NeRF in Implementation 3.

[0609] For example, to indicate that they are networks corresponding to the same frame 0, the first network (CoarseNW) and the second network (FineNW) in the 3D generative model of frame 0 are assigned the same 3D model frame id. Furthermore, to indicate that they are networks corresponding to the same 3D generative model, the first network (CoarseNW) and the second network (FineNW) are assigned the same model identifier, 3D model sub-model id.

[0610] Thus, the decoding device 1530 can identify the first network (CoarseNW) and the second network (FineNW) used during the generation of the 3D model. Furthermore, the first network (CoarseNW) and the second network (FineNW) are each assigned a ref_nerf_metadata_id, and the decoding device 1530 can perform decoding using the referenced NeRF metadata by referring to NeRFmetadata with the same nerf_metadata_id during decoding.

[0611] Next, the bit stream segmentation process in the decoding device 1530 will be explained.

[0612] The bitstream input to the decoding device 1530 contains data from various NeRF 3D model units. The decoding device 1530 first parses the header of the NeRF 3D model unit. If the Nerf_unit_type is FineNW Data Unit, the decoding device 1530 identifies the subsequent data as encoded data from the second network (FineNW) and decodes the data from the second network (FineNW). Similarly, the decoding device 1530 uses nerf_unit_type to determine which of the following is the data: coarse NWData Unit, NeRF Metadata, Sequence Parameter Set, Frame Parameter Set, or Network Parameter Set, and decodes the data represented by that type accordingly.

[0613] Next, the decoding processes of the first and second networks will be explained.

[0614] The decoding device 1530 parses the 3D model frame ID, 3D model sub-model ID, 3D modelspace ID, and ref_nerf_metadata_id to determine which frame, model, and space the encoded data from the first or second network belongs to, and pairs encoded data with the same identifier with each other. The decoding device 1530 can decode when all encoded data with the same identifier are available.

[0615] For example, when the decoding device 1530 receives a fine NW Data Unit with 3D model frame id=0 and 3D model sub modelid=1, it searches for a Coarse NW Data Unit with the same ID. When a Coarse NW Data Unit is found, it determines that it can be decoded and begins decoding.

[0616] Next, the constraints on the order of data arrangement will be explained.

[0617] If the second network (Fine Network) must reference the first network (Coarse Network), the encoding device 1520 can also be configured to transmit the encoded data of the first network (Coarse Network) before the encoded data of the second network (Fine Network). The first network (Coarse Network) is received and decoded by the decoding device 1530 before the second network (Fine Network). The decoding device 1530 receives the second network (Fine Network) after the first network (Coarse Network) and performs decoding, searching for the corresponding 3D model frame ID and 3D model sub-model ID.

[0618] Furthermore, if the received data is network encoded data, the decoding device 1530 parses the header of the network data unit and decodes the network using a prescribed method. For example, if it is an NNC data unit, the decoding device 1530 decodes it using a method specified by the NNC standard.

[0619] Next, an example of data segmentation will be described based on the process of segmenting three-dimensional data into more than one three-dimensional data as described in the above embodiments.

[0620] Figure 77 This is a diagram illustrating an example in Implementation 3 where the data of a frame is divided into three three-dimensional spaces. The diagram also shows an example of further dividing the data in one of the three three-dimensional spaces into two.

[0621] Figure 78 This is a diagram illustrating an example of an ID assigned to the segmented data in Implementation Method 3.

[0622] The figure shows an example of assigning a unique 3D model sub-model ID to each segmented data point. The 3D model sub-model ID is an ID used to identify the segmented data.

[0623] For example, a 3D model sub-model ID can be set to a unique value within the same subspace data. For instance, if four 3D data sets (sub_models) have the same space_id, the IDs can be assigned to the four 3D data sets (sub_models) in ascending order, starting from 0. Figure 78 In the middle, according to the order of the model in the upper left, the model in the upper center, the model in the lower center, and the model in the upper right, the 3D model sub-model id and 3D model space id are assigned as follows.

[0624] [3D model sub model id, 3D model space id]=[0, 0], [0, 1], [1, 1], [0, 2] Additionally, for example, in Figure 77 In the example, since it is divided into 4 three-dimensional data (sub_model), for each of the 4 three-dimensional data, id = 0, 1, 2, 3 can be assigned sequentially to the 4 three-dimensional data (sub_model) regardless of the subspace of the 3-dimensional data to which it belongs (i.e., space_id).

[0625] Additionally, for the four 3D data sets (sub_model), the same spaceid can be assigned to data within the same subspace.

[0626] Thus, the decoding device 1530 is able to identify the data units of the first network (coarse NW data unit) and the data units of the second network (fine NW data unit) after segmentation.

[0627] In addition, an example of assigning 3D model sub-model id and 3D model space id to both the data unit of the first network (coarse NW data unit) and the data unit of the second network (fine NW data unit) is given. However, when the two data units are mapped using 3D model sub-model id, it is also possible not to assign 3D model space id to either data unit.

[0628] Next, examples of the encoding method and the output of the corresponding decoding device will be explained.

[0629] Figure 79 This is a diagram used to illustrate the first example of the encoding method in Implementation Method 3. Figure 80 This is a diagram illustrating the first example of the output of the decoding device in Embodiment 3.

[0630] As shown in the figure, in the encoded data, the first network includes a density network for the first sampling point but does not include a color network for the first sampling point, and the second network includes a density network for the second sampling point and a color network for the second sampling point.

[0631] Therefore, the decoding device uses the first network to generate a second sampling point based on the density information for the first sampling point. Furthermore, the decoding device can use the second network to generate density and color information for the second sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and color information.

[0632] Figure 81 This is a diagram illustrating a second example of the encoding method in Implementation Method 3. Figure 82 This is a diagram illustrating a second example of the output of the decoding device in Embodiment 3.

[0633] As shown in the figure, in the encoded data, the first network includes a density network for the first sampling point and a color network for the first sampling point, and the second network includes a density network for the second sampling point and a color network for the second sampling point.

[0634] Therefore, the decoding device can use the first network to generate density and color information for the first sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and color information. Similarly, as in the first example, the decoding device can generate a second sampling point based on the density information generated using the first network, use the second network to generate density and color information for the second sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and color information.

[0635] Thus, the decoding device in the second example can output three-dimensional or two-dimensional data for the first sampling point, and can also output three-dimensional or two-dimensional data for the second sampling point.

[0636] Figure 83 This is a diagram illustrating a third example of the encoding method in Implementation Method 3. Figure 84 This is a diagram illustrating a third example of the output of the decoding device in Embodiment 3.

[0637] As shown in the figure, in the encoded data, the first network includes a density network for the first sampling point but does not include a color network for the first sampling point; the second network includes a density network and a color network for the second sampling point; and the third network includes a density network and a reflectance network for the second sampling point.

[0638] Therefore, the decoding device uses a first network to generate a second sampling point based on the density information for the first sampling point. Furthermore, the decoding device can use a second network to generate density and color information for the second sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and color information. Further, the decoding device can use a third network to generate density and reflectance information for the second sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and reflectance information. In this way, the decoding device can output three-dimensional data containing color information for the second sampling point and two-dimensional data containing color information, and can also output three-dimensional data containing reflectance information for the second sampling point and two-dimensional data containing reflectance information.

[0639] Figure 85 This is a diagram illustrating the fourth example of the encoding method in Implementation Method 3. Figure 86 This is a diagram illustrating the fourth example of the output of the decoding device in Embodiment 3.

[0640] The fourth example is an instance of generating a third sampling point that is denser than the second sampling point.

[0641] As shown in the figure, in the encoded data, the first network includes a density network for the first sampling point but does not include a color network for the first sampling point; the second network includes a density network for the second sampling point but does not include a color network for the second sampling point; and the third network includes both a density network and a color network for the third sampling point.

[0642] Therefore, the decoding device uses a first network to generate a second sampling point based on the density information for the first sampling point. Furthermore, the decoding device uses a second network to generate a third sampling point based on the density information for the second sampling point. The decoding device can also use a third network to generate density and color information for the third sampling point, and based on the density and color information, reproduce and output three-dimensional or two-dimensional data.

[0643] Figure 87 This is a diagram illustrating the fifth example of the encoding method in Implementation Method 3. Figure 88 This is a diagram illustrating the fifth example of the output of the decoding device in Embodiment 3.

[0644] The fifth example is similar to the fourth example, generating a third sampling point that is denser than the second sampling point.

[0645] As shown in the figure, in the encoded data, the first network includes a density network for the first sampling point and a color network for the first sampling point, the second network includes a density network for the second sampling point and a color network for the second sampling point, and the third network includes a density network for the third sampling point and a color network for the third sampling point.

[0646] Therefore, the decoding device can use a first network to generate density and color information for a first sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and color information. Furthermore, the decoding device can generate a second sampling point based on the density information generated using the first network, use a second network to generate density and color information for the second sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and color information. Additionally, the decoding device can generate a third sampling point based on the density information generated using the second network, use a third network to generate density and color information for the third sampling point, and reproduce and output three-dimensional or two-dimensional data based on the density and color information.

[0647] Thus, the decoding device in the fifth example can output three-dimensional or two-dimensional data for the first sampling point, and can output three-dimensional or two-dimensional data for the second sampling point, and can output three-dimensional or two-dimensional data for the third sampling point.

[0648] Next, the method for ensuring the consistency of the notification bitstream or the consistency of the decoding device will be explained.

[0649] Figure 89 This diagram illustrates the data exchange between the decoding unit and the control unit in Embodiment 3.

[0650] As shown in the figure, the decoding unit 1541 can store information indicating what kind of data is encoded in metadata such as SEI and send it to the control unit (application execution unit) 1542 (SEI information transmission).

[0651] The control unit 1542 can determine what kind of data is encoded in the bitstream or what kind of data the decoding device can output by decoding and parsing the SEI.

[0652] In addition, the control unit 1542 can specify which data to output to the decoding unit 1541 (output data specification), and the decoding unit 1541 can output the specific data specified by the control unit 1542 (specific data output).

[0653] Figure 90 This is a diagram used to illustrate the conformance point in implementation method 3. Figure 90 Decoding device 1530 and Figure 62The decoding device 1530 described herein is the same.

[0654] As shown in the figure, the data output by the decoding device 1530, or the block that the decoding device 1530 should process, can be defined in various ways.

[0655] For example, the first consistency point can be set to the data (first network, second network, and metadata) decoded by the decoding device 1530. The second consistency point can be set to the density information and color information corresponding to the second sampling point. The third consistency point can be set to the three-dimensional data reconstructed based on the first and second networks. The fourth consistency point can be set to the two-dimensional image rendered based on the first and second networks.

[0656] These coherence points can be information indicating which coherence point the decoding device 1530 is capable of outputting. Furthermore, a coherence point can be information indicating that the bitstream contains the data required for the decoding device 1530 to output the data set at that coherence point. That is, a coherence point can be information indicating whether the decoding device 1530 is capable of outputting the data represented by that coherence point, or it can be information indicating that the bitstream itself contains the data required to output the data represented by that coherence point.

[0657] For example, if the bitstream supports the output of a second consistency point, the consistency point in the bitstream, such as the SPS, indicates the second consistency point. In this case, additional information related to rendering can be omitted from signal transmission. Furthermore, the control unit 1542 can determine which data can be output by obtaining the consistency point from the bitstream.

[0658] Next, a variation of the fifth example of the encoding method and the output of the corresponding decoding device will be explained.

[0659] Figure 91 This is a diagram illustrating an example of a bitstream containing multiple networks in Implementation 3. Figure 92 This is a diagram illustrating an example of the syntax of the layered structure of multiple networks in Implementation 3.

[0660] As in the fifth example, when the encoding device 1520 encodes multiple networks, including density networks and color networks, it may also include an nw_level_id representing the network level in the header of each data unit. Therefore, the decoding device 1530 can identify the level from the header of the data unit.

[0661] Furthermore, the `nw_scalable_structure()` method, which represents the layer structure of multiple networks, can be included in the SPS or SEI. Thus, the decoding device 1530 can identify the layer structure before receiving data units and prepare for layered encoding.

[0662] A layer structure can contain a `layer_type`. `layer_type` can be an identifier representing the characteristics of a layer. For example, `layer_type` can indicate that the layer contains density, or it can indicate that it contains both density and color. `layer_type` can also be identification information used to identify the data contained in the layer corresponding to that `layer_type`.

[0663] Furthermore, the header of a data unit can contain data length information, density information data length, and color information data length.

[0664] Furthermore, metadata such as SPS can also contain flags representing the encoding of multiple networks, including density networks and color networks. These flags can then be used for signal transmission.

[0665] When the decoding device 1530 determines, based on metadata, that the encoded data contained in the bitstream is layered encoded data, it can select which data in the decoded data to output after receiving and decoding all the data. Furthermore, the decoding device 1530 can also, after determining which data to output, receive specific encoded data corresponding to the output data and decode the output data based on that encoded data.

[0666] For example, in the fifth example, when encoded data for the third sampling point is not required, the decoding device 1530 can decode the first and second networks but not the third network. For instance, the decoding device 1530 can determine which data units should not be decoded by referring to the nw_level_id contained in the header of the data unit. Thus, the encoding device 1520 unitizes the data according to each layer, generating a bitstream that assigns layer identifiers to the data units, thereby enabling the decoding device 1530 to perform partial decoding and layered decoding based on the bitstream. The decoding device 1530 can perform layered decoding, thereby enabling the decoding and output of three-dimensional or two-dimensional data with different resolutions.

[0667] Furthermore, by informing the decoding unit 1541 of the desired resolution, the control unit 1542 can decode only the encoded data used to decode the data at the desired resolution, thereby reducing the processing load. Depending on the use case, the control unit 1542 can decode and output low-resolution data when displaying thumbnails, and additionally decode and output high-resolution data when displaying the whole image. Therefore, it is possible to reduce the decoding processing load in the decoding device 1530.

[0668] Next, variations of the encoding and decoding devices will be described.

[0669] Figure 93 This is a block diagram illustrating an example of the structure of a modified version of the encoding device in Embodiment 3. Figure 94 This is a block diagram illustrating an example of the structure of a modified version of the decoding device in Embodiment 3.

[0670] The encoding device 1550 includes a three-dimensional generative model learning unit 1551, a three-dimensional generative model encoding unit 1553, and a bitstream data constructing unit 1529.

[0671] The difference between the three-dimensional generative model learning unit 1551 and the three-dimensional generative model learning unit 1521 is that it has a geometric generation unit 1552 instead of the first network learning unit 1522 and the sampling point determination unit 1523.

[0672] The geometry generation unit 1552 can generate a group of three-dimensional points (second sampling points) based on multiple two-dimensional images using photogrammetry technology. Alternatively, the geometry generation unit 1552 can generate second sampling points (a group of points) based on data separately obtained from a three-dimensional sensor (distance sensor), etc. The generated second sampling points are processed in the same way as the encoding device 1520.

[0673] The difference between the 3D generative model coding unit 1553 and the 3D generative model coding unit 1525 is that it has a point group coding unit 1554 instead of the first network coding unit 1526.

[0674] The dot group coding unit 1554 encodes the second sample points generated by the geometry generation unit 1552. The dot group coding unit 1554 may use, for example, MPEG G-PCC, MPEG V-PCC, Draco, etc., to encode the second sample points.

[0675] Furthermore, in the encoding device 1550, the constituent elements with the same markings as those in the encoding device 1520 have the same functions as the constituent elements of the encoding device 1520, so the description is omitted.

[0676] The decoding device 1560 includes a bitstream data segmentation unit 1561, a three-dimensional generation model decoding unit 1562, and a reconstruction unit 1564.

[0677] The bitstream data segmentation unit 1561 segments the input bitstream into second sampling points (point group), second network, and encoded data of metadata.

[0678] The difference between the 3D generative model decoding unit 1562 and the 3D generative model decoding unit 1532 is that it has a point group decoding unit 1563 instead of the first network decoding unit 1533.

[0679] The point group decoding unit 1563 decodes the second sample point (point group) based on the encoded data of the second sample point (point group). The point group decoding unit 1563 outputs the decoded second sample point (point group).

[0680] The difference between the reconstruction unit 1564 and the reconstruction unit 1536 is that the reconstruction unit 1564 does not have the density estimation unit 1537 and the sampling point determination unit 1538.

[0681] The attribute information estimation unit 1539 uses the learned second network output by the second network decoding unit 1534 and the second sampling point (point group) decoded by the point group decoding unit 1563 to estimate the density information and color information corresponding to the second sampling point.

[0682] Furthermore, in the decoding device 1560, the constituent elements with the same markings as those in the decoding device 1530 have the same functions as the constituent elements of the decoding device 1530, so the description is omitted.

[0683] (Structure Example 1) Figure 95 This is a diagram illustrating an example of the structure of the encoding device in Embodiment 3. Figure 96 This is a flowchart illustrating the first example of the encoding method of the encoding device in Embodiment 3.

[0684] The encoding device 1570 includes a circuit 1571 and a memory 1572 connected to the circuit 1571. The encoding device 1570 is an apparatus that implements the encoding devices 1500 and 1520.

[0685] Circuit 1571 performs the following actions.

[0686] Circuit 1571 acquires multiple networks constituting a three-dimensional data generation model (3D model) (S1521). Circuit 1571 encodes the multiple networks (S1522). Circuit 1571 outputs a bitstream containing the encoded multiple networks (S1523).

[0687] Therefore, the encoding device 1570 can output a bitstream containing encoded data obtained from multiple network encodings that will constitute a 3D generative model. This allows for a reduction in the storage capacity of 3D data or a reduction in the amount of 3D data transmitted.

[0688] For example, multiple networks include a first network and a second network. Furthermore, the output of the first network is used to generate the second network. Thus, the encoding device 1570 can use the output of the first network to generate the second network, thereby improving the accuracy of the 3D generated model while reducing the storage capacity or the amount of 3D data transmitted.

[0689] For example, the second network (FineNW) uses higher resolution sampling points for learning compared to the first network (CoarseNW). Therefore, the 3D generative model consists of the first network and the second network, which uses higher resolution sampling points for learning compared to the first network, thus improving the accuracy of the 3D generative model.

[0690] For example, the first network is used to output location information (point clusters). The second network is used to output attribute information (color information or reflectivity information) corresponding to the location information. Thus, the 3D generative model consists of the first network for outputting location information and the second network for outputting attribute information corresponding to the location information, thereby improving the accuracy of the 3D generative model.

[0691] For example, encoding multiple networks is done using NNC (Neural Network Coding) in the MPEG standard. This improves compatibility and allows decoding with a wide variety of decoding devices.

[0692] For example, circuit 1571 learns a 3D data generation model for each of the multiple viewpoints represented by the multiple viewpoint information by inputting multiple two-dimensional images, multiple viewpoint information corresponding to the multiple two-dimensional images, and a first sampling point into a first network, and outputs position information for the first sampling point. Furthermore, circuit 1571 learns a 3D data generation model for each of the multiple viewpoints by inputting a second sampling point determined based on the position information into a second network. Circuit 1571 also encodes the first and second networks obtained through learning. Therefore, learning can be performed using appropriate sampling points, thereby improving the accuracy of the 3D generation model.

[0693] For example, circuit 1571 also encodes metadata used for encoding. The metadata includes at least one of additional information, control information, and encoding parameters. This allows the metadata to be sent to the decoding device, reducing the load on the 3D generated model in the decoding device or improving the accuracy of the 3D generated model.

[0694] Figure 97 This is a diagram illustrating an example of the structure of the decoding device in Embodiment 3. Figure 98 This is a flowchart illustrating a first example of the decoding method of the decoding device in Embodiment 3.

[0695] The decoding device 1580 includes a circuit 1581 and a memory 1582 connected to the circuit 1581. The decoding device 1580 is a device that implements the decoding devices 1510 and 1530.

[0696] Circuit 1581 performs the following actions.

[0697] Circuit 1581 acquires a bitstream (S1531), which contains encoded multiple networks constituting a three-dimensional data generation model (three-dimensional model). Circuit 1581 decodes the multiple networks based on the bitstream (S1532).

[0698] Therefore, the decoding device 1580 is able to appropriately decode the bit stream that has achieved a reduction in transmission volume.

[0699] For example, multiple networks include a first network and a second network. The second network is generated using the output of the first network. Thus, the decoding device 1580 is able to decode the first network and the second network, thereby enabling it to decode the 3D generative model with improved accuracy.

[0700] For example, the second network (FineNW) uses higher resolution sampling points for learning compared to the first network (CoarseNW). Therefore, since the decoding device 1580 can decode both the first network and the second network (which uses higher resolution sampling points for learning compared to the first network), it can decode the 3D generative model with improved accuracy.

[0701] For example, the first network is used to output location information (point clusters). The second network is used to output attribute information (color information or reflectivity information) corresponding to the location information. Therefore, since the decoding device 1580 can decode both the first network used to output location information and the second network used to output attribute information corresponding to the location information, it can decode the 3D generated model with improved accuracy.

[0702] For example, decoding of multiple networks is performed using NNC (Neural Network Coding) in the MPEG standard. Thus, the decoding device 1580 is able to appropriately decode highly compatible bitstreams.

[0703] For example, the first network is a network for outputting position information for a first sampling point, comprising a 3D data generation model learned based on multiple two-dimensional images, multiple viewpoint information corresponding to the multiple two-dimensional images, and the first sampling point according to each of the multiple viewpoints represented by the multiple viewpoint information. The second network is a 3D data generation model learned based on a second sampling point determined according to the position information, according to each of the multiple viewpoints. Thus, the decoding device 1580 is able to decode the 3D generation model whose accuracy has been improved by learning using appropriate sampling points.

[0704] For example, the bitstream also contains metadata for the encoding. The metadata includes at least one of additional information, control information, and encoding parameters. Therefore, because the decoding device can obtain the metadata, it can reduce the overhead associated with the generation of the 3D generative model and can decode 3D generative models with improved accuracy.

[0705] (Structure Example 2) Encoding device 1570 can also implement encoding device 1550. Figure 99 This is a flowchart illustrating a second example of the encoding method of the encoding device in Embodiment 3.

[0706] Circuit 1571 can also perform the following operations.

[0707] Circuit 1571 generates a three-dimensional point group (S1541). Circuit 1571 inputs the generated three-dimensional point group (sampling points) into the network to generate a learned network (S1542). Circuit 1571 encodes the three-dimensional point group and the learned network (S1543). Circuit 1571 outputs a bitstream containing the encoded three-dimensional point group and the encoded network (S1544).

[0708] Therefore, the encoding device can, for example, output a bitstream containing encoded data obtained by encoding a sparse group of three-dimensional points and a network constituting a three-dimensional generative model. Thus, it is possible to reduce the storage capacity of three-dimensional data or reduce the amount of three-dimensional data transmitted.

[0709] Decoding device 1580 can also be implemented as decoding device 1560. Figure 100 This is a flowchart illustrating a second example of the decoding method of the decoding device in Embodiment 3.

[0710] Circuit 1581 obtains a bitstream containing the encoded 3D point group and the encoded network (S1551). Circuit 1581 decodes the 3D point group and the learned network based on the bitstream (S1552).

[0711] Therefore, the decoding device is able to appropriately decode bitstreams that have achieved reduced throughput.

[0712] (Structure Example 3) Figure 101 This is a flowchart illustrating a third example of the encoding method of the encoding device in Embodiment 3.

[0713] Circuit 1571 can also perform the following operations.

[0714] Circuit 1571 acquires multiple networks constituting a three-dimensional data generation model (three-dimensional model) (S1561). Circuit 1571 encodes the multiple networks (S1562). Circuit 1571 outputs a bit stream containing the encoded multiple networks (S1563). The bit stream contains: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames.

[0715] Therefore, the encoding device can output a bitstream containing encoded data obtained from multiple network encodings that will constitute a 3D generative model. This allows for a reduction in the storage capacity or the amount of 3D data transmitted.

[0716] For example, at least one of the first metadata and the second metadata includes a first parameter representing the number of layers in the first network of a plurality of networks and a second parameter representing the number of nodes in each layer.

[0717] For example, the second metadata may contain the same specific parameters as the first metadata. The specific parameters contained in the second metadata are used preferentially over the specific parameters contained in the first metadata.

[0718] For example, the first metadata contains a flag indicating whether the first parameter or the second parameter is contained in the first metadata or the second metadata.

[0719] For example, the header of each unit in the bitstream contains an identifier (3D model sub-model id), which indicates the frame number of the frame corresponding to the 3D data generation model. Therefore, a decoding device that has obtained the bitstream can identify whether it is a network targeting the same 3D model based on the identifier contained in the bitstream.

[0720] For example, a frame corresponding to a 3D data generation model is divided into multiple spaces. At least one of these spaces is further divided into multiple subspaces. Identifiers contain distinct values ​​corresponding to each of the subspaces.

[0721] For example, the header of each unit in the bitstream contains an identifier indicating the frame number of the frame corresponding to the 3D data generation model. Therefore, a decoding device that has obtained the bitstream can identify whether data units belong to the same frame based on the identifier contained in the bitstream.

[0722] For example, the header of each cell in the bitstream contains an identifier indicating which region in the multiple spaces obtained by segmenting the 3D data to generate the model. Therefore, a decoding device that has obtained the bitstream can identify whether the data belongs to the same region based on the identifiers contained in the bitstream.

[0723] For example, the first metadata contains parameters indicating the number of components that make up the bitstream and identifiers that identify the components.

[0724] Figure 102 This is a flowchart illustrating a third example of the decoding method of the decoding device in Embodiment 3.

[0725] Circuit 1581 can also perform the following operations.

[0726] Circuit 1581 acquires a bitstream containing multiple encoded networks constituting a three-dimensional data generation model (three-dimensional model) (S1571). The bitstream contains: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames. Circuit 1581 decodes the multiple networks based on the bitstream (S1572).

[0727] Therefore, the decoding device is able to appropriately decode bitstreams that have achieved reduced throughput.

[0728] For example, at least one of the first metadata and the second metadata includes a first parameter representing the number of layers in the first network of a plurality of networks and a second parameter representing the number of nodes in each layer.

[0729] For example, the second metadata may contain the same specific parameters as the first metadata. The specific parameters contained in the second metadata take precedence over the specific parameters contained in the first metadata.

[0730] For example, the first metadata contains a flag indicating which of the first or second parameters is included in the first and second metadata.

[0731] For example, the header of each unit in the bitstream contains an identifier (3D model sub-model id), which indicates the frame number of the frame corresponding to the 3D data generation model. Therefore, the decoding device that obtains the bitstream can identify whether it is a network targeting the same 3D model based on the identifier contained in the bitstream.

[0732] For example, a frame corresponding to a 3D data generation model is divided into multiple spaces. At least one of these spaces is further divided into multiple subspaces. Identifiers contain distinct values ​​corresponding to each of the subspaces.

[0733] For example, the header of each unit in the bitstream contains an identifier indicating the frame number of the frame corresponding to the 3D data generation model. Therefore, a decoding device that has obtained the bitstream can identify whether data units belong to the same frame based on the identifier contained in the bitstream.

[0734] For example, the header of each cell in the bitstream contains an identifier indicating which region in the multiple spaces obtained by segmenting the 3D data to generate the model. Therefore, a decoding device that has obtained the bitstream can identify whether the data belongs to the same region based on the identifiers contained in the bitstream.

[0735] For example, the first metadata contains parameters indicating the number of components that make up the bitstream and identifiers that identify the components.

[0736] (Structure Example 4) Figure 103 This is a flowchart illustrating a fourth example of the encoding method of the encoding device in Embodiment 3.

[0737] Circuit 1571 can also perform the following operations.

[0738] Circuit 1571 acquires multiple networks constituting a three-dimensional data generation model, including a first network and a second network (S1581). Circuit 1571 encodes the multiple networks (S1582). Circuit 1571 stores at least one of the following in the data units of the first network and the second network respectively: a density network for outputting density information and a color network for outputting color information, thereby outputting a bitstream containing the encoded multiple networks (S1583).

[0739] Therefore, the encoding device can output a bitstream containing encoded data obtained from multiple network encodings that will constitute a 3D generative model. This allows for a reduction in the storage capacity or the amount of 3D data transmitted.

[0740] For example, the data unit of the first network stores a density network for outputting density information for the first sampling point. The data unit of the second network stores a density network for outputting density information for the second sampling point, and a color network for outputting color information for the second sampling point.

[0741] Thus, the bitstream decoding device is able to reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network and color network for the second sample point.

[0742] For example, the data unit of the first network stores a density network for outputting density information for the first sampling point and a color network for outputting color information for the first sampling point. The data unit of the second network stores a density network for outputting density information for the second sampling point and a color network for outputting color information for the second sampling point.

[0743] Therefore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the first sampled data based on the density network and color network for the first sampled point. Furthermore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network and color network for the second sampled point.

[0744] For example, the multiple networks also include a third network. In the data unit of the first network, a density network for outputting density information for the first sampling point is stored. In the data unit of the second network, a density network for outputting density information for the second sampling point and a color network for outputting color information for the second sampling point are stored. In the data unit of the third network, a density network for outputting density information for the second sampling point and a reflectance network for outputting reflectance information for the second sampling point are stored.

[0745] Thus, the decoding device that obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network, color network, and reflectivity network for the second sample point.

[0746] For example, the multiple networks also include a third network. In the data unit of the first network, a density network for outputting density information for a first sampling point is stored. In the data unit of the second network, a density network for outputting density information for a second sampling point is stored. In the data unit of the third network, a density network for outputting density information for a third sampling point and a color network for outputting color information for the third sampling point are stored.

[0747] Thus, the decoding device that obtains the bitstream can, for example, further reproduce and output three-dimensional or two-dimensional data of the third sampled data based on the density network and color network for the dense third sampled points.

[0748] For example, the multiple networks also include a third network. In the data unit of the first network, a density network for outputting density information for a first sampling point and a color network for outputting color information for the first sampling point are stored. In the data unit of the second network, a density network for outputting density information for a second sampling point and a color network for outputting color information for the second sampling point are stored. In the data unit of the third network, a density network for outputting density information for a third sampling point and a color network for outputting color information for the third sampling point are stored.

[0749] Therefore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the first sampled data based on the density network and color network for the first sampled data. Furthermore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network and color network for the second sampled data. Additionally, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the third sampled data based on the density network and color network for the third sampled data.

[0750] For example, the bitstream also includes at least one of a first consistency point, a second consistency point, a third consistency point, and a fourth consistency point. The first consistency point indicates that the decoding device can perform processing on multiple networks decoded by the decoding device. The second consistency point indicates that the decoding device can perform processing on at least one of the density information and color information corresponding to the second sampling point. The third consistency point indicates that the decoding device can perform processing on three-dimensional data generated based on the first and second networks. The fourth consistency point indicates that the decoding device can perform processing on a two-dimensional image generated based on the first and second networks.

[0751] Therefore, by referring to the consistency points contained in the bitstream, the device that has obtained the bitstream can determine the data that the bitstream supports for output, or the data that the decoding device supports for output, and can determine the data that the device can output.

[0752] For example, the header of a data unit in the first network contains first-level information indicating the hierarchy of the first network. The header of a data unit in the second network contains second-level information indicating the hierarchy of the second network.

[0753] Figure 104 This is a flowchart illustrating a fourth example of the decoding method of the decoding device in Embodiment 3.

[0754] Circuit 1581 can also perform the following operations.

[0755] Circuit 1581 acquires a bitstream (S1591), which contains multiple encoded networks constituting a three-dimensional data generation model. The multiple networks include a first network and a second network. In the data units of the first network and the second network, at least one of a density network for outputting density information and a color network for outputting color information, which are learned networks, is stored respectively. Circuit 1581 decodes the multiple networks based on the bitstream (S1592).

[0756] Therefore, the decoding device is able to appropriately decode bitstreams that have achieved reduced throughput.

[0757] For example, the data unit of the first network stores a density network for outputting density information for the first sampling point. The data unit of the second network stores a density network for outputting density information for the second sampling point, and a color network for outputting color information for the second sampling point.

[0758] Thus, the bitstream decoding device is able to reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network and color network for the second sample point.

[0759] For example, the data unit of the first network stores a density network for outputting density information for the first sampling point and a color network for outputting color information for the first sampling point. The data unit of the second network stores a density network for outputting density information for the second sampling point and a color network for outputting color information for the second sampling point.

[0760] Therefore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the first sampled data based on the density network and color network for the first sampled point. Furthermore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network and color network for the second sampled point.

[0761] For example, the multiple networks also include a third network. In the data unit of the first network, a density network for outputting density information for the first sampling point is stored. In the data unit of the second network, a density network for outputting density information for the second sampling point and a color network for outputting color information for the second sampling point are stored. In the data unit of the third network, a density network for outputting density information for the second sampling point and a reflectance network for outputting reflectance information for the second sampling point are stored.

[0762] Thus, the decoding device that obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network, color network, and reflectivity network for the second sample point.

[0763] For example, the multiple networks also include a third network. In the data unit of the first network, a density network for outputting density information for a first sampling point is stored. In the data unit of the second network, a density network for outputting density information for a second sampling point is stored. In the data unit of the third network, a density network for outputting density information for a third sampling point and a color network for outputting color information for the third sampling point are stored.

[0764] Thus, the decoding device that obtains the bitstream can, for example, further reproduce and output three-dimensional or two-dimensional data of the third sampled data based on the density network and color network for the dense third sampled points.

[0765] For example, the multiple networks also include a third network. In the data unit of the first network, a density network for outputting density information for a first sampling point and a color network for outputting color information for the first sampling point are stored. In the data unit of the second network, a density network for outputting density information for a second sampling point and a color network for outputting color information for the second sampling point are stored. In the data unit of the third network, a density network for outputting density information for a third sampling point and a color network for outputting color information for the third sampling point are stored.

[0766] Therefore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the first sampled data based on the density network and color network for the first sampled data. Furthermore, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the second sampled data based on the density network and color network for the second sampled data. Additionally, the decoding device that has obtained the bitstream can reproduce and output three-dimensional or two-dimensional data of the third sampled data based on the density network and color network for the third sampled data.

[0767] For example, the bitstream also includes at least one of a first consistency point, a second consistency point, a third consistency point, and a fourth consistency point. The first consistency point indicates that the decoding device can perform processing on multiple networks decoded by the decoding device. The second consistency point indicates that the decoding device can perform processing on at least one of the density information and color information corresponding to the second sampling point. The third consistency point indicates that the decoding device can perform processing on three-dimensional data generated based on the first and second networks. The fourth consistency point indicates that the decoding device can perform processing on a two-dimensional image generated based on the first and second networks.

[0768] Therefore, by referring to the consistency points contained in the bitstream, the device that has obtained the bitstream can determine the data that the bitstream supports for output or the data that the decoding device supports for output, and can determine the data that the device can output.

[0769] For example, the header of a data unit in the first network contains first-level information indicating the hierarchy of the first network. The header of a data unit in the second network contains second-level information indicating the hierarchy of the second network.

[0770] Industrial applicability This disclosure can be applied to encoding devices capable of encoding multiple networks, or decoding devices capable of decoding multiple networks.

[0771] Explanation of reference numerals in the attached figures 1001 Three-Dimensional Data Encoding System 1002 3D Data Decoding System 1003 Sensor Terminal 1004 External Connection Part 1011 3D Data Generation System 1012 Reminder Department 1013 Coding Department 1014 Reuse Department 1015 Input / Output Section 1016 Control Department 1017 Sensor Information Acquisition Department 1018 3D Data Generation Department 1021 Sensor Information Acquisition Department 1022 Input / Output Section 1023 Demultiplexing Department 1024 Decoding Department 1025 Reminder Department 1026 User Interface 1027 Control Department 1031 3D Model Learning Department 1032 3D Model Coding Department 1033 3D Model Decoding Department 1034 Rendering and Reconstruction Department 1041 Data Segmentation Unit 1042 Coding Department 1051 Decoding Department 1052 Data Integration Section 1070 server 1071 Data Generation Department 1072-point group generation department 1073 Mesh Generation Department 1074 Model Generation Department 1075 Synchronization Unit 1076-point group coding department 1077 Mesh Coding Department 1078 Model Coding Department 1079 Reuse Department 1080 Data Extraction Department 1090 terminal 1091 Control Department 1092 Decoding Department 1093 Reminder Department 1101 point group sensor 1102 camera 1110 Data Generation Department 1111 dot group generation department 1112 Mesh Generation Department Model Generation Department 1113 1130 Decoding Device 1131 circuit 1132 memory 1140 encoding device 1141 circuit 1142 memory 1401 3D Data Generation Model 1402 Evaluation Function 1403 Three-Dimensional Data Generation Model 1420 encoding device 1421 Three-dimensional data generation model acquisition department 1422 Buffer Section 1423 Network Model Coding Department 1425 Decoding Device 1426 Network Model Decoding Department Rendering Department 1427 1430 Encoding Device 1431 Three-dimensional data generation model acquisition department 1432 Buffer Section 1433 Differential Calculation Department 1434 Network Model Encoding Department 1435 Decoding Device 1436 Network Model Decoding Department 1437 Addition Department 1438 Buffer Section 1439 Rendering Department 1450 encoding device 1451 Extended 3D Data Generation Model Acquisition Department 1452 Buffer Section 1453 Network Model Coding Department 1455 Decoding Device 1456 Network Model Decoding Department 1457 Rendering Department 1460 encoding device 1461 Extended 3D Data Generation Model Acquisition Department 1462 Buffer Section 1463 Differential Calculation Department 1464 Network Model Coding Department 1465 decoding device 1466 Network Model Decoding Department 1467 Addition Department 1468 Buffer Section 1469 Rendering Department 1470 encoding device 1471 circuit 1472 memory 1480 decoding device 1481 circuit 1482 memory 1490 encoding device 1491 processor 1492 memory 1495 Decoding Device 1496 processor 1497 memory 1500 encoding device 1501 NeRF 3D Generative Model Learning Department 1502 NeRF Generative Model Encoding Section 1510 Decoding Device 1511 NeRF Generative Model Decoding Unit 1512 Rendering Data Reconstruction Department 1520 encoding device 1521 3D Generative Model Learning Department 1522 First Online Learning Department 1522a First Network Department 1522b Comparison Section 1522c Rendering Department 1523 Sampling Point Decision Department 1524 Second Online Learning Department 1525 3D Generative Model Coding Department 1526 First Network Coding Department 1527 Second Network Coding Department 1528 Metadata Encoding Department 1529-bit stream data composition section 1530 Decoding Device 1531-bit stream data segmentation unit 1532 3D Generative Model Decoding Department 1533 First Network Decoding Department 1534 Second Network Decoding Department 1535 Metadata Decoding Department 1536 Reconstruction Department 1537 Density Estimation Section 1538 Sampling Points Decision Department 1539 Attribute Information Estimation Department 1540 Rendering Department 1541 Decoding Department 1542 Control Department 1550 encoding device 1551 3D Generative Model Learning Department 1552 Geometric Generating Unit 1553 3D Generative Model Coding Department 1554-point group coding department 1560 decoding device 1561-bit stream data segmentation unit 1562 3D Generative Model Decoding Department 1563-point group decoding department 1564 Reconstruction Department 1570 encoding device 1571 circuit 1572 memory 1580 decoding device 1581 circuit 1582 memory

Claims

1. An encoding device, characterized in that, have: Circuits; and The memory is connected to the circuit. The circuit acquires multiple networks constituting a 3D data generation model during operation, encodes the multiple networks, and outputs a bitstream containing the encoded multiple networks. The bitstream includes: (1) first metadata, which stores parameters common to the sequence; and (2) second metadata, which stores parameters common to frames, access units, or multiple frames.

2. The encoding device according to claim 1, characterized in that, At least one of the first metadata and the second metadata includes a first parameter representing the number of layers in the first network of the plurality of networks and a second parameter representing the number of nodes in each layer.

3. The encoding device according to claim 2, characterized in that, The second metadata contains the same specific parameters as those contained in the first metadata. The specific parameter contained in the second metadata takes precedence over the specific parameter contained in the first metadata.

4. The encoding device according to claim 2, characterized in that, The first metadata contains a flag indicating which of the first or second parameters is included in the first metadata and the second metadata.

5. The encoding device according to any one of claims 1 to 4, characterized in that, The header of each cell in the bitstream contains an identifier representing the three-dimensional data generation model.

6. The encoding device according to claim 5, characterized in that, A frame corresponding to the three-dimensional data generation model is divided into multiple spaces. At least one of the plurality of spaces is divided into a plurality of subspaces. The identifier contains different values ​​corresponding to the plurality of subspaces respectively.

7. The encoding device according to any one of claims 1 to 4, characterized in that, The header of each unit in the bitstream contains an identifier that represents the frame number of the frame corresponding to the three-dimensional data generation model.

8. The encoding device according to any one of claims 1 to 4, characterized in that, The header of each unit in the bitstream contains an identifier indicating which region of the multiple spaces into which the three-dimensional data generation model is segmented.

9. The encoding device according to any one of claims 1 to 4, characterized in that, The first metadata includes a parameter representing the number of components constituting the bitstream and an identifier that identifies the components.

10. A decoding device, characterized in that, have: Circuits; and The memory is connected to the circuit. The circuit acquires a bitstream during operation, which contains multiple encoded networks that constitute a three-dimensional data generation model. The bitstream includes: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames. The circuit decodes the plurality of networks based on the bit stream during operation.

11. The decoding device according to claim 10, characterized in that, At least one of the first metadata and the second metadata includes a first parameter representing the number of layers in the first network of the plurality of networks and a second parameter representing the number of nodes in each layer.

12. The decoding device according to claim 11, characterized in that, The second metadata contains the same specific parameters as those contained in the first metadata. The specific parameter contained in the second metadata takes precedence over the specific parameter contained in the first metadata.

13. The decoding device according to claim 11, characterized in that, The first metadata contains a flag indicating which of the first or second parameters is included in the first metadata and the second metadata.

14. The decoding apparatus according to any one of claims 10 to 13, characterized in that, The header of each cell in the bitstream contains an identifier representing the three-dimensional data generation model.

15. The decoding apparatus according to claim 14, characterized in that, A frame corresponding to the three-dimensional data generation model is divided into multiple spaces. At least one of the plurality of spaces is divided into a plurality of subspaces. The identifier contains different values ​​corresponding to the plurality of subspaces respectively.

16. The decoding apparatus according to any one of claims 10 to 13, characterized in that, The header of each unit in the bitstream contains an identifier that represents the frame number of the frame corresponding to the three-dimensional data generation model.

17. The decoding apparatus according to any one of claims 10 to 13, characterized in that, The header of each unit in the bitstream contains an identifier indicating which region of the multiple spaces into which the three-dimensional data generation model is segmented.

18. The decoding apparatus according to any one of claims 10 to 13, characterized in that, The first metadata includes a parameter representing the number of components constituting the bitstream and an identifier that identifies the components.

19. An encoding method, executed by an encoding device, characterized in that, Obtain the three-dimensional data to generate the model. The multiple networks constituting the three-dimensional data generation model are encoded. The output contains a bitstream of the encoded networks. The bitstream includes: (1) first metadata, which stores parameters common to the sequence; and (2) second metadata, which stores parameters common to frames, access units, or multiple frames.

20. A decoding method, executed by a decoding device, characterized in that, Obtain a bitstream containing multiple encoded networks. The bitstream includes: (1) first metadata, storing parameters common to the sequence; and (2) second metadata, storing parameters common to frames, access units, or multiple frames. The multiple networks are decoded based on the bitstream.

Citation Information

Patent Citations

  • Map display device

    WO2014020663A1