Encoding device, decoding device, encoding method, and decoding method
By using encoding and decoding devices and a neural network model to generate bitstreams, the problem of low compression efficiency of 3D data is solved, and the amount of data is reduced and 3D data output at different resolutions is achieved.
Patent Information
- Application Number
- CN202480040504.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-30
- Filing Date
- 2024-06-25
- Publication Date
- 2026-01-13
AI Technical Summary
In existing technologies, the large amount of 3D data leads to low compression efficiency during accumulation or transmission, especially when generating motion images from arbitrary viewpoints, where the data volume requirement is high.
By using an encoding and decoding device and a neural network learning model to generate and decode bitstreams, a two-dimensional image containing viewpoint information is generated, thereby achieving compression of three-dimensional data.
It effectively reduces the storage capacity and network bandwidth requirements for storing and transmitting motion image data, and enables the output of 3D data at different resolutions.
Smart Images

Figure CN121336408A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to encoding devices, decoding devices, encoding methods, and decoding methods. Background Technology
[0002] Devices and services utilizing 3D data are expected to become increasingly common in a wide range of fields, including computer vision, mapping information, surveillance, infrastructure inspection, and image distribution, which enable autonomous movement of vehicles or robots. 3D data is acquired through various methods, such as distance sensors like rangefinders, stereo cameras, or combinations of multiple monocular cameras.
[0003] As one method of representing 3D data, there is the so-called point cloud method, which uses a group of points in 3D space to represent the shape of a 3D structure. For point clouds, the position and color of the point groups are preserved. It is envisioned that point clouds will become the mainstream method of representing 3D data, but the data volume of point groups is extremely large. Therefore, in the accumulation or transmission of 3D data, similar to 2D moving images (for example, MPEG-4 AVC or HEVC, which are standardized through MPEG), data compression based on encoding is necessary.
[0004] In addition, point cloud compression is partially supported by publicly available libraries that perform point cloud association processing (such as Point Cloud Library).
[0005] In addition, there are known technologies that use three-dimensional map data to retrieve and display facilities located around a vehicle (for example, see Patent Document 1).
[0006] Existing technical documents Patent documents Patent Document 1: International Publication No. 2014 / 020663 Non-patent literature Non-patent document 1: ISO / IEC 15938-17:2022 (Information technology - Multimedia content description interface - Part 17: Compression of neural networks for multimedia content description and analysis (https / / www.iso.org / standard / 78480.html)) Summary of the Invention
[0007] The problem that the invention aims to solve The purpose of this disclosure is to provide an encoding device, etc., that can reduce the amount of data obtained from motion images from any viewpoint.
[0008] Methods for solving problems An encoding device disclosed herein comprises: a circuit; and a memory connected to the circuit. During operation, the circuit acquires a first three-dimensional data generation model corresponding to a first time moment and a second three-dimensional data generation model corresponding to a second time moment. It generates a bitstream by encoding the acquired first and second three-dimensional data generation models. The first and second three-dimensional data generation models output two-dimensional images of the subject as observed from the viewpoint and the viewing direction, respectively, when input with viewpoint information including the viewpoint and the viewing direction.
[0009] The decoding device involved in one of the technical solutions disclosed herein includes: a circuit; and a memory connected to the circuit. During operation, the circuit acquires a bitstream and decodes from the bitstream a first three-dimensional data generation model corresponding to a first time moment and a second three-dimensional data generation model corresponding to a second time moment. The first three-dimensional data generation model and the second three-dimensional data generation model output two-dimensional images of the subject observed from the viewpoint and the viewing direction, respectively, when input with viewpoint information including viewpoint and viewing direction.
[0010] Furthermore, these general or specific technical solutions can be implemented through systems, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or through any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0011] Invention Effects The decoding device disclosed herein can output three-dimensional data at different resolutions. Attached Figure Description
[0012] Figure 1 This is a diagram illustrating an example of the configuration of a three-dimensional data encoding and decoding system according to Embodiment 1.
[0013] Figure 2 This is a diagram illustrating the structure of the point group data in Implementation Method 1.
[0014] Figure 3 This is a diagram illustrating an example of the structure of a data file that describes information about point group data in Embodiment 1.
[0015] Figure 4 This is a diagram illustrating the composition of the three-dimensional mesh data in Implementation Method 1.
[0016] Figure 5 This is a diagram illustrating an example of the structure of a data file containing information about three-dimensional mesh data in Embodiment 1.
[0017] Figure 6 This is a diagram used to illustrate the three-dimensional model in Implementation Method 1.
[0018] Figure 7 This is a diagram illustrating the types of three-dimensional data in Implementation Method 1.
[0019] Figure 8 This is a diagram used to illustrate the encoding process of three-dimensional data in Implementation Method 1.
[0020] Figure 9 This is a diagram used to illustrate the decoding process of three-dimensional data in Implementation Method 1.
[0021] Figure 10 This is a two-dimensional schematic diagram showing the tiles and slices of the three-dimensional data in Embodiment 1.
[0022] Figure 11 This is a block diagram illustrating an example of the functional configuration of the server and terminal in Embodiment 1.
[0023] Figure 12 This is a block diagram illustrating another example of the data generation unit of the server in Embodiment 1.
[0024] Figure 13 This is a diagram used to illustrate the relationship between the three-dimensional space and the encoded data in Implementation Method 1.
[0025] Figure 14 This is a diagram illustrating an example of the syntax of the encoding method unit in Implementation Method 1.
[0026] Figure 15 This is a diagram illustrating an example of the syntax of the encoded point group in Implementation Method 1.
[0027] Figure 16 This is a diagram illustrating an example of the syntax of the encoding grid in Implementation Method 1.
[0028] Figure 17 This is a diagram illustrating an example of the syntax of the encoded three-dimensional model in Implementation Method 1.
[0029] Figure 18 This is a diagram illustrating an example of the syntax of three-dimensional data information in Implementation Method 1.
[0030] Figure 19 This is a diagram used to illustrate the data structure of the coding point group in Implementation Method 1.
[0031] Figure 20 This is a diagram used to illustrate the data structure of the coded grid in Implementation Method 1.
[0032] Figure 21 This is a diagram used to illustrate the data structure of the encoded three-dimensional model in Implementation Method 1.
[0033] Figure 22 This is a two-dimensional diagram illustrating an example of multiple three-dimensional spaces in Embodiment 1.
[0034] Figure 23 This is a diagram illustrating an example of a boundary box in Embodiment 1.
[0035] Figure 24 This is a diagram illustrating an example of the syntax for three-dimensional spatial information in Implementation Method 1.
[0036] Figure 25 This is a flowchart illustrating an example of partial decoding in Implementation Method 1.
[0037] Figure 26 This is a diagram illustrating an example of a three-dimensional spatial region that becomes part of the decoded object in Embodiment 1.
[0038] Figure 27 This is a diagram illustrating an example of the data structure of a partially decoded group of encoded points in Implementation Method 1.
[0039] Figure 28 This is a diagram illustrating an example of the data structure of the partially decoded encoded grid in Embodiment 1.
[0040] Figure 29 This is a diagram illustrating an example of the data structure of a partially decoded encoded 3D model in Implementation Method 1.
[0041] Figure 30 This is a diagram illustrating an example of the configuration of the decoding device in Embodiment 1.
[0042] Figure 31 This is a flowchart illustrating an example of a decoding method using the decoding apparatus in Embodiment 1.
[0043] Figure 32 This is a flowchart illustrating another example of a decoding method using a decoding device.
[0044] Figure 33 This is a diagram illustrating an example of the configuration of an encoding device.
[0045] Figure 34 This is a flowchart illustrating an example of an encoding method using an encoding device.
[0046] Figure 35 This is a diagram used to illustrate the processing during the learning of the three-dimensional generative model in Implementation Method 2.
[0047] Figure 36This diagram illustrates the process of generating still images of a subject from any viewpoint using a three-dimensional generative model in Embodiment 2.
[0048] Figure 37 This is a diagram used to illustrate the motion image generation method using the three-dimensional data generation model of Embodiment 1 in Embodiment 2.
[0049] Figure 38 This is a diagram showing a first example of the configuration of the encoding device in Embodiment 1 of Embodiment 2.
[0050] Figure 39 This is a diagram showing a first example of the configuration of the decoding device in Embodiment 1 of Embodiment 2.
[0051] Figure 40 This is a diagram illustrating a second example of the configuration of the encoding device in Embodiment 1 of Embodiment 2.
[0052] Figure 41 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 1 of Embodiment 2.
[0053] Figure 42 This is a diagram illustrating the motion image generation method using the extended three-dimensional data generation model of Embodiment 2 in Embodiment 2.
[0054] Figure 43 This is a diagram showing a first example of the configuration of the encoding device in Embodiment 2 of Embodiment 2.
[0055] Figure 44 This is a diagram showing a first example of the configuration of the decoding device in Embodiment 2 of Embodiment 2.
[0056] Figure 45 This is a diagram illustrating a second example of the configuration of the encoding device in Embodiment 2 of Implementation 2.
[0057] Figure 46 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 2 of Embodiment 2.
[0058] Figure 47 This is a diagram illustrating a motion image generation method using an extended three-dimensional data generation model, which is used to explain a variation of Embodiment 2.
[0059] Figure 48 This is a diagram illustrating a motion image generation method using a three-dimensional data generation model, which is used to explain a variation of Embodiment 2.
[0060] Figure 49 This is a diagram showing an example of the configuration of the encoding device in Embodiment 2.
[0061] Figure 50 This is a flowchart illustrating an example of the encoding method of the encoding device in Embodiment 2.
[0062] Figure 51 This is a diagram showing an example of the configuration of the decoding device in Embodiment 2.
[0063] Figure 52 This is a flowchart illustrating an example of a decoding method of the decoding device in Embodiment 2.
[0064] Figure 53 This is a diagram illustrating an example of the configuration of an encoding device.
[0065] Figure 54 This is a diagram illustrating an example of the configuration of a decoding device. Detailed Implementation
[0066] The encoding device of the first technical solution disclosed herein includes: a circuit; and a memory connected to the circuit. During operation, the circuit acquires a first three-dimensional data generation model corresponding to a first time moment and a second three-dimensional data generation model corresponding to a second time moment. It generates a bitstream by encoding the acquired first three-dimensional data generation model and the second three-dimensional data generation model. When the first three-dimensional data generation model and the second three-dimensional data generation model are input with viewpoint information including a viewpoint and a viewing direction, they respectively output a two-dimensional image of the subject as viewed from the viewpoint and the viewing direction.
[0067] Therefore, it is possible to generate a bitstream containing a first three-dimensional data generation model that generates a two-dimensional image corresponding to a first moment based on arbitrary viewpoint information and a second three-dimensional data generation model that generates a two-dimensional image corresponding to a second moment. Thus, it is possible to generate a bitstream that compresses the data of the motion image obtained from the arbitrary viewpoint. Therefore, it is possible to reduce the storage capacity used to store the data of the motion image obtained from the arbitrary viewpoint, or the network bandwidth used to transmit the data.
[0068] The encoding device involved in the second technical solution of this disclosure, in the encoding device involved in the first technical solution, the first three-dimensional data generation model and the second three-dimensional data generation model respectively use neural network learning models.
[0069] The encoding apparatus of the third technical solution disclosed herein, in the encoding apparatus of the first or second technical solution, includes a bit stream containing first time information representing the first time moment and second time information representing the second time moment.
[0070] The encoding apparatus of the fourth technical solution disclosed herein, in the encoding apparatus of the third technical solution, includes a bit stream containing a first frame number corresponding to the first time moment and a second frame number corresponding to the second time moment.
[0071] The encoding apparatus of the fifth technical solution disclosed herein, in any of the encoding apparatuses of the first to fourth technical solutions, the bit stream includes frame rate information related to the frame rate of a plurality of learning images used in the generation of the first three-dimensional data generation model and the second three-dimensional data generation model, wherein the plurality of learning images are two-dimensional images obtained by shooting at a plurality of different timings.
[0072] The encoding apparatus involved in the sixth technical solution of this disclosure, in any of the encoding apparatuses involved in the first to fourth technical solutions, the bit stream contains viewpoint information, which includes the viewpoint and viewing direction of multiple learning images used in the generation of the first three-dimensional data generation model and the second three-dimensional data generation model.
[0073] The encoding device involved in the seventh technical solution of this disclosure, in the encoding device involved in the sixth technical solution, wherein the plurality of learning images are two-dimensional images obtained by taking pictures of the subject from mutually different viewpoints and viewing directions, and the viewpoint information includes the mutually different viewpoints and viewing directions.
[0074] The encoding device involved in the eighth technical solution of this disclosure, in the encoding device involved in any of the first to seventh technical solutions, in the encoding of the second three-dimensional data generation model, the circuit calculates difference information representing the difference between the first three-dimensional data generation model and the second three-dimensional data generation model, and the bit stream contains the difference information.
[0075] The encoding device involved in the ninth technical solution of this disclosure, in the encoding device involved in the eighth technical solution, the difference includes the difference between the weight parameters corresponding to the nodes contained in the first three-dimensional data generation model and the second three-dimensional data generation model.
[0076] The encoding apparatus involved in the tenth technical solution of this disclosure, in the encoding apparatus involved in the eighth or ninth technical solution, the bit stream includes reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.
[0077] In the encoding apparatus of the eleventh technical solution of this disclosure, in the encoding apparatus of any of the first to tenth technical solutions, the first moment corresponds to a random access point, and the first three-dimensional data generation model is encoded by intra-frame prediction or by inter-frame prediction with a prediction value of 0.
[0078] The encoding device involved in the twelfth technical solution of this disclosure, in the encoding device involved in the eleventh technical solution, the first three-dimensional data generation model and the second three-dimensional data generation model are included in one of a plurality of sets, and the first three-dimensional data generation model is the beginning of the data order among the plurality of three-dimensional data generation models included in the set.
[0079] The encoding apparatus of the thirteenth technical solution of this disclosure, in the encoding apparatus of the twelfth technical solution, wherein the bit stream contains licensing information indicating whether the three-dimensional data generation model is permitted to refer to the three-dimensional data generation models contained in other sets during the encoding of the plurality of three-dimensional data generation models.
[0080] In the encoding apparatus of the fourteenth technical solution of this disclosure, in any of the encoding apparatuses of the first to thirteenth technical solutions, the first three-dimensional data generation model corresponds to a first period including the first time moment, and the second three-dimensional data generation model corresponds to a second period including the second time moment.
[0081] The encoding apparatus of the fifteenth technical solution of this disclosure, in the encoding apparatus of the fourteenth technical solution, uses multiple first learning images in the generation of the first three-dimensional data generation model as two-dimensional images obtained by shooting at multiple different timings during the first period.
[0082] The encoding apparatus of the sixteenth technical solution of this disclosure, in the encoding apparatus of the fourteenth or fifteenth technical solution, outputs a two-dimensional image of the subject at the input time when the first three-dimensional data generation model is input at a time included in the first period.
[0083] The encoding apparatus involved in the seventeenth technical solution of this disclosure, in the encoding apparatus involved in any of the fourteenth to sixteenth technical solutions, the bit stream contains quantity information indicating the upper limit number of images that the first three-dimensional data generation model can generate.
[0084] The encoding apparatus of the eighteenth technical solution of this disclosure, in the encoding apparatus of the fifteenth technical solution, the bit stream includes first information related to the plurality of first learning images, the first information including multiple viewpoints and multiple line-of-sight directions corresponding to the plurality of first learning images, and multiple different timings.
[0085] In the encoding device of the nineteenth technical solution of this disclosure, in the encoding device of any one of the fourteenth to eighteenth technical solutions, the first period or the second period is dynamically determined according to the subject.
[0086] In the encoding device involved in the twentieth technical solution of this disclosure, in any of the encoding devices involved in the first to nineteenth technical solutions, the circuit saves the generated first three-dimensional data generation model in the memory, and the circuit generates the second three-dimensional data generation model based on the first three-dimensional data generation model saved in the memory.
[0087] In the encoding device involved in the twenty-first technical solution of this disclosure, in any of the encoding devices involved in the first to nineteenth technical solutions, the circuit saves the generated first three-dimensional data generation model and the second three-dimensional data generation model in the memory, the circuit generates an initial model based on the first three-dimensional data generation model and the second three-dimensional data generation model saved in the memory, and the circuit generates a third three-dimensional data generation model corresponding to the third time step based on the initial model.
[0088] The decoding device involved in the twenty-second technical solution of this disclosure includes: a circuit; and a memory connected to the circuit. The circuit, in operation, acquires a bit stream and decodes from the bit stream a first three-dimensional data generation model corresponding to a first time moment and a second three-dimensional data generation model corresponding to a second time moment. The first three-dimensional data generation model and the second three-dimensional data generation model respectively output a two-dimensional image of the subject observed from the viewpoint and the viewing direction when viewpoint information including the viewpoint and the viewing direction is input.
[0089] Therefore, it is possible to generate a bitstream containing a first three-dimensional data generation model that generates a two-dimensional image corresponding to a first moment based on arbitrary viewpoint information and a second three-dimensional data generation model that generates a two-dimensional image corresponding to a second moment. Thus, it is possible to generate a bitstream that compresses the data of the motion image obtained from the arbitrary viewpoint. Therefore, it is possible to reduce the storage capacity used to store the data of the motion image obtained from the arbitrary viewpoint, or the network bandwidth used to transmit the data.
[0090] The decoding device involved in the twenty-third technical solution of this disclosure, in the decoding device involved in the twenty-second technical solution, the first three-dimensional data generation model and the second three-dimensional data generation model respectively use neural network learning models.
[0091] The decoding apparatus involved in the twenty-fourth technical solution of this disclosure, in the decoding apparatus involved in the twenty-second or twenty-third technical solutions, the bit stream includes first time information representing the first time moment and second time information representing the second time moment.
[0092] The decoding apparatus involved in the twenty-fifth technical solution of this disclosure, in the decoding apparatus involved in the twenty-fourth technical solution, the bit stream includes a first frame number corresponding to the first time moment and a second frame number corresponding to the second time moment.
[0093] The decoding apparatus involved in the twenty-sixth technical solution of this disclosure, in the decoding apparatus involved in any of the twenty-second to twenty-fifth technical solutions, the bit stream contains frame rate information related to the frame rate of a plurality of learning images used in the generation of the first three-dimensional data generation model and the second three-dimensional data generation model, the plurality of learning images being two-dimensional images obtained by shooting at a plurality of different timings.
[0094] The decoding device involved in the twenty-seventh technical solution of this disclosure, and the decoding device involved in any of the twenty-second to twenty-fifth technical solutions, wherein the bit stream contains viewpoint information, which includes the viewpoint and viewing direction of multiple learning images used in the generation of the first three-dimensional data generation model and the second three-dimensional data generation model.
[0095] The decoding device involved in the twenty-eighth technical solution of this disclosure, in the decoding device involved in the twenty-seventh technical solution, the plurality of learning images are two-dimensional images obtained by taking pictures of the subject from different viewpoints and viewing directions, and the viewpoint information includes the different viewpoints and viewing directions.
[0096] The decoding apparatus involved in the twenty-ninth technical solution of this disclosure, in the decoding apparatus involved in any of the twenty-second to twenty-eighth technical solutions, the bit stream includes differential information representing the difference between the first three-dimensional data generation model and the second three-dimensional data generation model.
[0097] In the decoding device of the thirtieth technical solution of this disclosure, the difference includes the difference between the nodes contained in the first three-dimensional data generation model and the second three-dimensional data generation model and the decoding device of the twenty-ninth technical solution.
[0098] In the decoding apparatus of the thirty-first technical solution of this disclosure, in the decoding apparatus of the twenty-ninth or thirtieth technical solution, the bit stream includes reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.
[0099] In the decoding apparatus involved in the thirty-second technical solution of this disclosure, in the decoding apparatus involved in any of the twenty-second to thirty-first technical solutions, the first moment corresponds to a random access point, and the first three-dimensional data generation model is decoded by intra-frame prediction or by inter-frame prediction with a prediction value of 0.
[0100] The decoding device involved in the thirty-third technical solution of this disclosure, in the decoding device involved in the thirty-second technical solution, the first three-dimensional data generation model and the second three-dimensional data generation model are included in one of a plurality of sets, and the first three-dimensional data generation model is the beginning of the data order among the plurality of three-dimensional data generation models included in the set.
[0101] The decoding apparatus involved in the thirty-fourth technical solution of this disclosure, in the decoding apparatus involved in the thirty-third technical solution, the bit stream contains permission information, which indicates whether the three-dimensional data generation model is permitted to refer to the three-dimensional data generation models contained in other sets during the decoding of the plurality of three-dimensional data generation models.
[0102] In the decoding device involved in the thirty-fifth technical solution of this disclosure, in the decoding device involved in any of the twenty-second to thirty-fourth technical solutions, the first three-dimensional data generation model corresponds to the first period including the first time moment, and the second three-dimensional data generation model corresponds to the second period including the second time moment.
[0103] The decoding device involved in the thirty-sixth technical solution of this disclosure, in the decoding device involved in the thirty-fifth technical solution, uses multiple first learning images in the generation of the first three-dimensional data generation model as two-dimensional images obtained by shooting at multiple different timings during the first period.
[0104] The decoding device involved in the thirty-seventh technical solution of this disclosure, in the decoding device involved in the thirty-fifth or thirty-sixth technical solutions, when the first three-dimensional data generation model is input at a time included in the first period, outputs a two-dimensional image of the subject at the input time.
[0105] The decoding apparatus involved in the thirty-eighth technical solution of this disclosure, in the decoding apparatus involved in any of the thirty-fifth to thirty-seventh technical solutions, the bit stream contains quantity information indicating the upper limit number of images that the first three-dimensional data generation model can generate.
[0106] The decoding device involved in the thirty-ninth technical solution of this disclosure, in the decoding device involved in the thirty-sixth technical solution, the bit stream includes first information related to the plurality of first learning images, the first information including multiple viewpoints and multiple line-of-sight directions corresponding to the plurality of first learning images, and multiple different timings.
[0107] In the decoding device involved in the fortieth technical solution of this disclosure, in the decoding device involved in any of the thirty-fifth to thirty-ninth technical solutions, the first period or the second period is dynamically determined according to the subject.
[0108] The decoding device involved in the forty-first technical solution of this disclosure, in the decoding device involved in any of the technical solutions from the twenty-second to the fortieth technical solutions, wherein the circuit saves the generated first three-dimensional data generation model in the memory, and the circuit generates the second three-dimensional data generation model based on the first three-dimensional data generation model saved in the memory.
[0109] The decoding device involved in the forty-second technical solution of this disclosure, in the decoding device involved in any of the technical solutions from the twenty-second to the fortieth technical solutions, wherein the circuit saves the generated first three-dimensional data generation model and the second three-dimensional data generation model in the memory, the circuit generates an initial model based on the first three-dimensional data generation model and the second three-dimensional data generation model saved in the memory, and the circuit generates a third three-dimensional data generation model corresponding to the third time step based on the initial model.
[0110] Furthermore, these general or specific technical solutions can be implemented through systems, integrated circuits, computer programs, or computer-readable recording media such as CD-ROMs, or through any combination of systems, methods, integrated circuits, computer programs, and recording media.
[0111] Hereinafter, embodiments will be described in detail with reference to the accompanying drawings. Furthermore, the embodiments described below represent specific examples of this disclosure. The numerical values, shapes, materials, constituent elements, arrangement positions of constituent elements, connection methods, steps, and order of steps shown in the following embodiments are examples and are not intended to limit this disclosure. Additionally, constituent elements in the following embodiments that are not described in the independent claims representing the highest-level concept will be described as arbitrary constituent elements.
[0112] (Implementation Method 1) The configuration of the three-dimensional data encoding and decoding system of this embodiment will be described. Figure 1 This is a diagram illustrating an example of the configuration of the three-dimensional data encoding and decoding system of this embodiment. (See diagram for example.) Figure 1 As shown, the three-dimensional data encoding and decoding system includes a three-dimensional data encoding system 1001, a three-dimensional data decoding system 1002, a sensor terminal 1003, and an external connection unit 1004.
[0113] The 3D data encoding system 1001 generates encoded data or reused data by encoding 3D data. Furthermore, the 3D data encoding system 1001 can be a single device or a system implemented by multiple devices. Additionally, the 3D data encoding device may include a portion of the multiple processing units included in the 3D data encoding system 1001.
[0114] The 3D data encoding system 1001 includes a 3D data generation system 1011, a prompting unit 1012, an encoding unit 1013, a multiplexing unit 1014, an input / output unit 1015, and a control unit 1016. The 3D data generation system 1011 includes a sensor information acquisition unit 1017 and a 3D data generation unit 1018.
[0115] The sensor information acquisition unit 1017 acquires sensor signals from the sensor terminal 1003 and outputs the sensor signals to the three-dimensional data generation unit 1018. The three-dimensional data generation unit 1018 generates three-dimensional data based on the sensor signals and outputs the three-dimensional data to the encoding unit 1013.
[0116] The prompting unit 1012 prompts the user with sensor signals or three-dimensional data. For example, the prompting unit 1012 displays information or images based on sensor signals or three-dimensional data.
[0117] The encoding unit 1013 encodes (compresses) the three-dimensional data and outputs the resulting encoded data, control information obtained during the encoding process, and other additional information to the multiplexing unit 1014. The additional information may include, for example, sensor signals.
[0118] The multiplexing unit 1014 generates multiplexed data by multiplexing the encoded data, control information, and additional information input from the encoding unit 1013. The format of the multiplexed data is, for example, a file format for storage or a packet format for transmission.
[0119] The input / output unit 1015 (e.g., a communication unit or interface) outputs multiplexed data to the outside. Alternatively, the multiplexed data is stored in an internal memory or other storage unit. The control unit 1016 (or application execution unit) controls each processing unit. That is, the control unit 1016 performs control such as encoding and multiplexing. The control unit 1016 can also perform control such as demultiplexing, decoding, or prompting.
[0120] Alternatively, sensor signals can be input to the encoding unit 1013 or the multiplexing unit 1014. Furthermore, the input / output unit 1015 can directly output 3D data or encoded data to the outside.
[0121] The transmission signal (multiplexed data) output from the three-dimensional data encoding system 1001 is input to the three-dimensional data decoding system 1002 via the external connection unit 1004.
[0122] The 3D data decoding system 1002 generates 3D data by decoding encoded or multiplexed data. Furthermore, the 3D data decoding system 1002 can be a single device or a system comprised of multiple devices. Additionally, the 3D data decoding device may include a portion of the multiple processing units included in the 3D data decoding system 1002.
[0123] The 3D data decoding system 1002 includes a sensor information acquisition unit 1021, an input / output unit 1022, a demultiplexing unit 1023, a decoding unit 1024, a prompting unit 1025, a user interface 1026, and a control unit 1027.
[0124] The sensor information acquisition unit 1021 acquires sensor signals from the sensor terminal 1003.
[0125] The input / output unit 1022 acquires the transmission signal, decodes the multiplexed data (file format or packet) from the transmission signal, and outputs the multiplexed data to the demultiplexing unit 1023.
[0126] The demultiplexing unit 1023 obtains encoded data, control information, and additional information from the multiplexed data, and outputs the encoded data, control information, and additional information to the decoding unit 1024.
[0127] The decoding unit 1024 reconstructs the point group data by decoding the encoded data.
[0128] The prompting unit 1025 prompts the user with dot group data. For example, the prompting unit 1025 displays information or images based on the dot group data. The user interface 1026 obtains instructions based on the user's operation. The control unit 1027 (or application execution unit) controls each processing unit. That is, the control unit 1027 performs demultiplexing, decoding, and prompting control, etc.
[0129] Furthermore, the input / output unit 1022 can directly acquire point group data or encoded data from external sources. Additionally, the prompting unit 1025 can acquire additional information such as sensor signals and provide prompts based on this additional information. Furthermore, the prompting unit 1025 can also provide prompts based on user instructions obtained from the user interface 1026.
[0130] The sensor terminal 1003 generates the information obtained from the sensor, namely the sensor signal. The sensor terminal 1003 is a terminal equipped with a sensor or camera, such as a moving object like a car, a flying object like an airplane, a portable terminal, or a camera.
[0131] Sensor signals that can be acquired by sensor terminal 1003 include, for example: (1) signals obtained from LIDAR, millimeter-wave radar, or infrared sensors indicating the distance between sensor terminal 1003 and an object or the reflectivity of the object; (2) signals obtained from multiple monocular camera images or stereo camera images indicating the distance between a camera and an object or the reflectivity of the object. Additionally, sensor signals may also include sensor posture, orientation, gyroscope (angular velocity), position (GPS information or altitude), velocity, or acceleration. Furthermore, sensor signals may also include temperature, air pressure, humidity, or magnetism.
[0132] The external connection unit 1004 is realized through communication with an integrated circuit (LSI or IC), an external storage unit, a cloud server via the Internet, or broadcasting.
[0133] Next, the point group data will be explained. Figure 2 It is a diagram showing the composition of point group data. Figure 3 This is a diagram illustrating an example of the structure of a data file that records information about point group data.
[0134] Point cluster data comprises data from multiple points. Each point's data includes location information (3D coordinates) and attribute information related to that location. A cluster of these points is called a point cluster. For example, a point cluster indicates the 3D shape of an object.
[0135] Position information, such as three-dimensional coordinates, is sometimes referred to as geometric information. Additionally, the data for each point can also contain attribute information across multiple attribute categories. Attribute categories could include, for example, color or reflectivity.
[0136] One attribute can be mapped to one location, or multiple attributes with different attribute categories can be mapped to one location. Furthermore, multiple mappings can be established between attribute categories and one location.
[0137] Figure 3 The example data file shown illustrates a one-to-one correspondence between location information and attribute information, illustrating the location and attribute information of N points constituting the point group data.
[0138] Positional information includes, for example, information along the x, y, and z axes. Attribute information includes, for example, RGB color information. Representative data files include .ply files, etc.
[0139] Next, the three-dimensional mesh data will be explained. Figure 4 It is a diagram showing the composition of three-dimensional mesh data. Figure 5 This is a diagram illustrating an example of the structure of a data file containing information about three-dimensional mesh data.
[0140] 3D mesh data is a data format used in Computer Graphics (CG) that indicates the three-dimensional shape of an object through a collection of face information. These face information points to polygons such as triangles or quadrilaterals. 3D mesh data is also known as polygonal data or polygonal meshes.
[0141] The constituent elements are a group of three-dimensional points, vertices of the multiple three-dimensional points forming the group, edges connecting two vertices of the multiple three-dimensional points, and a set of faces enclosed by the multiple edges. A group of three-dimensional points is a collection of points that contains positional information in three-dimensional space and attribute information corresponding to that positional information. Furthermore, a three-dimensional point can also be simply referred to as a point.
[0142] Vertices can also possess attributes such as color, reflectivity, and normal vectors specific to a 3D point. The relationships between vertices forming an edge or face can also be represented by connectivity. Furthermore, vertices can also be represented by position. The face and back faces can be represented by the orientation of the normal vectors relative to 3D points. Additionally, vertices can also possess surface-specific attribute information.
[0143] Grid data files can take the form of object files, for example. Figure 5 In the mesh data file shown, the position information G(1)~G(N) of the N vertices constituting the mesh, and the attribute information A(1)~A(N) of the vertices are represented as vertex information. In the mesh data file, vertex information may also not include attribute information.
[0144] Furthermore, attribute information does not necessarily have to correspond one-to-one with vertices. Figure 5 The mesh data file shows an example of three-dimensional mesh data with M attribute information A2.
[0145] Face information is represented by a combination of vertex indices. n[1,3,4] represents the face of a triangle formed by vertices n=1, n=3, and n=4.
[0146] Additionally, m[2,4,6] indicates that the attribute information in attribute information A2 with m=2, m=4, and m=6 correspond to three vertices respectively. Furthermore, an example of a face consisting of 3 vertices is shown here, but a face can have any number of vertices greater than or equal to 3, not limited to 3. For example, if the face is a quadrilateral, the number of vertices is 4; if the face is a polygon, the number of vertices is the same as the number of vertices constituting the polygon.
[0147] Furthermore, attribute information A2 can be represented by a file different from the mesh data file, or it can include its pointer information. For example, attribute information can be stored in a two-dimensional attribute map file, and the attribute map filename and the two-dimensional coordinates in the attribute map can also be represented by attribute information A2 from the mesh data file. In this way, attribute information A2 can be contained in the mesh data file or represented by a file different from the mesh data file; regardless of the method used, attribute information for three-dimensional points can be specified.
[0148] Next, the three-dimensional model will be explained. Figure 6 It is a diagram used to illustrate a three-dimensional model.
[0149] A 3D model is a model generated based on 2D or 3D data.
[0150] The 3D model learning unit 1031 generates a network model, i.e., a 3D model, by learning 2D data (2D images) or 3D data (point groups or meshes) and using neural networks to learn 3D shapes and corresponding attribute information.
[0151] The 3D model learning unit 1031 can also generate a 3D model by learning from a 2D image using NeRF (Neural Radiance Fields). Alternatively, the 3D model learning unit 1031 can generate a 3D model after transforming a 2D image into 3D data through photogrammetry using the 2D image. The 3D model can also be generated using 3D data obtained from a sensor (distance sensor).
[0152] Three-dimensional model data consists of the elements that constitute a three-dimensional model, containing information indicating the structure of the network model, feature quantities, etc. For example, three-dimensional model data contains information related to the constituent elements of a neural network. This information includes, for example, multiple layers such as the input layer, intermediate layers, and output layer, the nodes in each layer, the weight coefficients for each node, and the transformation functions of each node.
[0153] The 3D model encoding unit 1032 can also encode 3D model data and transmit the encoded 3D model data.
[0154] The 3D model decoding unit 1033 receives the transmitted encoded 3D model data and decodes the 3D model based on the encoded 3D model data.
[0155] The rendering and reconstruction unit 1034 reconstructs (generates) two-dimensional data (two-dimensional images) or three-dimensional data (point groups or meshes) based on the decoded three-dimensional model. For example, when using a three-dimensional model modeled via NeRF, the rendering and reconstruction unit 1034 obtains viewpoint position or line-of-sight vector information, generates rendered two-dimensional data (two-dimensional images) based on the three-dimensional model and the viewpoint position or line-of-sight vector, and outputs the two-dimensional data. The generated two-dimensional data indicates a two-dimensional image of a three-dimensional object observed from the viewpoint position or from the line of sight indicated by the line-of-sight vector. The three-dimensional object is a three-dimensional object that serves as the source of the two-dimensional or three-dimensional data input to the three-dimensional model learning unit 1031.
[0156] Next, the types of three-dimensional data will be explained. Figure 7 This is a diagram illustrating the types of three-dimensional data. For example... Figure 7 As shown, there are static and dynamic objects in 3D data.
[0157] Static objects are 3D data at any given time (a specific moment). Dynamic objects are 3D data that changes over time. Hereinafter, the point group data at a specific moment will be referred to as a PCC frame or frame. Additionally, grid data at any given time will be referred to as a grid frame or frame.
[0158] The object can be three-dimensional data with the area restricted to some extent, such as typical image data, or it can be three-dimensional data with the area unrestricted, such as map information.
[0159] In addition, points of various densities can exist, including sparse point clusters (sparse grid data) and dense point clusters (dense grid data).
[0160] The details of each processing unit are described below. Sensor information is obtained through various methods such as distance sensors like LIDAR or rangefinders, stereo cameras, or combinations of multiple monocular cameras. The 3D data generation unit 1018 generates point group data based on the sensor information obtained by the sensor information acquisition unit 1017. The 3D data generation unit 1018 generates position information (geometric information) as point group data and adds attribute information specific to that position information to the position information.
[0161] The 3D data generation unit 1018 can also process point group data during the generation of position information or the addition of attribute information. For example, the 3D data generation unit 1018 can reduce the amount of data by deleting point groups with overlapping positions. Furthermore, the 3D data generation unit 1018 can transform position information (position shifting, rotation, or normalization, etc.) and process point group data to generate mesh data. Additionally, the 3D data generation unit 1018 can also render attribute information.
[0162] In addition, Figure 1 In this system, the three-dimensional data generation system 1011 is included in the three-dimensional data encoding system 1001, but it can also be set independently outside the three-dimensional data encoding system 1001.
[0163] The encoding unit 1013 generates encoded data by encoding the three-dimensional data based on a predefined encoding method. Regarding encoding methods, there are G-PCC (encoding method using position information), V-PCC (encoding method using video codec), Draco (grid coding method), and V-DMC (grid coding method). The encoding method is not limited to these methods; for example, it can also be a method for encoding dynamic grids, or other methods combining these methods.
[0164] The decoding unit 1024 decodes the three-dimensional data by decoding the encoded data based on a predefined encoding method.
[0165] The multiplexing unit 1014 generates multiplexed data by multiplexing encoded data using existing multiplexing methods. The generated multiplexed data is transmitted or stored. In addition to multiplexing encoded data of 3D data, the multiplexing unit 1014 also multiplexes other media such as images, sounds, subtitles, applications, and files, or reference time information. Furthermore, the multiplexing unit 1014 can also multiplex attribute information associated with sensor information or point group data.
[0166] As multiplexing methods or file formats, there are ISOBMFF, ISOBMFF-based transmission methods such as MPEG-DASH, MMT, MPEG-2 TS Systems, and RTP.
[0167] The demultiplexing unit 1023 extracts encoded data, other media, and time information from the multiplexed data.
[0168] The input / output unit 1015 transmits multiplexed data using a method that matches the transmission medium or storage medium, such as broadcasting or communication. The input / output unit 1015 can communicate with other devices via the Internet, or with storage units such as cloud servers.
[0169] Use HTTP, FTP, TCP, or UDP as the communication protocol. You can also use either the PULL or PUSH communication method.
[0170] Either wired or wireless transmission can be used. For wired transmission, Ethernet (registered trademark), USB, RS-232C, HDMI (registered trademark), or coaxial cable can be used. For wireless transmission, wireless LAN, Wi-Fi (registered trademark), Bluetooth (registered trademark), or millimeter wave can be used.
[0171] In addition, broadcast methods can be used, such as DVB-T2, DVB-S2, DVB-C2, ATSC3.0, or ISDB-S3.
[0172] Next, the process of dividing 3D data into more than one 3D data segment will be explained. Figure 8 It is a diagram used to illustrate the encoding and processing of three-dimensional data. Figure 9 This is a diagram used to illustrate the decoding process of three-dimensional data.
[0173] like Figure 8 As shown, the data segmentation unit 1041 segments the three-dimensional data into one or more three-dimensional spaces, generating one or more segmented three-dimensional data (i.e., one or more segmented three-dimensional data). The encoding unit 1042 can also encode one or more segmented three-dimensional data to generate encoded data. The data segmentation unit 1041 and the encoding unit 1042 are components of an encoding device and can be included in one encoding device or in different devices.
[0174] One or more 3D spaces can be individually labeled as tiles or intervals. A 3D space can be, for example, a bounding box. Furthermore, the 3D data contained within each of the segmented 3D spaces can also be represented as slices. A slice is segmented 3D data, including any of the following: a group of points with location information (Geometry) or attribute information (Attribute), a mesh, or a 3D model. Each slice in the multiple slices is encoded by the encoding unit 1042 according to each constituent element and is output as encoded data. The encoded data includes the encoded slices.
[0175] like Figure 9 As shown, in the decoding process, the decoding unit 1051 decodes one or more segmented 3D data (one or more slices) based on the encoded data. The data combining unit 1052 combines the one or more segmented 3D data to restore (generate) 3D data. The decoding unit 1051 and the data combining unit 1052 are components of a decoding device and can be included in that decoding device or in different devices. The one or more segmented 3D data decoded by the decoding unit 1051 may not be combined. The decoding unit 1051 may also decode a portion of the one or more segmented 3D data based on a portion of the encoded data and output the decoded portion of the segmented 3D data. In this case, the decoding device may not have the data combining unit 1052.
[0176] Figure 10 It is a two-dimensional schematic diagram that illustrates the tiles and slices of three-dimensional data.
[0177] When encoding multiple slices, the encoding device can encode using the dependencies between the slices or without using dependencies. When encoding without dependencies, the encoding device can encode each slice independently, reducing processing time by encoding multiple slices in parallel. Similarly, when the decoding device encodes multiple slices without dependencies, it can decode each slice independently, reducing processing time by decoding multiple slices in parallel. Furthermore, the decoding device can reduce processing load by partially decoding only a portion of the multiple slices.
[0178] When encoding using dependency relationships, the encoding device transmits signals to identifiers indicating dependency relationships and encodes sequentially starting from the dependent data. When encoding multiple slices using dependency relationships, the decoding device decodes sequentially starting from the dependent data based on the identifiers.
[0179] In 3D data segmentation, any number of segments and any segmentation method can be used. The shape of an object can also be determined, and multiple 3D points can be segmented for each object. Alternatively, segmentation can be based on the number of 3D points contained in a slice; that is, an upper limit can be determined for the number of 3D points a slice can contain. Furthermore, 3D data can be segmented using map or location information, based on whether it is contained in 3D space (tile information). Multiple tile shapes can also overlap.
[0180] By dividing 3D data into multiple segmented 3D data in this way, adaptive encoding corresponding to the content or object can be performed, and parallel processing can be performed during decoding.
[0181] Next, the method for selecting the 3D data to be prompted or transmitted from multiple 3D data sets will be explained.
[0182] The server stores multiple 3D data sets for the same space. For example, the server stores point cluster data and mesh data for the same space. The server is an example of an encoding device. The terminal, based on its purpose, switches the 3D data obtained from the server and displays the switched 3D data. The terminal can also be a terminal that parses 3D data. In this case, the terminal can also switch the 3D data to be displayed based on purposes such as parsing or displaying, or user operations. The terminal is an example of a decoding device.
[0183] In switching between 3D data, the focus can be on whether to use a group of prompt points or a grid as the 3D data. Alternatively, the focus can be on whether to transmit a group of prompt points or a grid as the 3D data. For example, the terminal can send the user's selection to the server, receive (download) 3D data based on that selection from the server, and then provide prompts for the received 3D data. The 3D data (point group or grid) may or may not be encoded on the server. If the 3D data is encoded, the terminal can receive the encoded 3D data from the server, decode the 3D data based on the received encoded data, and then provide prompts for the decoded 3D data.
[0184] Next, the configuration of server 1070 and terminal 1090 will be explained. Figure 11 This is a block diagram illustrating an example of the functional configuration of a server and a terminal.
[0185] Server 1070 includes a data generation unit 1071, a synchronization unit 1075, a point group coding unit 1076, a grid coding unit 1077, a model coding unit 1078, a multiplexing unit 1079, and a data extraction unit 1080.
[0186] The data generation unit 1071 generates three-dimensional data based on at least one of two-dimensional data and three-dimensional data. The generated three-dimensional data includes at least two of point group data, mesh data, and three-dimensional model data. The data generation unit 1071 has a point group generation unit 1072, a mesh generation unit 1073, and a model generation unit 1074. The data generation unit 1071 only needs to have at least two of the point group generation unit 1072, mesh generation unit 1073, and model generation unit 1074. The point group generation unit 1072 generates point group data based on at least one of two-dimensional data and three-dimensional data. The mesh generation unit 1073 generates mesh data based on at least one of two-dimensional data and three-dimensional data. The model generation unit 1074 generates three-dimensional model data by performing machine learning based on at least one of two-dimensional data and three-dimensional data.
[0187] The two-dimensional data input to the data generation unit 1071 can also be two-dimensional images acquired by a camera. The three-dimensional data input to the data generation unit 1071 can, for example, be point data of spaces such as a building site, factory, or office acquired by a sensor such as LiDAR. The data generation unit 1071 can also generate color information as attribute information for each point contained in the point data of the three-dimensional data using the two-dimensional image of the two-dimensional data. The three-dimensional data generated by the data generation unit 1071 can also be divided into arbitrary spaces. Point data, mesh data, and three-dimensional model data can also be divided into arbitrary spaces respectively.
[0188] The synchronization unit 1075 acquires the spatial position or time (reproduction time, decoding time, acquisition time, etc.) of the point group data, mesh data, and 3D model data generated by the data generation unit 1071. The time of each data point is the reproduction time, decoding time, acquisition time, etc. Alternatively, the synchronization unit 1075 may not acquire the synchronization of the point group data, mesh data, and 3D model data, but instead generate synchronization information for acquiring synchronization. Furthermore, the synchronization unit 1075 may perform processing to acquire the synchronization of at least two types of 3D data from the point group data, mesh data, and 3D model data generated by the data generation unit 1071, or generate synchronization information (synchronization signal) for acquiring synchronization; alternatively, it may not perform the processing for acquiring the synchronization of all three types of 3D data (synchronization processing).
[0189] The point group encoding unit 1076 encodes the point group data that has been synchronized by the synchronization unit 1075. Alternatively, the point group encoding unit 1076 may not encode the point group data. The point group data may be pre-encoded or encoded upon request from the terminal 1090.
[0190] The grid encoding unit 1077 encodes the grid data that has been synchronized by the synchronization unit 1075.
[0191] The model encoding unit 1078 encodes the three-dimensional model data that has been synchronized by the synchronization unit 1075.
[0192] The multiplexing unit 1079 uses a prescribed format or a prescribed multiplexing method to multiplex the encoded point group data (coded point group), the encoded mesh data (coded mesh data), the encoded 3D model data, and the synchronization information. Alternatively, multiplexing based on the multiplexing unit 1079 may not be performed. In this case, the server 1070 may not have the multiplexing unit 1079.
[0193] The data extraction unit 1080 extracts a portion of the multiplexed 3D data corresponding to the request from the terminal 1090 and sends the extracted portion of 3D data to the terminal 1090. Alternatively, data extraction based on the data extraction unit 1080 may not be performed. In this case, the server 1070 may not have the data extraction unit 1080. Without data extraction based on the data extraction unit 1080, the server 1070 can send the 3D data multiplexed by the multiplexing unit 1079 to the terminal 1090. Furthermore, even without multiplexing based on the multiplexing unit 1079, the server 1070 can send encoded point group data (encoded point group), encoded mesh data (encoded mesh), encoded 3D model data (encoded 3D model), and synchronization information to the terminal 1090, or send a bitstream containing encoded point group data (encoded point group), encoded mesh data (encoded mesh), encoded 3D model data (encoded 3D model), and synchronization information to the terminal 1090.
[0194] The terminal 1090 includes a control unit 1091, a decoding unit 1092, and a prompting unit 1093.
[0195] The control unit 1091 will send a request for a portion of the 3D data to the server 1070. The control unit 1091 can also process user operations to determine a portion of the 3D data.
[0196] The decoding unit 1092 decodes a portion of the 3D data based on the bit stream (encoded data) obtained from the server 1070.
[0197] The prompting unit 1093 renders a portion of the decoded 3D data to provide a prompt.
[0198] Figure 11 The data generation unit 1071 can also be... Figure 12 The data generation unit 1110 shown is used to achieve this. Figure 12 This is a block diagram illustrating another example of the data generation unit of a server.
[0199] The data generation unit 1110 includes a point group generation unit 1111, a grid generation unit 1112, and a model generation unit 1113.
[0200] The point group generation unit 1111 has the same function as the point group generation unit 1072. The point group generation unit 1111 acquires point group data from the point group sensor 1101 and a two-dimensional image from the camera 1102, and generates point group data based on the point group data and the two-dimensional image. The point group data generated by the point group generation unit 1111 includes position information of each point and attribute information (such as color information) extracted from the two-dimensional image, wherein the attribute information corresponds to each point indicated by the position information.
[0201] The mesh generation unit 1112 generates mesh data based on the point group data generated by the point group generation unit 1111.
[0202] The model generation unit 1113 has the same function as the model generation unit 1074. The model generation unit 1113 acquires point group data from the point group sensor 1101 and two-dimensional images from the camera 1102, and performs machine learning based on the point group data and the two-dimensional images to generate three-dimensional model data.
[0203] like Figure 11 As explained, point cluster data, grid data, and 3D model data can also be generated independently. For example... Figure 12 As explained, grid data can also be generated from point cluster data. Furthermore, point cluster data can also be generated from grid data.
[0204] Mesh can be generated from point groups, and point groups can also be generated from mesh.
[0205] Furthermore, point cluster data, mesh data, and 3D model data can be generated by the server 1070, or by sensors or a terminal 1090 equipped with sensors. Sensors include, for example, a point cluster sensor 1101 and a camera 1102.
[0206] Next, the relationship between three-dimensional space and coded data will be explained. Figure 13 It is a diagram used to illustrate the relationship between three-dimensional space and coded data.
[0207] As mentioned above, three-dimensional data includes, for example, point group data, grid data, and any of the three-dimensional models.
[0208] like Figure 13As shown, when 3D data is divided into 3 3D data units in 3 3D spaces (tiles or intervals), the encoding device encodes each of the 3 3D data units separately and adds a header to perform data unitization. The header contains the identifier (Space_ID) of the space to which the encoded data of the data unit belongs, and the identifier (DataUnit_ID) of the data unit.
[0209] The data unit is further given a header containing the identifier of the data unit or the length information of the data unit, and the encoding method unit is generated through unitization.
[0210] Next, the syntax of the encoding method unit will be explained. Figure 14 This is a diagram illustrating an example of the syntax of an encoding scheme unit. Figure 15 This is a diagram illustrating an example of the syntax for encoding point groups. Figure 16 This is a diagram illustrating an example of the syntax of a coding grid. Figure 17 This is a diagram illustrating an example of the syntax for encoding a three-dimensional model.
[0211] The `unit_type` directive indicates the category of the data unit stored in the encoding mode unit. Thus, the category of the data unit stored in the encoding mode unit is specified.
[0212] length indicates the length of a data unit.
[0213] The data() function indicates the body of a data unit.
[0214] exist Figure 15 In this context, when `unit_type` is 0, it indicates that the data unit is the location information (geometric information) of a group of encoded points. When `unit_type` is 1, it indicates that the data unit is the attribute information of a group of encoded points. When `unit_type` is 2, it indicates that the data unit is the metadata of a group of encoded points.
[0215] exist Figure 16 In this context, when `unit_type` is 0, it indicates that the data unit is encoding the location information (geometric information) of the mesh. When `unit_type` is 1, it indicates that the data unit is encoding the attribute information of the mesh. When `unit_type` is 2, it indicates that the data unit is encoding the metadata of the mesh.
[0216] exist Figure 17 In the context of `unit_type` being 0, the data unit indicates that it encodes element 1 of the 3D model. With `unit_type` being 1, the data unit indicates that it encodes element 2 of the 3D model. With `unit_type` being 2, the data unit indicates that it encodes metadata of the 3D model.
[0217] in addition, Figures 15-17 The syntax shown is an example and is not limited to the above configuration. These syntaxes can be constructed using a portion of the syntax, or using types (categories) not mentioned above, and the order of the syntactic components can be rearranged. For example, in the syntax of the encoding mode unit, it can also be as follows: Figure 14 The configuration of the encoding mode unit, which is common to multiple encoding modes, is shown in that way. Figures 15-17 The unit_type, length, and data() shown are shown.
[0218] Additionally, a header can be assigned to the encoding unit to indicate its category. The categories of encoding units include, for example, `point_cloud_codec_unit` for point cloud data, `mesh_codec_unit` for mesh data, and `model_codec_unit` for 3D model data. This allows for the comprehensive processing of multiple encoding methods.
[0219] Figure 18 This is a diagram illustrating an example of the syntax for three-dimensional data information.
[0220] In terms of syntax, when multiple encoding methods are saved in a single format, the number of 3D data contained in that format (number_of_3Dformat) and the type of 3D data (format_type) can also be indicated, and data in each format can be saved. Therefore, it is possible to comprehensively process multiple encoding methods or 3D data, and to recognize multiple encoding methods or 3D data.
[0221] 3Ddata_info indicates the format structure information for storing multiple 3D data sets.
[0222] number_of_3Dformat indicates the number of 3D formats used.
[0223] `format_type` indicates the category of the format of the saved 3D data. For example, a number for `format_type` and the corresponding format can be determined as follows: `format_type` of 0 indicates that the saved 3D data is in the format of point cloud. `format_type` of 1 indicates that the saved 3D data is in the format of mesh. `format_type` of 2 indicates that the saved 3D data is in the format of G-PCC (g-pcc). `format_type` of 3 indicates that the saved 3D data is in the format of V-DMC (v-dmc). `format_type` of 4 indicates that the saved 3D data is in the format of 3D model.
[0224] Next, the data structure of the encoded data of multiple three-dimensional data will be explained according to each type of three-dimensional data. Figure 19 It is a diagram used to illustrate the data structure of a group of coded points. Figure 20 It is a diagram used to illustrate the data structure of the coded grid. Figure 21 It is a diagram used to illustrate the data structure of a coded 3D model.
[0225] For each type of three-dimensional data, the encoding device divides the three-dimensional data into multiple three-dimensional data according to each of the multiple spatial regions, and encodes the multiple three-dimensional data (i.e. multiple segmented three-dimensional data) separately to generate encoded data.
[0226] Each encoded data is assigned a header, and at least one of the data_unit_id and space_id is stored.
[0227] Here, `data_unit_id` is an identifier that identifies a data unit within the encoded data and is unique within the encoded data. Additionally, `space_id` indicates identification information for a spatial region. If either `data_unit_id` or `space_id` is common across multiple 3D datasets, it indicates the same value across all three-dimensional datasets.
[0228] exist Figures 19-21 In the example, data units with data_unit_id=0 in the encoded point group, data units with data_unit_id=3 in the encoded mesh, and data units with data_unit_id=0 in the encoded 3D model are all assigned space_id=1. This means that the 3D data is contained in the common 3D space indicated by Space_ID#1.
[0229] Data, including headers, can be contained in bitstream structures such as data units or encoding methods, or stored in file formats specified by ISOBMFF, such as various BOXes.
[0230] Next, the three-dimensional spatial information will be explained. Figure 22 It is a diagram that represents an example of multiple three-dimensional spaces in a two-dimensional way. Figure 23 This is a diagram showing an example of a bounding box. Figure 24 This is a diagram illustrating an example of syntax for three-dimensional spatial information.
[0231] In the syntax of three-dimensional spatial information, 3Dspace_info indicates information about the segmented three-dimensional space. 3Dspace_info can be used for partial decoding.
[0232] number_of_space indicates the number of three-dimensional spaces after partitioning.
[0233] space_id indicates the identifier of the segmented three-dimensional space.
[0234] Three-dimensional spatial information includes bounding box information as a specification Figure 23 Information about the bounding box shown.
[0235] The bounding box information includes bounding_box_xyz and bounding_box_whd.
[0236] `bounding_box_xyz` indicates the coordinates of the reference point of the bounding box. Figure 23 In the example, the coordinates of x, y, and z (x0, y0, z0) are used to represent it.
[0237] `bounding_box_whd` indicates the size of the bounding box. Figure 23 In the example, it can be represented by width w, height h, and depth d (w0, h0, d0).
[0238] Additionally, three-dimensional spatial information may include an identifier for each data unit of the encoded data. Furthermore, three-dimensional spatial information may also omit this identifier; that is, the identifier may not be transmitted as a signal.
[0239] pointcloud_id indicates the identifier of the data unit of the coded point group in the space corresponding to space_id.
[0240] mesh_id indicates the identifier of the data cell of the coded grid of the space corresponding to space_id.
[0241] model_id indicates the identifier of the data unit of the coded 3D model corresponding to space_id.
[0242] Furthermore, if the data unit shows a data_unit_id instead of a space_id, the identifier of each encoded data unit can be stored in the information indicating each space in the three-dimensional spatial information. This allows for the establishment of a correspondence between the three-dimensional spatial information and the segmented three-dimensional encoded data.
[0243] Alternatively, if the space_id is shown in the data unit, the three-dimensional spatial information can be mapped to an identifier for each data unit of the encoded data using the space_id. In this case, it is also possible not to store the identifier for each data unit of the encoded data.
[0244] Alternatively, the 3D spatial information of the point group data and the grid data can be commonalized by making the segmentation method, the origin of each segmented space, and the size of the bounding box the same in both the grid data and the point group data. Alternatively, the same 3D spatial information can be used in both the point group data and the grid data. This allows for the commonalization of 3D spatial information across different types of 3D data, enabling the use of the same 3D spatial information. By making 3D spatial information commonalized, switching between different types of 3D data (e.g., switching prompts or transmissions) becomes easier. Furthermore, in formats that integrate multiple 3D data sets, it is possible to utilize a single 3D spatial information across all 3D data sets instead of setting 3D spatial information for each set, thus reducing the amount of 3D spatial information.
[0245] In addition to point group data and grid data, it can also synchronize the three-dimensional spatial information of the three-dimensional model with other types of three-dimensional data, and can also make the three-dimensional spatial information of other types of three-dimensional data common.
[0246] Next, the relationship between the data structure of 3D data and partial decoding will be explained. Figure 25 This is a flowchart illustrating an example of partial decoding. Figure 26 This is a diagram illustrating an example of a three-dimensional spatial region of an object that is partially decoded. Figure 27 This is a diagram illustrating an example of the data structure of a partially decoded group of coded points. Figure 28 This is a diagram illustrating an example of a data structure for a partially decoded encoded grid. Figure 29 This is a diagram illustrating an example of the data structure of a partially decoded encoded 3D model.
[0247] In partial decoding, firstly, the decoding device determines the three-dimensional spatial region of the object to be partially decoded (S1001).
[0248] Next, the decoding device uses three-dimensional spatial information (3Dspace_info) to determine the region that overlaps with the three-dimensional spatial region of the object based on the bounding box information of multiple three-dimensional spatial regions, and obtains the space_id corresponding to the determined region (S1002).
[0249] Next, the decoding device obtains a data unit with the obtained space_id from the encoded data and decodes it (S1003). Thus, the decoding device performs partial decoding, decoding only a portion of the three-dimensional data. In partial decoding, the decoding device does not decode the entire three-dimensional data, but only a portion of it.
[0250] For example, such as Figure 26 As shown, when the three-dimensional space region of the object being partially decoded is shown in thick lines, the space_id of the obtained three-dimensional space is determined to be #2 based on the three-dimensional space information.
[0251] Then, as Figures 27-29 As shown, the encoded data of various three-dimensional data were used to establish corresponding data units with Space_id=#2 and then decoded.
[0252] In addition, the decoding device can also obtain the data unit ID from the three-dimensional spatial information instead of obtaining the space_id, obtain the data unit with the obtained data unit ID, and perform partial decoding.
[0253] In the above embodiments, point group data, mesh data, and 3D model data are exemplified as 3D data representing 3D objects, but the methods are not limited to these. For example, a 3D object may also be represented by multiple groups, each containing line-of-sight information indicating a line of sight and a 2D image obtained when viewing the 3D object from that line of sight. That is, data containing these multiple groups can also be processed as a type of 3D data. In addition, 3D data may also be data in other formats such as Gaussian splatting data.
[0254] Figure 30 This is a diagram illustrating an example of the configuration of a decoding device. Figure 31 This is a flowchart illustrating an example of a decoding method performed by a decoding device.
[0255] The decoding device 1130 includes a circuit 1131 and a memory 1132 connected to the circuit 1131.
[0256] Circuit 1131 performs the following actions.
[0257] Circuit 1131 acquires encoded data (S1021), the encoded data including: encoding method information (format), indicating one encoding method containing first data representing a three-dimensional object and second data representing the three-dimensional object; and identification information, indicating the three-dimensional space containing the three-dimensional object. Next, based on the encoded data, circuit 1131 decodes the first data and second data corresponding to the three-dimensional space (S1022). Next, circuit 1131 renders the first data to generate first prompt data for prompting (S1023). Next, circuit 1131 renders the second data to generate second prompt data for prompting (S1024). Next, circuit 1131 switches from the generated second prompt data to the first prompt data for prompting (S1025). Furthermore, the first prompt data and the second prompt data are, for example, two-dimensional data or three-dimensional data generated by the rendering and reconstruction unit 1034.
[0258] Therefore, based on the first and second data corresponding to the three-dimensional space, first and second prompt data are generated, and prompts are given by switching from the second prompt data to the first prompt data. This allows for prompting in a way that does not produce spatial deviation during the switching between the two data representing the three-dimensional object. Thus, the first and second prompt data can be appropriately used to provide prompts.
[0259] For example, the first data is point group data representing the three-dimensional object.
[0260] Therefore, since the prompt is given by switching from the second prompt data to the first prompt data based on the point group data, the prompt can be given in a way that does not produce spatial deviation in the switching between the two data representing the 3D object.
[0261] For example, the second data is mesh data representing the three-dimensional object.
[0262] Therefore, by switching from the second cue data based on grid data to the first cue data for cues, it is possible to switch cues in a way that does not produce spatial deviation when switching between the two data representing the 3D object.
[0263] For example, the second data is three-dimensional model data representing the three-dimensional object. The three-dimensional model data indicates a machine learning model obtained by performing machine learning on multiple sets of views and two-dimensional images.
[0264] Therefore, by switching from the second cue data based on the 3D model data to the first cue data for prompting, it is possible to switch the cue data in a way that does not produce spatial deviation when switching between the two data representing the 3D object.
[0265] For example, the second data is a two-dimensional image obtained when the three-dimensional object is viewed from a specified line of sight.
[0266] Therefore, by switching from second cue data based on a two-dimensional image to first cue data for cues, it is possible to switch between the two data representing a three-dimensional object in a way that does not produce spatial deviation.
[0267] For example, the circuit also receives a switching request for prompt data from the user. In the prompt, the circuit switches from the second prompt data to the first prompt data according to the switching request.
[0268] Therefore, it can switch at a time specified by the user.
[0269] For example, the circuit also receives an operation from the user to change the style of the prompt. In the prompt, the circuit changes the style of the prompt according to the operation, switching from the second prompt data to the first prompt data based on the change.
[0270] Therefore, it can switch at timed intervals corresponding to user actions.
[0271] For example, in the acquisition process, the circuit acquires the encoded data from the encoding device via a communication network. In the prompting process, the circuit switches from the second prompting data to the first prompting data based on the bandwidth of the communication network.
[0272] Therefore, it can switch according to the bandwidth of the communication network. For example, when the bandwidth of the communication network changes from less than the specified bandwidth to more than the specified bandwidth, it can switch from the second prompt data to the first prompt data to make a prompt.
[0273] For example, in the prompt, the circuit switches from the second prompt data to the first prompt data based on the capabilities of the circuit that is available.
[0274] Therefore, it is possible to switch according to the capability of the available circuit. For example, when the capability of the available circuit changes from less than the specified capability to more than the specified capability, it is possible to switch from the second prompt data to the first prompt data to provide a prompt.
[0275] For example, the encoded data includes synchronization information for synchronizing the coordinate system of the first data with the coordinate system of the second data. The circuit, in the prompt, provides prompts for both the first and second prompt data based on the synchronization information.
[0276] Therefore, it is possible to switch from the second prompt data to the first prompt data based on matching the coordinate systems of the first and second prompt data. Thus, it is possible to provide prompts in a way that minimizes spatial deviation when switching between the two data representing a 3D object.
[0277] For example, the circuit further determines whether to synchronize the coordinate system of the first data with the coordinate system of the second data. If the circuit determines that it is necessary to synchronize the coordinate system of the first data with the coordinate system of the second data, the circuit provides a prompt based on the synchronization information for both the first and second prompt data in the prompt.
[0278] Therefore, synchronous processing can be performed when needed and skipped when not needed. This could potentially reduce the processing load.
[0279] For example, the first data and the second data have a common structure in the first data and the second data, respectively.
[0280] Therefore, it is possible to reduce the amount of encoded data. Therefore, it is possible to reduce communication capacity.
[0281] For example, the encoded data includes spatial information for determining the three-dimensional space containing the three-dimensional object. The circuit also obtains an object region indicating a portion of the three-dimensional space. Based on the spatial information, the circuit determines first overlapping data, which is a portion of the first data and overlaps with the object region. In the decoding, the circuit decodes the determined first overlapping data.
[0282] Therefore, for example, the amount of data acquired can be reduced by acquiring only the first overlapping data. This reduces communication capacity. Furthermore, for example, only the first overlapping data can be decoded. This reduces processing load.
[0283] Alternatively, circuit 1131 can also be like Figure 32 The decoding method shown in the flowchart is followed. Figure 32 This is a flowchart illustrating another example of a decoding method performed by a decoding device.
[0284] Circuit 1131 decodes the encoding information representing the three-dimensional object and indicating a second encoding method different from the first encoding method of the first data (S1031). Circuit 1131 decodes the second data indicating the second encoding method by the encoding information (S1032). The second data is used to generate second prompt data for prompting.
[0285] Therefore, by decoding the second data of the second encoding method indicated by the encoding method information obtained through decoding, it is possible to obtain the second data of the second prompt data used to generate appropriate prompts.
[0286] Figure 33 This is a diagram illustrating an example of the configuration of an encoding device. Figure 34 This is a flowchart illustrating an example of an encoding method performed by an encoding device.
[0287] The encoding device 1140 includes a circuit 1141 and a memory 1142 connected to the circuit 1141.
[0288] Circuit 1141 performs the following actions.
[0289] Circuit 1141 generates encoding information representing the three-dimensional object and indicating a second encoding method different from the first encoding method of the first data (S1041). Circuit 1141 generates second data indicating the second encoding method of the encoding information (S1042). Circuit 1141 generates a bitstream containing the encoding information and the second data (S1043). The second data is used to generate second prompt data for prompting.
[0290] Therefore, since a bitstream containing encoding method information and second data is generated, the decoding device that obtained the bitstream can obtain second data for generating appropriate second prompt data.
[0291] (Implementation Method Two) A method for generating still images of a subject (three-dimensional object) observed from any viewpoint in still space using a learning-based model, i.e., a three-dimensional data generation model, is explained.
[0292] Figure 35 This is a diagram used to illustrate the processing during the learning of the three-dimensional generative model in Implementation Method 2. Figure 36 This diagram illustrates the process of generating still images of a subject from any viewpoint using a three-dimensional generative model in Embodiment 2.
[0293] Information processing devices acquire 3D data generation models through learning, thereby enabling the generation of still images observed from any viewpoint in static space. For example, there are 3D data generation models generated using methods such as Neural Radiance Fields (NeRF).
[0294] During learning, for example, the information processing device acquires learning data, which includes an image of viewpoint A (correct value) obtained from any viewpoint A, and viewpoint information (camera pose, etc.) of viewpoint A when the image was obtained. The viewpoint information may include viewpoint A and the direction of the line of sight from viewpoint A. The information processing device, for example, uses an evaluation function 1402 to optimize the parameters of the network included in the 3D data generation model, such that the difference between the generated image of viewpoint A output from the 3D data generation model 1401 by inputting the viewpoint information from the aforementioned learning data and the image of viewpoint A as the input image corresponding to viewpoint A is minimized. By performing this learning process using multiple learning data corresponding to multiple different viewpoints, the information processing device can obtain a more accurate 3D data generation model. Learning processing is performed on the learning data corresponding to each of the multiple viewpoints. That is, the same processing as the learning processing for viewpoint A is performed on each viewpoint.
[0295] During generation, if the information processing device inputs viewpoint information, for example, viewpoint B, into the learned 3D data generation model 1403, it outputs a generated image of viewpoint B. If it inputs viewpoint information of viewpoint Z, which is different from viewpoint B, it outputs a generated image of viewpoint Z. The viewpoint information of viewpoint B may include viewpoint B and the viewing direction from viewpoint B. The viewpoint information of viewpoint Z may include viewpoint Z and the viewing direction from viewpoint Z.
[0296] Thus, by learning and obtaining a 3D data generation model 1403, it is possible to generate still images observed from any viewpoint in static space. However, it is not possible to directly generate moving images.
[0297] In addition, Figure 36 The example shown is a 3D data generation model that generates an image of a given viewpoint if that viewpoint information is input. However, this is not a limitation, and the data output from the 3D data generation model can be in any form. For example, the 3D data generation model could also be a network model that outputs learned 3D data of the object space as point cluster data or mesh data. Thus, users can stereoscopically view 3D data such as point cluster data or mesh data in the object space, and can also use the point cluster data or mesh data to measure the dimensions of objects in the object space as 3D data output.
[0298] [Example 1] Figure 37This diagram illustrates the motion image generation method using the three-dimensional data generation model of Embodiment 1 in Embodiment 2. Furthermore, this embodiment describes an example of the configuration of an apparatus and a method for encoding or decoding the three-dimensional data generation models NNt0~NNt5 generated corresponding to times t0~t5. However, it is not limited to this; it can also be applied to an apparatus and method for encoding or decoding the three-dimensional data generation models at any time within any period.
[0299] This embodiment illustrates a method for generating a moving image of an object (subject) observed from any viewpoint using a three-dimensional data generation model. In this method, for example... Figure 37 As shown, by obtaining 3D data generation models corresponding to each time point, still images of an object observed from any viewpoint at each time point can be generated. By arranging the generated still images in chronological order, motion images can be generated. More specifically, in generating motion images for times t0 to t5, multiple 3D data generation models NNt0 to NNt5 corresponding to times t0 to t5 are generated through learning. The viewpoint information (camera pose, etc.) of the viewpoint A from which the motion images are to be generated is input into the generated 3D data generation models NNt0 to NNt5 corresponding to times t0 to t5. Thus, the generated images of viewpoint A at times t0 to t5 are output by the 3D data generation models NNt0 to NNt5. By concatenating these images in time, motion images of the object observed from viewpoint A at times t0 to t5 can be generated.
[0300] However, in this case, maintaining multiple 3D data generation models corresponding to multiple time points requires either a large storage capacity for storing the data of these multiple 3D data generation models on a storage device or a large network bandwidth for transmitting the data of these multiple 3D data generation models over a network. Therefore, the data size can also be reduced by using, for example, NNC (Neural Network Coding) in the MPEG (Moving Picture Experts Group) standard to encode the data of the multiple 3D data generation models corresponding to multiple time points. In this disclosure, a method for more effectively compressing this data is described.
[0301] NNC is shown in Non-Patent Document 1.
[0302] Figure 38 This is a diagram illustrating a first example of the configuration of the encoding device in Embodiment 1 of Implementation Method 2.
[0303] The encoding device 1420 includes a three-dimensional data generation model acquisition unit 1421, a buffer unit 1422, and a network model encoding unit 1423.
[0304] The 3D data generation model acquisition unit 1421 acquires learning data from times t0 to t5, and uses this learning data to generate 3D data generation models NNt0 to NNt5 for times t0 to t5 through learning. The learning data includes multiple viewpoint images obtained by photographing an object from one or more viewpoint positions along one or more viewing directions at each time t0 to t5, and one or more viewpoint information indicating one or more viewpoint positions and one or more viewing directions corresponding to the multiple viewpoint images. The one or more viewpoint information may also be the camera position and pose when each of the multiple viewpoint images was photographed. Furthermore, the learning data is not limited to this and may also include information obtained from other sensors. For example, the learning data may also include point cluster data and depth images obtained using LiDAR or TOF sensors at each time. This improves the accuracy of the 3D data generation model obtained through learning.
[0305] The buffer unit 1422 stores the three-dimensional data generation model at time t generated by the three-dimensional data generation model acquisition unit 1421. The buffer unit 1422 is implemented by a storage device such as a memory. The three-dimensional data generation model at time t stored in the buffer unit 1422 can also be used as an initial model when the three-dimensional data generation model acquisition unit 1421 acquires (generates) the three-dimensional data generation model after time t through learning. As a result, the learning time can be shortened and the accuracy of the three-dimensional data generation model after time t can be improved.
[0306] Furthermore, the buffer unit 1422 can also store multiple 3D data generation models corresponding to multiple time points. Therefore, for example, an initial model can be generated based on the multiple 3D data generation models stored in the buffer unit 1422, for example, through averaging or other processing. The 3D data generation model acquisition unit 1421 learns 3D data generation models after time t using this initial model, and can obtain a high-precision 3D data generation model. Furthermore, if the 3D data generation model acquisition unit 1421 does not refer to past 3D data generation models during learning, the encoding device 1420 may not need to include the buffer unit 1422. This reduces the amount of storage used as the buffer unit 1422.
[0307] The network model encoding unit 1423 encodes the three-dimensional data generation models NNt0~NNt5 obtained by the three-dimensional data generation model acquisition unit 1421 and outputs a bit stream.
[0308] Furthermore, as a network model encoding method, the data size can be reduced, for example, by using NNC data encoding from the MPEG standard. That is, the network model encoding unit 1423 uses NNC to encode the three-dimensional data generation models NNt0 to NNt5 and appends the encoding result to the bitstream. In other words, the network model encoding unit 1423 generates encoded data as the encoding result and generates a bitstream containing the encoded data.
[0309] Specifically, the network model encoding unit 1423 first encodes the 3D data generation model NNt0 at time t0 using NNC, and appends the encoding result to the bitstream. Next, the network model encoding unit 1423 encodes the 3D data generation model NNt1 at time t1 using NNC, and appends the encoding result to the bitstream. In this way, the network model encoding unit 1423 can also reduce the amount of encoding by sequentially encoding the 3D data generation model at each time using NNC and appending each encoding result to the bitstream.
[0310] Furthermore, at this time, the network model encoding unit 1423 can also append time information, which indicates the time corresponding to the encoded 3D data generation model, as metadata to the bitstream. Thus, by decoding and referring to the metadata contained in the bitstream, the decoding device can determine the time corresponding to the decoded 3D data generation model and can appropriately generate motion images of the object from any viewpoint.
[0311] In addition, metadata is not limited to time information; it can also include information related to the acquisition (generation) of learning data, or information required by the decoding device to generate motion images.
[0312] For example, the network model encoding unit 1423 may also attach information related to the camera's frame rate when acquiring (generating) learning data as metadata. Thus, the decoding device can decode the frame rate of the generated motion image from the bitstream and appropriately set that frame rate.
[0313] Alternatively, the network model encoding unit 1423 may append the frame number corresponding to each time moment as metadata to the bitstream instead of the time information, and use other parameters to associate each frame number with the time information. For example, the network model encoding unit 1423 may append the time information and frame rate of the first frame as metadata, and the decoding device may calculate the time information of each frame based on this metadata, thereby reducing the amount of encoding corresponding to the time information of each frame.
[0314] Furthermore, the network model encoding unit 1423 can also append viewpoint information from viewpoint images used during learning to the bitstream. Thus, the decoding device can, for example, generate high-quality motion images by preferentially selecting viewpoints close to the viewpoint positions corresponding to the images used during learning. This is because the closer the viewpoint position or time is to the time of learning, the more likely the 3D data generation model is to generate higher-quality viewpoint images.
[0315] Figure 39 This is a diagram illustrating a first example of the configuration of the decoding device in Embodiment 1 of Embodiment 2.
[0316] The decoding device 1425 includes a network model decoding unit 1426 and a rendering unit 1427.
[0317] The network model decoding unit 1426 acquires the bit stream and, based on the acquired bit stream, decodes the three-dimensional data generation model NNt0~NNt5 and metadata such as time information for times t0~t5.
[0318] The rendering unit 1427 uses the 3D data generation models NNt0~NNt5 decoded by the network model decoding unit 1426 and metadata such as time information to generate motion images of viewpoint A based on viewpoint information specified by the user or system. Specifically, the rendering unit 1427 inputs the viewpoint information of viewpoint A into the 3D data generation model NNt0 at time t0 to generate image IMGt0 of viewpoint A at time t0. Next, it inputs the viewpoint information of viewpoint A into the 3D data generation model NNt1 at time t1 to generate image IMGt1 of viewpoint A at time t1. The rendering unit 1427 applies the generation processing of these images at each time to each time t2~t5 to generate images IMGt2~IMGt5 of viewpoint A at times t2~t5. Furthermore, the rendering unit 1427 uses images IMGt0~IMGt5 and metadata such as time information to generate motion images of objects observed from viewpoint A at times t0~t5. The moving image may include, for example, images IMGt0~IMGt5 and cue time information for calculating the cue times of images IMGt0~IMGt5 based on times t0~t5.
[0319] Furthermore, the viewpoint information can change according to time. For example, viewpoint information of viewpoint A can be input into the 3D data generation model NNt0~NNt3 at times t0~t3, and viewpoint information of viewpoint B can be input into the 3D data generation model NNt4~NNt5 at times t4~t5. Thus, the rendering unit 1427 generates multiple images of the object observed from viewpoint A at times t0~t3, and generates multiple images of the object observed from viewpoint B at times t4~t5. In other words, the rendering unit 1427 can generate motion images of the observed object, where the viewpoint switches from viewpoint A to viewpoint B at time t4.
[0320] Furthermore, the rendering unit 1427 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time. Thus, the user can switch between generating moving images and generating still images depending on the application.
[0321] Furthermore, the rendering unit 1427 is not limited to generating moving or still images based on a 3D data generation model. For example, the rendering unit 1427 can also generate point group data or mesh data based on a 3D data generation model, and output the generated point group data or mesh data as dynamic point group data or dynamic mesh data. Thus, users can use dynamic 3D data of audiovisual objects such as HMDs (Head Mount Displays), and can also use the dynamic 3D data to measure the amount of motion of objects.
[0322] Figure 40 This is a second example of the configuration of the encoding device in Embodiment 1 of Embodiment 2.
[0323] The encoding device 1430 includes a three-dimensional data generation model acquisition unit 1431, a buffer unit 1432, a difference calculation unit 1433, and a network model encoding unit 1434.
[0324] The three-dimensional data generation model acquisition unit 1431 is the same as the three-dimensional data generation model acquisition unit 1421 of the encoding device 1420.
[0325] The buffer unit 1432 is the same as the buffer unit 1422 of the encoding device 1420, but it differs from the buffer unit 1422 in that it inputs the three-dimensional data generation model stored in the memory or the like as a reference three-dimensional data generation model into the difference calculation unit 1433.
[0326] The difference calculation unit 1433 calculates difference information, which represents the difference between the three-dimensional data generation models NNt0~NNt5 generated by the three-dimensional data generation model acquisition unit 1431 at times t0~t5 and the three-dimensional data generation models (hereinafter referred to as reference three-dimensional data generation models) generated by the three-dimensional data generation model acquisition unit 1431 before each time step. Here, the difference information may include the difference in the weight parameters of the nodes of each network model, etc. For example, the difference calculation unit 1433 obtains the three-dimensional data generation model NNt5 at time t5 from the three-dimensional data generation model acquisition unit 1431 and obtains the three-dimensional data generation model NNt4 at time t4 from the buffer unit 1432 as a reference three-dimensional data generation model.
[0327] Alternatively, the difference calculation unit 1433 can use the three-dimensional data generation model NNt5 and the three-dimensional data generation model NNt4, for example, to calculate the difference (change) between the weight parameters of the nodes in the network model of the three-dimensional data generation model NNt5 and the weight parameters of the nodes in the network model of the three-dimensional data generation model NNt4, and input the difference information representing the difference to the network model encoding unit 1434. Thus, the difference information is encoded by the network model encoding unit 1434. That is, the encoding device 1430 can also perform predictive encoding by encoding the difference between the predicted value and the information related to the network model in the three-dimensional data generation model NNt5 based on the prediction of the three-dimensional data generation model NNt4, thereby reducing the amount of data. Through such predictive encoding, for example, in cases where the changes in the three-dimensional data generation model are small over time, such as when the object hardly moves, the value of the encoded difference becomes smaller, thus improving encoding efficiency. For example, the encoding device 1430 can also be set to RNNt0 = 0 and RNNtn = NNt(n-1) (n is an integer value from 1 to 5), using the previous three-dimensional data generation model as a reference three-dimensional data generation model, and reducing the number of bits through predictive coding.
[0328] Furthermore, in the second example, the encoding device 1430 performs predictive encoding on information related to the network model in the three-dimensional data generation model NNt5 based on information related to the network model in the three-dimensional data generation model NNt4, but is not limited to this. For example, the encoding device 1430 may select a reference three-dimensional data generation model for prediction from one or more three-dimensional data generation models stored in the buffer 1432, and perform predictive encoding using the selected three-dimensional data generation model. In this case, the encoding device 1430 may append information representing the selected three-dimensional data generation model (reference three-dimensional data generation model information) to the bitstream in order to pass the selected three-dimensional data generation model to the decoding device. Thus, the encoding device 1430 can select the optimal reference three-dimensional data generation model from the viewpoint of encoding efficiency, thereby improving encoding efficiency. Furthermore, by decoding the reference three-dimensional data generation model information, the decoding device can appropriately decode the bitstream, which has improved encoding efficiency.
[0329] Furthermore, when the encoding device 1430 performs predictive coding with reference to two or more three-dimensional data generation models stored in the buffer 1432, it can also append information representing the two or more reference three-dimensional data generation models to the bitstream. Thus, the encoding device 1430 can use two or more reference three-dimensional data generation models to improve the coding efficiency of predictive coding. Moreover, the decoding device can appropriately decode the bitstream with improved coding efficiency.
[0330] Furthermore, when the reference 3D data generation model is not stored in the buffer 1432, for example, when encoding the initial 3D data generation model (initial frame) in data order, the encoding device 1430 may encode the 3D data generation model of the processing object without calculating the difference from the predicted value (hereinafter referred to as intra-frame prediction), or it may encode by calculating the difference from the predicted value set to 0. Additionally, when the encoding device 1430 sets a certain time t as a random access point, it can encode the 3D data generation model corresponding to time t through intra-frame prediction, or it may encode by calculating the difference from the predicted value set to 0. Therefore, the decoding device can start decoding the 3D data generation model from the initial 3D data generation model (initial frame) or the random access point in data order, improving the functionality during playback.
[0331] Furthermore, a set of multiple 3D data generation models (multiple frames) can be defined (hereinafter referred to as GOF (Group of Frame)). The first frame of the GOF can also be encoded through intra-frame prediction. Thus, the decoding device can randomly access the first frame of the GOF. In addition, by decoding the first frame of the GOF, functionality such as fast-forward playback can be improved.
[0332] Furthermore, the encoding device 1430 may also append permission information indicating whether inter-GOF prediction referencing is permitted to the bitstream. For example, if the bitstream contains permission information indicating that inter-GOF prediction referencing is prohibited, the decoding device can determine that multiple GOFs can be decoded in parallel. Additionally, for example, by permitting inter-GOF prediction referencing, encoding efficiency can be improved.
[0333] The network model encoding unit 1434 is the same as the network model encoding unit 1423 of the encoding device 1420, but it differs in that it encodes the difference information d0~d5 of the three-dimensional data generation model NNt0~NNt5 input from the difference calculation unit 1433 and outputs a bit stream.
[0334] Furthermore, the encoding device 1430 includes a difference calculation unit 1433 and a network model encoding unit 1434 separately, but it is not limited to this. For example, it may be configured such that the difference calculation unit 1433 is included within the network model encoding unit 1434. That is, the network model encoding unit 1434 may also perform the processing of the difference calculation unit 1433.
[0335] Furthermore, the encoding device 1430 may also append prediction coding information to the bitstream, indicating whether the 3D data generation model was encoded using intra-frame prediction or using a reference 3D data generation model for prediction coding (hereinafter referred to as inter-frame prediction). Thus, by decoding the prediction coding information, the decoding device can appropriately determine whether intra-frame prediction or inter-frame prediction should be used to decode the 3D data generation model.
[0336] Figure 41 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 1 of Implementation Method 2.
[0337] The decoding device 1435 includes a network model decoding unit 1436, an addition unit 1437, a buffer unit 1438, and a rendering unit 1439.
[0338] The network model decoding unit 1436 acquires the bit stream and, based on the acquired bit stream, decodes the difference information d0~d5 and other metadata such as time information of the three-dimensional data generation model NNt0~NNt5 at times t0~t5.
[0339] The addition unit 1437 adds the difference information d0~d5 of the three-dimensional data generation model corresponding to times t0~t5, which is decoded by the network model decoding unit 1436, and the reference three-dimensional data generation model RNNt0~RNNt5 obtained from the buffer unit 1438 at the corresponding times to calculate the three-dimensional data generation model NNt0~NNt5. In this way, the decoding device 1435 can also be set to RNNt0=0, RNNtn=NNt(n-1) (n is a value of 1~5), and use the three-dimensional data generation model of the previous time as the reference three-dimensional data generation model for prediction decoding.
[0340] Furthermore, in the second example, the decoding device 1435 separately describes the addition unit 1437 and the network model decoding unit 1436, but it is not limited to this. For example, it could also be a structure in which the addition unit 1437 is included within the network model decoding unit 1436. That is, the network model decoding unit 1436 can also perform the processing of the addition unit 1437.
[0341] Furthermore, if the buffer 1438 does not store a reference 3D data generation model, for example, when decoding the initial 3D data generation model (the first frame) in data order, the decoding device 1435 may perform decoding without adding the difference information to the reference 3D data generation model via the addition unit 1437 and without prediction (hereinafter referred to as intra-frame prediction), or it may add the prediction value set to 0 to the difference information for decoding. Additionally, if the decoding device 1435 sets a certain time t as a random access point, it can decode the 3D data generation model corresponding to time t using intra-frame prediction, or it may add the prediction value set to 0 to the difference information for decoding. Furthermore, if the bitstream contains prediction encoding information indicating that the 3D data generation model to be decoded has been encoded using intra-frame prediction, it can decode the 3D data generation model using intra-frame prediction, or it may add the prediction value set to 0 to the difference information for decoding. Therefore, the decoding device 1435 can begin decoding the 3D data generation model from the 3D data generation model that starts with the data sequence (starting frame), random access points, or 3D data generation models that have been encoded by intra-frame prediction, thereby improving the functionality during reproduction.
[0342] In addition, the decoding device 1435 in the second example performs predictive decoding on information related to the network model in the three-dimensional data generation model NNt5 based on information related to the network model in the three-dimensional data generation model NNt4, but is not limited to this. For example, the decoding device 1435 may select a reference three-dimensional data generation model for prediction from one or more three-dimensional data generation models stored in the buffer 1438, and perform predictive decoding using the selected three-dimensional data generation model. In this case, the decoding device 1435 may also decode the information representing the selected three-dimensional data generation model (reference three-dimensional data generation model information) from the bitstream. Thus, the decoding device 1435 decodes the reference three-dimensional data generation model information from the bitstream generated by the encoding device 1430, which selects the reference three-dimensional data generation model that is optimal from the viewpoint of encoding efficiency, thereby enabling appropriate decoding of the bitstream that improves encoding efficiency.
[0343] Furthermore, when performing predictive decoding with reference to two or more three-dimensional data generation models stored in the buffer 1438, the decoding device 1435 can also decode information representing two or more reference three-dimensional data generation models from the bitstream. Thus, the decoding device 1435 can appropriately decode the bitstream, which improves the coding efficiency of predictive coding, using two or more reference three-dimensional data generation models.
[0344] The rendering unit 1439 is the same as the rendering unit 1427 of the decoding device 1425. The rendering unit 1439 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time.
[0345] [Example 2] Figure 42 This diagram illustrates the motion image generation method using the extended three-dimensional data generation model of Embodiment 2 in Embodiment 2. Furthermore, this embodiment describes an example of the configuration and method for encoding or decoding extended three-dimensional data generation models NNt0-2 and NNt3-5, which are generated corresponding to periods t0-t2 and t3-t5 respectively, but it is not limited to this; it can also be applied to an apparatus and method for encoding or decoding extended three-dimensional data generation models in any period.
[0346] This embodiment illustrates a method for generating a moving image of an object (subject) observed from any viewpoint using a three-dimensional data generation model. In this method, for example, as... Figure 42In this way, by obtaining a three-dimensional data generation model (hereinafter referred to as the extended three-dimensional data generation model) capable of generating images from any viewpoint within a certain time range (period), it is possible to generate still images of objects observed from any viewpoint at any time within each period. By arranging the generated still images in chronological order, moving images can be generated. Similar to the three-dimensional data generation model in Embodiment 1, the extended three-dimensional data generation model is, for example, a three-dimensional data generation model generated by methods such as NeRF.
[0347] More specifically, when generating motion images from time t0 to t5, an extended 3D data generation model NNt0-2 capable of representing the period from t0 to t2 and an extended 3D data generation model NNt3-5 capable of representing the period from t3 to t5 are generated through learning. The viewpoint information (camera pose, etc.) of the viewpoint A from which the motion images are to be generated is input into the generated extended 3D data generation models NNt0-2 and NNt3-5. Thus, the generated images of viewpoint A from time t0 to t5 are output by the extended 3D data generation models NNt0-2 and NNt3-5. By concatenating these images temporally, motion images from time t0 to t5, in which the object is observed from viewpoint A, can be generated.
[0348] However, in this case, maintaining the extended 3D data generation model corresponding to each period (time period) requires either a large storage capacity to store the extended 3D data generation model data in a storage device or a large network bandwidth to transmit the data of multiple 3D data generation models over a network. Therefore, the data size can also be reduced by using, for example, NNC (Neural Network Coding) in the MPEG (Moving Picture Experts Group) standard to encode the extended 3D data generation model corresponding to each period. In this disclosure, a method for more effectively compressing this data is described.
[0349] Furthermore, based on the above configuration, the information processing device can generate any viewpoint image at any time within the period t0-t5. For example, when acquiring the extended 3D data generation model NNt0-2, the information processing device generates the extended 3D data generation model NNt0-2 by using multi-viewpoint images captured at times t0, t1, and t2 as learning data, and by machine learning based on the camera poses corresponding to the multi-viewpoints. Moreover, when generating a motion image of viewpoint A, the information processing device can generate not only viewpoint images A at times t0, t1, and t2, but also images of any viewpoint at times t0.5 and t1.5 between times t0, t1, and t2. Time t0.5 is the time between time t0 and time t1, and time t1.5 is the time between time t1 and time t2.
[0350] Therefore, not only the time corresponding to the image during learning, the information processing device can also generate an image of any viewpoint corresponding to a time offset from the time corresponding to the image during learning, thus enabling the generation of motion images of viewpoint A at a high frame rate.
[0351] Furthermore, as learning data for the extended 3D data generation model NNt0-2, the information processing device can learn not only the learning data at times t0, t1, and t2, but also, for example, the learning data at time t3. Thus, it is possible to generate viewpoint images from any viewpoint after time t2 with high precision, such as an image from any viewpoint at time t2.5.
[0352] Furthermore, as learning data for the extended 3D data generation model NNt3-5, the information processing device can learn not only the learning data corresponding to times t3, t4, and t5, but also, for example, supplement the learning data corresponding to times t2 and t6. Thus, the information processing device can generate images from any viewpoint before time t3 or from any viewpoint after time t5 with high precision. Moreover, as a switching point of the extended 3D data generation model, for example, in the above example, when generating a viewpoint image at time 2.5 between time t2 and t3, which is the switching point between the extended 3D data generation model NNt0-2 and the extended 3D data generation model NNt3-5, the information processing device can also generate viewpoint images at time t2.5 using both the extended 3D data generation model NNt0-2 and the extended 3D data generation model NNt3-5, and generate the average image of the two generated viewpoint images at time t2.5 as the viewpoint image at time t2.5. Thus, a high-precision viewpoint image at time t2.5 can be generated.
[0353] In this way, by specifying the time and viewpoint information within the period corresponding to the extended three-dimensional data generation model, the information processing device can generate an image of the object observed from the specified viewpoint at the specified time.
[0354] Figure 43 This is a diagram illustrating the first example of the configuration of the encoding device in Embodiment 2 of Implementation Method 2.
[0355] The encoding device 1450 includes an extended three-dimensional data generation model acquisition unit 1451, a buffer unit 1452, and a network model encoding unit 1453.
[0356] The extended 3D data generation model acquisition unit 1451 acquires learning data for each period t0~t2 and t3~t5, from time t0 to t5. Using the acquired learning data for each period, it generates an extended 3D data generation model NNt0-2 for period t0~t2 and an extended 3D data generation model NNt3-5 for period t3~t5 through learning. The learning data includes multiple viewpoint images of an object captured from one or more viewpoint positions along one or more viewing directions at each time t0~t5, and one or more viewpoint information representing one or more viewpoint positions and one or more viewing directions corresponding to the multiple viewpoint images. The one or more viewpoint information may be the position and pose of the camera when capturing each of the multiple viewpoint images. Furthermore, the learning data is not limited to this and may also include information obtained from other sensors. For example, the learning data may also include point cluster data and depth images acquired at each time using a LiDAR or TOF sensor. As a result, the accuracy of the extended 3D data generation model obtained through learning can be improved.
[0357] The buffer unit 1452 stores the extended three-dimensional data generation model for the period tm-n, from time tm (m is an integer) to time tn (n is an integer greater than m), generated by the extended three-dimensional data generation model acquisition unit 1451. The buffer unit 1452 is implemented using a storage device such as a memory. The extended three-dimensional data generation model for the period tm-n stored in the buffer unit 1452 can also be used as an initial model when the extended three-dimensional data generation model acquisition unit 1451 acquires (generates) the extended three-dimensional data generation model for the period after tm-n through learning. As a result, the learning time can be shortened and the accuracy of the extended three-dimensional data generation model for the period after tm-n can be improved.
[0358] Furthermore, the buffer unit 1452 can also store multiple extended 3D data generation models corresponding to multiple periods. Thus, for example, an initial model can be generated based on the multiple extended 3D data generation models stored in the buffer unit 1452, for example, through averaging or other processing. The extended 3D data generation model acquisition unit 1451 learns extended 3D data generation models for periods after period tm-n by using this initial model, and can obtain a high-precision extended 3D data generation model. Furthermore, if the extended 3D data generation model acquisition unit 1451 does not refer to extended 3D data generation models of past periods during learning, the encoding device 1450 may not need to include the buffer unit 1452. This reduces the amount of storage used as the buffer unit 1452.
[0359] The network model encoding unit 1453 encodes the extended three-dimensional data generation model NNt0-2 and the extended three-dimensional data generation model NNt3-5 obtained by the extended three-dimensional data generation model acquisition unit 1451 and outputs a bit stream.
[0360] Furthermore, as a network model encoding method, the data size can be reduced, for example, by using NNC data encoding from the MPEG standard. That is, the network model encoding unit 1453 uses NNC to encode the extended 3D data generation model NNt0-2 and the extended 3D data generation model NNt3-5, and appends the encoding result to the bitstream. In other words, the network model encoding unit 1453 generates encoded data as the encoding result and generates a bitstream containing the encoded data.
[0361] Specifically, the network model encoding unit 1453 first uses NNC to encode the extended 3D data generation model NNt0-2 for periods t0 to t2, and appends the encoding result to the bitstream. Next, the network model encoding unit 1453 uses NNC to encode the extended 3D data generation model NNt3-5 for periods t3 to t5, and appends the encoding result to the bitstream. In this way, the network model encoding unit 1453 can sequentially encode the extended 3D data generation model for each period using NNC and append each encoding result to the bitstream, thereby reducing the amount of encoding.
[0362] Furthermore, at this time, the network model encoding unit 1453 can also append time information, representing the period to which the encoded extended 3D data generation model corresponds, as metadata to the bitstream. Thus, by decoding and referring to the metadata contained in the bitstream, the decoding device can determine which period the decoded extended 3D data generation model corresponds to, and can appropriately generate motion images of the object from any viewpoint.
[0363] Furthermore, the network model encoding unit 1453 can also generate information as time information indicating which period of viewpoint image the extended 3D data generation model can generate, and append the generated time information as metadata to the bitstream. Thus, in the decoding device, by decoding this metadata, the period during which the extended 3D data generation model can generate viewpoint images can be determined, and motion images can be generated appropriately.
[0364] In addition, metadata is not limited to time information; it can also include information related to the acquisition (generation) of learning data, or information required by the decoding device to generate motion images.
[0365] For example, the network model encoding unit 1453 may also attach information related to the camera's frame rate when acquiring (generating) the learning data as metadata. Thus, the decoding device can decode the frame rate of the generated motion image from the bitstream and appropriately set that frame rate.
[0366] Alternatively, the network model encoding unit 1453 may append the frame number corresponding to each period as metadata to the bitstream instead of the time information, and use other parameters to associate each frame number with the time information. For example, the network model encoding unit 1453 may append the time information and frame rate of the first frame as metadata, and the decoding device may calculate the time information of each frame based on this metadata, thereby reducing the amount of encoding corresponding to the time information of each frame.
[0367] Furthermore, the network model encoding unit 1453 can also append viewpoint information from viewpoint images used in the learning process, or time information indicating the time when the viewpoint image was captured, to the bitstream. Thus, for example, the decoding device can generate high-quality motion images by preferentially selecting viewpoints close to the viewpoint positions corresponding to images used in the learning process, or times close to the times corresponding to images used in the learning process. This is because the closer the viewpoint position or time is to the time of learning, the more likely the extended 3D data generation model is to generate higher-quality viewpoint images.
[0368] Figure 44 This is a diagram illustrating the first example of the configuration of the decoding device in Embodiment 2 of Implementation Method 2.
[0369] The decoding device 1455 includes a network model decoding unit 1456 and a rendering unit 1457.
[0370] The network model decoding unit 1456 acquires the bit stream and, based on the acquired bit stream, decodes metadata such as the extended three-dimensional data generation model NNt0-2 during period t0~t2 and the extended three-dimensional data generation model NNt3-5 during period t3~t5, as well as the time information corresponding to these extended three-dimensional data generation models NNt0-2 and NNt3-5.
[0371] The rendering unit 1457 uses metadata such as the extended 3D data generation models NNt0-2 and NNt3-5 decoded by the network model decoding unit 1456 and time information to generate motion images of viewpoint A based on viewpoint information specified by the user or system. Specifically, the rendering unit 1457 inputs the viewpoint information of viewpoint A and the time within the period t0~t2 into the extended 3D data generation model NNt0-2, and generates images IMGt0 of viewpoint A at time t0, IMGt1 of viewpoint A at time t1, and IMGt2 of viewpoint A at time t2. The rendering unit 1457 applies the image generation processing of the period t0~t2 to the extended 3D data generation model NNt3-5 for the period t3~t5 to generate images IMGt3~IMGt5 of viewpoint A at times t3~t5. Furthermore, the rendering unit 1457 uses metadata such as images IMGt0~IMGt5 and time information to generate motion images of the object observed from viewpoint A at times t0~t5. The motion image may, for example, include images IMGt0~IMGt5 and prompt time information for calculating prompt times based on the images IMGt0~IMGt5 at times t0~t5.
[0372] Furthermore, the viewpoint information can change according to time. For example, viewpoint information of viewpoint A can be input into the extended 3D data generation model NNt0-2 during the period t0~t2, and viewpoint information of viewpoint B can be input into the extended 3D data generation model NNt3-5 during the period t3~t5. As a result, the rendering unit 1457 generates multiple images of the object observed from viewpoint A during time t0~t2, and generates multiple images of the object observed from viewpoint B during time t3~t5. That is, the rendering unit 1457 can generate motion images of the object observed, with the viewpoint switching from viewpoint A to viewpoint B at time t3.
[0373] Furthermore, the rendering unit 1457 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time. Thus, the user can switch between generating moving images and generating still images depending on the application.
[0374] Furthermore, the rendering unit 1457 is not limited to generating moving or still images based on the extended 3D data generation model. For example, the rendering unit 1457 can also generate point group data or mesh data that the extended 3D data generation model can represent, and output the generated point group data or mesh data as dynamic point group data or dynamic mesh data. Thus, users can use dynamic 3D data of audiovisual objects such as HMDs (Head Mount Displays), and can also use the dynamic 3D data to measure the amount of motion of objects.
[0375] Figure 45 This is a diagram illustrating a second example of the configuration of the encoding device in Embodiment 2 of Implementation Method 2.
[0376] The encoding device 1460 includes an extended three-dimensional data generation model acquisition unit 1461, a buffer unit 1462, a difference calculation unit 1463, and a network model encoding unit 1464.
[0377] The extended three-dimensional data generation model acquisition unit 1461 is the same as the extended three-dimensional data generation model acquisition unit 1451 of the encoding device 1450.
[0378] The buffer unit 1462 is the same as the buffer unit 1452 of the encoding device 1450, but it differs from the buffer unit 1452 in that it inputs the extended three-dimensional data generation model stored in the memory or the like as a reference extended three-dimensional data generation model into the difference calculation unit 1463.
[0379] The difference calculation unit 1463 calculates difference information, which represents the difference between the extended 3D data generation model NNt0-2 generated by the extended 3D data generation model acquisition unit 1461 for periods t0 to t2 and the extended 3D data generation model NNt3-5 for periods t3 to t5, respectively, and the extended 3D data generation model generated by the extended 3D data generation model acquisition unit 1461 before each period (hereinafter referred to as the reference extended 3D data generation model). Here, the difference information may include the difference in the weight parameters of the nodes of each network model, etc. For example, the difference calculation unit 1463 obtains the extended 3D data generation model NNt3-5 for periods t3 to t5 from the extended 3D data generation model acquisition unit 1461, and obtains the extended 3D data generation model NNt0-2 for periods t0 to t2 from the buffer unit 1462 as the reference extended 3D data generation model.
[0380] The difference calculation unit 1463 can also use extended 3D data generation models NNt3-5 and NNt0-2, for example, to calculate the difference (change) between the weight parameters of the nodes in the network model of extended 3D data generation model NNt3-5 and the weight parameters of the nodes in the network model of extended 3D data generation model NNt0-2, and input the difference information representing the difference to the network model encoding unit 1464. Thus, the difference information is encoded by the network model encoding unit 1464. That is, the encoding device 1460 can also reduce the amount of data by predictive encoding, which encodes the difference between the predicted and predicted values, based on the information related to the network model in extended 3D data generation model NNt0-2 predicted by extended 3D data generation model NNt3-5. Through such predictive encoding, for example, in cases where the changes in the extended 3D data generation model are small over time, such as when the object hardly moves, the value of the encoded difference becomes smaller, thus improving encoding efficiency. For example, the encoding device 1460 can also be set to RNNt0-2 = 0 and RNNt3-5 = NNt0-2, and the extended three-dimensional data generation model of the previous time period can be used as a reference extended three-dimensional data generation model to reduce the number of bits through predictive coding.
[0381] Furthermore, in the second example, the encoding device 1460 performs predictive encoding on information related to the network model in the extended 3D data generation model NNt3-5 based on information related to the network model in the extended 3D data generation model NNt0-2, but is not limited to this. For example, the encoding device 1460 may select a reference extended 3D data generation model for prediction from one or more extended 3D data generation models stored in the buffer 1462, and perform predictive encoding using the selected extended 3D data generation model. In this case, in order to pass the selected extended 3D data generation model to the decoding device, the encoding device 1460 may also append information representing the selected extended 3D data generation model (reference extended 3D data generation model information) to the bitstream. Thus, the encoding device 1460 can select the optimal reference extended 3D data generation model from the viewpoint of encoding efficiency, thereby improving encoding efficiency. Furthermore, by decoding the reference extended 3D data generation model information, the decoding device can appropriately decode the bitstream with improved encoding efficiency.
[0382] Furthermore, when the encoding device 1460 performs predictive coding with reference to two or more extended three-dimensional data generation models stored in the buffer 1462, it can also append information representing the two or more reference extended three-dimensional data generation models to the bitstream. Thus, the encoding device 1460 can use two or more reference extended three-dimensional data generation models to improve the coding efficiency of predictive coding. Furthermore, the decoding device can appropriately decode the bitstream with improved coding efficiency.
[0383] Furthermore, when the reference extended 3D data generation model is not stored in the buffer 1462, for example, when encoding the initial extended 3D data generation model (initial frame) in data order, the encoding device 1460 can encode the extended 3D data generation model of the processing object without calculating the difference from the predicted value (hereinafter referred to as intra-frame prediction), or it can calculate the difference from the predicted value set to 0 for encoding. Additionally, when a certain period tm-n is set as a random access point, the encoding device 1460 can encode the extended 3D data generation model corresponding to period tm-n through intra-frame prediction, or it can calculate the difference from the predicted value set to 0 for encoding. Therefore, the decoding device can decode the extended 3D data generation model starting from the initial extended 3D data generation model (initial frame) or the random access point in data order, improving functionality during playback.
[0384] Furthermore, a set of multiple extended 3D data generation models (multiple frames) can be defined (hereinafter referred to as GOF (Group of Frame)). The first frame of the GOF can also be encoded through intra-frame prediction. Thus, the decoding device can randomly access the first frame of the GOF. In addition, by decoding the first frame of the GOF, functionality such as fast-forward playback can be improved.
[0385] Furthermore, the encoding device 1460 may also append permission information indicating whether inter-GOF prediction referencing is permitted to the bitstream. For example, if the bitstream contains permission information indicating that inter-GOF prediction referencing is prohibited, the decoding device can determine that multiple GOFs can be decoded in parallel. Additionally, for example, by permitting inter-GOF prediction referencing, encoding efficiency can be improved.
[0386] The network model encoding unit 1464 is the same as the network model encoding unit 1453 of the encoding device 1450, but it differs in that it encodes the difference information d0-2 and d3-5 of the extended three-dimensional data generation models NNt0-2 and NNt3-5 input from the difference calculation unit 1463 and outputs a bit stream.
[0387] Furthermore, the encoding device 1460 includes a difference calculation unit 1463 and a network model encoding unit 1464 separately, but it is not limited to this. For example, it may be configured to include the difference calculation unit 1463 within the network model encoding unit 1464. That is, the network model encoding unit 1464 may also perform the processing of the difference calculation unit 1463.
[0388] Furthermore, the encoding device 1460 may also append prediction coding information to the bitstream, indicating whether the extended 3D data generation model was encoded using intra-frame prediction or using a reference extended 3D data generation model for prediction coding (hereinafter referred to as inter-frame prediction). Thus, by decoding the prediction coding information, the decoding device can appropriately determine whether intra-frame prediction or inter-frame prediction should be used to decode the extended 3D data generation model.
[0389] Figure 46 This is a diagram illustrating a second example of the configuration of the decoding device in Embodiment 2 of Implementation Method 2.
[0390] The decoding device 1465 includes a network model decoding unit 1466, an addition unit 1467, a buffer unit 1468, and a rendering unit 1469.
[0391] The network model decoding unit 1466 acquires the bit stream and, based on the acquired bit stream, decodes the difference information d0-2, d3-5, and time information of the extended three-dimensional data generation model NNt0-2 and the extended three-dimensional data generation model NNt3-5 during the period t0~t2.
[0392] The addition unit 1467 adds the difference information d0-2, d3-5 of the extended 3D data generation models NNt0-2 and NNt3-5 corresponding to periods t0~t2 and t3~t5, decoded by the network model decoding unit 1466, and the reference extended 3D data generation models RNNt0-2 and RNNt3-5 obtained from the buffer unit 1468 during the corresponding periods to calculate the extended 3D data generation models NNt0-2 and NNt3-5. Thus, the decoding device 1465 can also be set to RNNt0-2 = 0 and RNNt3-5 = NNt0-2, using the extended 3D data generation model from the previous time period as a reference extended 3D data generation model for prediction decoding.
[0393] Furthermore, in the second example, the decoding device 1465 separately describes the addition unit 1467 and the network model decoding unit 1466, but it is not limited to this. For example, it could be configured to include the addition unit 1467 within the network model decoding unit 1466. That is, the network model decoding unit 1466 could also perform the processing of the addition unit 1467.
[0394] Furthermore, if the buffer 1468 does not store a reference extended 3D data generation model, for example, when decoding the initial extended 3D data generation model (the first frame) in data order, the decoding device 1465 may perform decoding without adding the difference information to the reference extended 3D data generation model by the addition unit 1467 and without prediction (hereinafter referred to as intra-frame prediction), or it may add the prediction value set to 0 to the difference information to perform decoding. Additionally, if the decoding device 1465 sets a certain period tm-n as a random access point, it can decode the extended 3D data generation model corresponding to period tm-n through intra-frame prediction, or it may add the prediction value set to 0 to the difference information to perform decoding. Furthermore, if the bitstream contains prediction encoding information indicating that the extended 3D data generation model of the decoding target has been encoded through intra-frame prediction, it can decode the extended 3D data generation model through intra-frame prediction, or it may add the prediction value set to 0 to the difference information to perform decoding. Therefore, the decoding device 1465 can start decoding the extended three-dimensional data generation model from the extended three-dimensional data generation model (starting frame) that begins in data order, random access points, or extended three-dimensional data generation models that have been encoded by intra-frame prediction, thereby improving the functionality during reproduction.
[0395] In addition, the decoding device 1465 in the second example performs predictive decoding on information related to the network model in the extended 3D data generation model NNt3-5 based on information related to the network model in the extended 3D data generation model NNt0-2, but is not limited to this. For example, the decoding device 1465 may also select a reference extended 3D data generation model for prediction from one or more extended 3D data generation models stored in the buffer 1468, and perform predictive decoding using the selected extended 3D data generation model. In this case, the decoding device 1465 may also decode information representing the selected extended 3D data generation model (reference extended 3D data generation model information) from the bitstream. Thus, the decoding device 1465 decodes the reference extended 3D data generation model information from the bitstream generated by the encoding device 1460, which selects the reference extended 3D data generation model that is optimal from the viewpoint of encoding efficiency, thereby enabling appropriate decoding of the bitstream with improved encoding efficiency.
[0396] Furthermore, when performing predictive decoding with reference to two or more extended three-dimensional data generation models stored in the buffer 1468, the decoding device 1465 can also decode information representing two or more reference extended three-dimensional data generation models from the bitstream. Thus, the decoding device 1465 can appropriately decode the bitstream that improves the coding efficiency of predictive coding using two or more reference extended three-dimensional data generation models.
[0397] The rendering unit 1469 is the same as the rendering unit 1427 of the decoding device 1425. The rendering unit 1469 does not necessarily need to generate moving images; it can also generate still images with specified viewpoint information at a specified time.
[0398] [Variation Example] Additionally, the encoding device 1460 may include information related to the number of images that the extended 3D data generation model can generate (i.e., the upper limit of the number of images) in the metadata attached to the bitstream of the extended 3D data generation model during the period tm-n. Thus, the decoding device 1465 can know the number of images that the decoded extended 3D data generation model can generate, for example, by appropriately setting the frame rate of the motion picture to be generated, and also, for example, by calculating the number of delayed frames until the motion picture is displayed.
[0399] Additionally, the encoding device 1460 can also append information indicating the time unit (i.e., the smallest time unit) to which the extended 3D data generation model can generate viewpoint images as time information appended to the bitstream. For example, as time information, the encoding device 1460 can append information such as whether the viewpoint image can be generated up to a time unit of 1 msec or 1 μmsec to the bitstream. Thus, the decoding device 1465 can determine the time unit to which the viewpoint information is generated and can accordingly generate high frame rate motion images or 3D data.
[0400] Furthermore, the encoding device 1460 can also append information about the extended 3D data generation model to the metadata of the bitstream attached to the extended 3D data generation model during the period tm-n. For example, by appending the timing information or viewpoint information of the image used for learning as metadata to the bitstream, the encoding device 1460 can, by decoding the metadata, know the timing or viewpoint information at which the extended 3D data generation model can generate viewpoint images with high quality, thereby enabling the production of high-quality motion pictures.
[0401] Furthermore, the width of the viewpoint image generated by the extended 3D data generation model can also be... Figure 47 The switching can be done dynamically as shown. Specifically, the width of the period can also be switched according to the subject. Figure 47 This is a diagram illustrating a motion image generation method using an extended three-dimensional data generation model, which is used to explain a variation of Embodiment 2.
[0402] For example, in scenes with many stationary objects in the subject (scenes where the number of stationary objects in multiple subjects is a first number or more, or scenes where the volume (area) occupied by the stationary objects in multiple subjects is a first amount or more), the encoding device 1460 can generate an extended 3D data generation model with a long duration capable of generating viewpoint images with high image quality by expanding the width of the learning data used during learning (i.e., extending the duration). For example, in scenes with many moving objects in the subject (scenes where the number of moving objects in multiple subjects is a first number or more, or scenes where the volume (area) occupied by the moving objects in multiple subjects is a first amount or more), the encoding device 1460 can generate an extended 3D data generation model with a narrow duration capable of generating viewpoint images with high image quality, even for moving objects, by narrowing the width of the learning data used during learning.
[0403] Alternatively, the encoding device 1460 can use learning data for a certain period tm-n (e.g., a Group of Frames (GOF) representing a set of frames within period tm-n in a learning image) to generate an extended 3D data generation model NNtm-n for that period tm-n. In this case, the encoding device 1460 buffers the learning image frames for period tm-n to generate the extended 3D data generation model and performs compressed transmission, thus generating a transmission delay of the GOF size. The encoding device 1460 can also append information related to this transmission delay, such as the number of GOF frames and the number of delayed frames, to the bitstream. Therefore, the decoding device 1465 can obtain the delay information by decoding the bitstream and can appropriately reproduce the motion image or 3D data taking the delay into account.
[0404] Furthermore, the above embodiments illustrate an example of generating a still image of an arbitrary viewpoint at a given time or period using a 3D data generation model or an extended 3D data generation model, but are not necessarily limited to this. For example, other 3D data generation models, such as... Figure 48 As shown, it generates (outputs) 3D data such as point cluster data or mesh data at a specific moment within a certain period. This allows users to perform dimensional measurements of objects or obtain 3D data with higher audiovisual detail. Figure 48 This is a diagram used to illustrate a motion image generation method based on a modified example of embodiment two involving a three-dimensional data generation model.
[0405] Furthermore, encoding devices 1420 and 1460 can also include recommended output formats corresponding to the use case in the metadata of the bitstream, indicating the output format of images, point group data, grid data, etc. Thus, the user can select the recommended output format based on the use case.
[0406] Furthermore, encoding devices 1420 and 1460 can also attach more than one viewpoint information to the metadata of the bitstream appended to the 3D data generation model or the extended 3D data generation model. For example, encoding devices 1420 and 1460 can consider including recommended viewpoint information for audiovisual objects or user viewpoint information when acquiring learning data in the metadata. Thus, decoding devices 1425 and 1465 can generate motion graphics or 3D data using viewpoint information selected from more than one viewpoint information appended to the bitstream based on user intent, etc.
[0407] Alternatively, a default viewpoint can be predetermined based on more than one viewpoint. Alternatively, if no user-specified viewpoint is provided, the decoding devices 1425 and 1465 can use the predetermined default viewpoint to generate motion images or 3D data. Thus, the decoding devices 1425 and 1465 can automatically generate motion images or 3D data even without user specification.
[0408] As an example of using this implementation method, there are the following usage methods.
[0409] First, the encoding devices 1420 and 1460 use cameras or sensors to acquire data of a dynamic object that they want to send to a distance, and use the data of the dynamic object as learning data to generate a three-dimensional data generation model or an extended three-dimensional data generation model of the dynamic object.
[0410] Next, the encoding devices 1420 and 1460 encode the three-dimensional data generation model or the extended three-dimensional data generation model using the encoding method described in this embodiment, and transmit the bit stream containing the encoding result to a remote location.
[0411] Then, decoding devices 1425 and 1465 decode the bitstream received from a distance, and use the decoded 3D data of the dynamic object to generate a 3D model or an extended 3D data generation model to generate motion images or 3D data from any viewpoint. The generated 3D data can then be used for appreciation or measurement purposes. In this way, this embodiment can also be applied to all use cases of remotely sharing information in a certain space.
[0412] Furthermore, when there are more than one object in a certain space that needs to be sent to a distant location, the 3D data generation modeling, encoding and transmission, decoding, and rendering processes described in this embodiment can be applied to each object separately. For example, dynamic objects in the foreground and static objects in the background existing in a certain space can be generated and modeled in 3D data and encoded and transmitted separately. As a result, the optimal 3D data generation modeling or encoding method can be applied to each object, thereby improving encoding efficiency.
[0413] Furthermore, it is not necessarily limited to this; multiple objects can also be treated as a single object, and the 3D data generation, modeling, encoding, transmission, decoding, and rendering processes described in this embodiment can be applied to each object separately. This allows for the transmission of multiple objects to a remote location while minimizing processing overhead.
[0414] Figure 49 This is a diagram illustrating an example of the configuration of the encoding device in Embodiment 2. Figure 50 This is a flowchart illustrating an example of the encoding method of the encoding device in Embodiment 2.
[0415] The encoding device 1470 includes a circuit 1471 and a memory 1472. The encoding device 1470 is a device that implements the encoding devices 1420 and 1460.
[0416] Circuit 1471 performs the following actions.
[0417] Circuit 1471 acquires a first three-dimensional data generation model (e.g., three-dimensional data generation model NNt0) corresponding to a first time point (e.g., time t0) and a second three-dimensional data generation model (e.g., three-dimensional data generation model NNt1) corresponding to a second time point (e.g., time t1) (S1401). Circuit 1471 generates a bitstream by encoding the acquired first and second three-dimensional data generation models (S1402). The first and second three-dimensional data generation models output two-dimensional images of the subject as viewed from the viewpoint and the viewing direction, respectively, when input with viewpoint information including the viewpoint and the viewing direction.
[0418] Therefore, it is possible to generate a bitstream containing a first three-dimensional data generation model that generates a two-dimensional image corresponding to a first moment based on arbitrary viewpoint information and a second three-dimensional data generation model that generates a two-dimensional image corresponding to a second moment. Thus, it is possible to generate a bitstream that compresses the data of the motion image obtained from an arbitrary viewpoint. Therefore, it is possible to reduce the storage capacity used to store the data of the motion image obtained from an arbitrary viewpoint, or the network bandwidth used to transmit the data.
[0419] For example, the first 3D data generation model and the second 3D data generation model are learning models that use neural networks.
[0420] For example, the bitstream includes first time information representing the first time moment and second time information representing the second time moment.
[0421] For example, the bitstream includes a first frame number corresponding to the first time point and a second frame number corresponding to the second time point.
[0422] For example, the bitstream contains frame rate information related to the frame rate of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model. The multiple learning images are 2D images obtained by capturing images at multiple different timings.
[0423] For example, the bitstream contains viewpoint information, which includes the viewpoint and line-of-sight direction of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model.
[0424] For example, the plurality of learning images are two-dimensional images obtained by photographing the subject from mutually different viewpoints and viewing directions. The viewpoint information includes the mutually different viewpoints and viewing directions.
[0425] For example, circuit 1471 calculates difference information representing the difference between the first three-dimensional data generation model and the second three-dimensional data generation model in the encoding of the second three-dimensional data generation model. The bitstream contains the difference information.
[0426] For example, the difference includes the difference between the weight parameters corresponding to the nodes contained in the first 3D data generation model and the second 3D data generation model.
[0427] For example, the bitstream includes reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.
[0428] For example, the first time point corresponds to a random access point. The first 3D data generation model is encoded either by intra-frame prediction or by inter-frame prediction with a prediction value of 0.
[0429] For example, the first 3D data generation model and the second 3D data generation model are contained in one of a plurality of sets. The first 3D data generation model is the first in the data order among the plurality of 3D data generation models contained in the set.
[0430] For example, the bitstream contains licensing information indicating whether the 3D data generation model is permitted to reference other 3D data generation models in the respective encodings of the plurality of 3D data generation models.
[0431] For example, the first three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt0-2) corresponds to a first period (e.g., period t0~t2) that includes the first time point (e.g., time t0). The second three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt3-5) corresponds to a second period (e.g., period t3~t5) that includes the second time point (e.g., time t3).
[0432] For example, the multiple first learning images used in the generation of the first three-dimensional data generation model are two-dimensional images obtained by taking pictures at multiple different timings during the first period.
[0433] For example, when the first three-dimensional data generation model is input with a time contained in the first period, it outputs a two-dimensional image of the subject at the input time.
[0434] For example, the bitstream contains information indicating the upper limit of the number of images that the first 3D data generation model can generate.
[0435] For example, the bitstream contains first information associated with the plurality of first learning images. The first information includes multiple viewpoints and multiple line-of-sight directions corresponding to the plurality of first learning images, as well as multiple different timings.
[0436] For example, the first period or the second period is dynamically determined based on the subject.
[0437] For example, circuit 1471 saves the generated first three-dimensional data generation model in memory 1472. Based on the first three-dimensional data generation model saved in memory 1472, circuit 1471 generates the second three-dimensional data generation model.
[0438] For example, circuit 1471 stores the generated first 3D data generation model and the second 3D data generation model in memory 1472. Based on the first 3D data generation model and the second 3D data generation model stored in memory 1472, circuit 1471 generates an initial model. Based on the initial model, circuit 1471 generates a third 3D data generation model (e.g., 3D data generation model NNt2) corresponding to a third time point (e.g., time t2).
[0439] Figure 51 This is a diagram illustrating an example of the configuration of the decoding device in Embodiment 2. Figure 52 This is a flowchart illustrating an example of a decoding method using the decoding device in Embodiment 2.
[0440] The decoding device 1480 includes circuitry 1481 and memory 1482. The decoding device 1480 is a device that implements the decoding devices 1425 and 1465.
[0441] Circuit 1481 performs the following actions.
[0442] Circuit 1481 acquires the bitstream (S1411). Circuit 1481 decodes from the bitstream a first three-dimensional data generation model (e.g., three-dimensional data generation model NNt0) corresponding to a first time moment (e.g., time t0) and a second three-dimensional data generation model (e.g., three-dimensional data generation model NNt1) corresponding to a second time moment (e.g., time t1) (S1412). When the first three-dimensional data generation model and the second three-dimensional data generation model are input with viewpoint information including the viewpoint and the viewing direction, they output a two-dimensional image of the subject as viewed from the viewpoint and the viewing direction, respectively.
[0443] Therefore, based on the compressed bitstream of motion image data obtained from any viewpoint, it is possible to decode a first three-dimensional data generation model that generates a two-dimensional image corresponding to a first moment based on arbitrary viewpoint information, and a second three-dimensional data generation model that generates a two-dimensional image corresponding to a second moment. Thus, it is possible to appropriately decode the bitstream that reduces the storage capacity for storing the motion image data obtained from any viewpoint or the network bandwidth for transmitting the data.
[0444] For example, the first 3D data generation model and the second 3D data generation model are learning models that use neural networks.
[0445] For example, the bitstream includes first time information representing the first time moment and second time information representing the second time moment.
[0446] For example, the bitstream includes a first frame number corresponding to the first time point and a second frame number corresponding to the second time point.
[0447] For example, the bitstream contains frame rate information related to the frame rate of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model. The multiple learning images are 2D images obtained by capturing images at multiple different timings.
[0448] For example, the bitstream contains viewpoint information, which includes the viewpoint and line-of-sight direction of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model.
[0449] For example, the plurality of learning images are two-dimensional images obtained by photographing the subject from mutually different viewpoints and viewing directions. The viewpoint information includes the mutually different viewpoints and viewing directions.
[0450] For example, the bitstream contains difference information representing the difference between the first 3D data generation model and the second 3D data generation model.
[0451] For example, the difference includes the difference between the weight parameters corresponding to the nodes contained in the first 3D data generation model and the second 3D data generation model.
[0452] For example, the bitstream includes reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.
[0453] For example, the first time point corresponds to a random access point. The first 3D data generation model is encoded either by intra-frame prediction or by inter-frame prediction with a prediction value of 0.
[0454] For example, the first 3D data generation model and the second 3D data generation model are contained in one of a plurality of sets. The first 3D data generation model is the first in the data order among the plurality of 3D data generation models contained in the set.
[0455] For example, the bitstream contains licensing information indicating whether the 3D data generation model is permitted to reference other 3D data generation models in the respective encodings of the plurality of 3D data generation models.
[0456] For example, the first three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt0-2) corresponds to a first period (e.g., period t0~t2) that includes the first time point (e.g., time t0). The second three-dimensional data generation model (e.g., extended three-dimensional data generation model NNt3-5) corresponds to a second period (e.g., period t3~t5) that includes the second time point (e.g., time t3).
[0457] For example, the multiple first learning images used in the generation of the first three-dimensional data generation model are two-dimensional images obtained by taking pictures at multiple different timings during the first period.
[0458] For example, when the first three-dimensional data generation model is input with a time period included in the first period, it outputs a two-dimensional image of the subject at the input time.
[0459] For example, the bitstream contains information indicating the upper limit of the number of images that the first 3D data generation model can generate.
[0460] For example, the bitstream contains first information associated with the plurality of first learning images. The first information includes multiple viewpoints and multiple line-of-sight directions corresponding to the plurality of first learning images, as well as multiple different timings.
[0461] For example, the first period or the second period is dynamically determined based on the subject.
[0462] For example, circuit 1471 saves the generated first three-dimensional data generation model in memory 1472. Based on the first three-dimensional data generation model saved in memory 1472, circuit 1471 generates the second three-dimensional data generation model.
[0463] For example, circuit 1471 stores the generated first 3D data generation model and the second 3D data generation model in memory 1472. Based on the first 3D data generation model and the second 3D data generation model stored in memory 1472, circuit 1471 generates an initial model. Based on the initial model, circuit 1471 generates a third 3D data generation model (e.g., 3D data generation model NNt2) corresponding to a third time point (e.g., time t2).
[0464] (other) In one embodiment, a method for generating motion images from a predetermined viewpoint is disclosed. The generation of motion images is achieved, for example, by a device including a memory and circuitry connected to the memory. In one example of this device, a three-dimensional data generation model (Neural Network) generated through learning is stored in the memory, and the circuitry retrieves the stored three-dimensional data generation model (Neural Network) and generates motion images based on the three-dimensional data generation model. Alternatively, the three-dimensional data generation model or an extended three-dimensional data generation model may not be stored in memory. For example, encoding devices 1420 and 1460 may retrieve specified information from a URL on a specified network and obtain the three-dimensional data generation model based on that specified information.
[0465] Figure 53 This is a diagram illustrating an example of the configuration of an encoding device.
[0466] The encoding device 1490 includes a processor 1491 and a memory 1492.
[0467] Processor 1491 is a circuit that performs information processing and is capable of accessing memory 1492. For example, processor 1491 is a dedicated or general-purpose electronic circuit that encodes a 3D data generation model. Processor 1491 can also be a processor like a CPU. Alternatively, processor 1491 can be an assembly of multiple electronic circuits. Furthermore, for example, processor 1491 can also function as multiple components of the aforementioned encoding device, excluding the component for storing information.
[0468] Memory 1492 is a dedicated or general-purpose memory that stores information used by processor 1491 to encode the 3D data generation model. Memory 1492 can be an electronic circuit or connected to processor 1491. Alternatively, memory 1492 can be contained within processor 1491. Alternatively, memory 1492 can be an assembly of multiple electronic circuits. Alternatively, memory 1492 can be a disk or optical disk, or it can be a storage device or recording medium. Alternatively, memory 1492 can be non-volatile memory or volatile memory.
[0469] For example, memory 1492 may store the encoded 3D data generation model, or it may store the stream corresponding to the encoded 3D data generation model. Additionally, memory 1492 may also store the program used by processor 1491 to encode the 3D data generation model.
[0470] Furthermore, in the encoding device 1490, it is possible to omit all of the aforementioned components of the encoding device, and it is also possible to omit all of the aforementioned processes. A portion of the components may be included in other devices, and a portion of the aforementioned processes may be performed by other devices.
[0471] Figure 54 This is a diagram illustrating an example of the configuration of a decoding device.
[0472] The decoding device 1495 includes a processor 1496 and a memory 1497.
[0473] Processor 1496 is a circuit that performs information processing and is capable of accessing memory 1497. For example, processor 1496 is a dedicated or general-purpose electronic circuit for decoding streams. Processor 1496 can also be a processor like a CPU. Alternatively, processor 1496 can be an assembly of multiple electronic circuits. Furthermore, for example, processor 1496 can also function as multiple components of the aforementioned decoding device, excluding the component for storing information.
[0474] Memory 1497 is a dedicated or general-purpose memory that stores information used by processor 1496 to decode the stream. Memory 1497 can be an electronic circuit or connected to processor 1496. Alternatively, memory 1497 can be contained within processor 1496. Alternatively, memory 1497 can be an assembly of multiple electronic circuits. Alternatively, memory 1497 can be a magnetic disk or optical disk, or it can be a storage device or recording medium. Alternatively, memory 1497 can be non-volatile memory or volatile memory.
[0475] For example, the memory 1497 can store a 3D data generation model or a stream. Furthermore, the memory 1497 can also store a program for the processor 1496 to decode the stream.
[0476] Furthermore, in the decoding device 1495, it is possible to omit all of the aforementioned components of the decoding device, and also to omit all of the aforementioned processes. A portion of the components may be included in other devices, and a portion of the aforementioned processes may be performed by other devices.
[0477] Industrial applicability This disclosure can be applied to encoding devices and the like that can output three-dimensional data at different resolutions.
[0478] Explanation of reference numerals in the attached figures 1001 Three-Dimensional Data Encoding System 1002 3D Data Decoding System 1003 Sensor Terminal 1004 External Connection Part 1011 3D Data Generation System 1012 Reminder Department 1013 Coding Department 1014 Reuse Department 1015 Input / Output Section 1016 Control Department 1017 Sensor Information Acquisition Department 1018 3D Data Generation Department 1021 Sensor Information Acquisition Department 1022 Input / Output Section 1023 Demultiplexing Department 1024 Decoding Department 1025 Reminder Department 1026 User Interface 1027 Control Department 1031 3D Model Learning Department 1032 3D Model Coding Department 1033 3D Model Decoding Department 1034 Rendering and Reconstruction Department 1041 Data Segmentation Unit 1042 Coding Department 1051 Decoding Department 1052 Data Integration Section 1070 server 1071 Data Generation Department 1072-point group generation department 1073 Mesh Generation Department 1074 Model Generation Department 1075 Synchronization Unit 1076-point group coding department 1077 Mesh Coding Department 1078 Model Coding Department 1079 Reuse Department 1080 Data Extraction Department 1090 terminal 1091 Control Department 1092 Decoding Department 1093 Reminder Department 1101 point group sensor 1102 camera 1110 Data Generation Department 1111 dot group generation department 1112 Mesh Generation Department Model Generation Department 1113 1130 Decoding Device 1131 circuit 1132 memory 1140 encoding device 1141 circuit 1142 memory 1401 3D Data Generation Model 1402 Evaluation Function 1403 Three-Dimensional Data Generation Model 1420 encoding device 1421 Three-dimensional data generation model acquisition department 1422 Buffer Section 1423 Network Model Coding Department 1425 Decoding Device 1426 Network Model Decoding Department Rendering Department 1427 1430 Encoding Device 1431 Three-dimensional data generation model acquisition department 1432 Buffer Section 1433 Differential Calculation Department 1434 Network Model Coding Department 1435 Decoding Device 1436 Network Model Decoding Department 1437 Addition Department 1438 Buffer Section 1439 Rendering Department 1450 encoding device 1451 Extended 3D Data Generation Model Acquisition Department 1452 Buffer Section 1453 Network Model Coding Department 1455 Decoding Device 1456 Network Model Decoding Department 1457 Rendering Department 1460 encoding device 1461 Extended 3D Data Generation Model Acquisition Department 1462 Buffer Section 1463 Differential Calculation Department 1464 Network Model Coding Department 1465 decoding device 1466 Network Model Decoding Department 1467 Addition Department 1468 Buffer Section 1469 Rendering Department 1470 encoding device 1471 circuit 1472 memory 1480 decoding device 1481 circuit 1482 memory 1490 encoding device 1491 processor 1492 memory 1495 Decoding Device 1496 processor 1497 memory
Claims
1. An encoding device, characterized in that, have: Circuits; and The memory is connected to the circuit. The circuit, during operation, Obtain the first three-dimensional data generation model corresponding to the first time step and the second three-dimensional data generation model corresponding to the second time step. A bitstream is generated by encoding the obtained first three-dimensional data generation model and the second three-dimensional data generation model. The first 3D data generation model and the second 3D data generation model output a 2D image of the subject as viewed from the viewpoint and the viewing direction when they are input with viewpoint information including the viewpoint and the viewing direction.
2. The encoding device according to claim 1, characterized in that, The first 3D data generation model and the second 3D data generation model are both learning models that use neural networks.
3. The encoding device according to claim 1, characterized in that, The bitstream contains first time information representing the first time and second time information representing the second time.
4. The encoding device according to claim 3, characterized in that, The bitstream includes a first frame number corresponding to the first time point and a second frame number corresponding to the second time point.
5. The encoding device according to any one of claims 1 to 4, characterized in that, The bitstream contains frame rate information related to the frame rate of the multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model. The multiple learning images are two-dimensional images obtained by taking pictures at multiple different time points.
6. The encoding device according to any one of claims 1 to 4, characterized in that, The bitstream contains viewpoint information, which includes the viewpoint and line-of-sight direction of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model.
7. The encoding device according to claim 6, characterized in that, The multiple learning images are two-dimensional images obtained by photographing the subject from different viewpoints and line-of-sight directions. The viewpoint information includes the different viewpoints and line-of-sight directions.
8. The encoding device according to any one of claims 1 to 4, characterized in that, In the encoding of the second three-dimensional data generation model, the circuit calculates the difference information representing the difference between the first three-dimensional data generation model and the second three-dimensional data generation model. The bitstream contains the differential information.
9. The encoding device according to claim 8, characterized in that, The difference includes the difference between the weight parameters corresponding to the nodes contained in the first 3D data generation model and the second 3D data generation model.
10. The encoding device according to claim 8, characterized in that, The bitstream contains reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.
11. The encoding device according to any one of claims 1 to 4, characterized in that, The first moment corresponds to a random access point. The first 3D data generation model is encoded either by intra-frame prediction or by inter-frame prediction with a prediction value of 0.
12. The encoding device according to claim 11, characterized in that, The first 3D data generation model and the second 3D data generation model are contained in one of a plurality of sets. The first three-dimensional data generation model is the first in the data order among the multiple three-dimensional data generation models contained in the set.
13. The encoding device according to claim 12, characterized in that, The bitstream contains licensing information indicating whether the 3D data generation model is permitted to reference other 3D data generation models in the encoding of each of the plurality of 3D data generation models.
14. The encoding device according to any one of claims 1 to 4, characterized in that, The first three-dimensional data generation model corresponds to the first period including the first moment. The second three-dimensional data generation model corresponds to the second period that includes the second time point.
15. The encoding device according to claim 14, characterized in that, The multiple first learning images used in the generation of the first three-dimensional data generation model are two-dimensional images obtained by taking pictures at multiple different timings during the first period.
16. The encoding device according to claim 14, characterized in that, When the first three-dimensional data generation model is input with respect to the time included in the first period, it outputs a two-dimensional image of the subject at the input time.
17. The encoding device according to claim 14, characterized in that, The bitstream contains information indicating the upper limit of the number of images that the first 3D data generation model can generate.
18. The encoding device according to claim 15, characterized in that, The bitstream contains first information related to the plurality of first learning images. The first information includes multiple viewpoints and multiple line-of-sight directions corresponding to the multiple first learning images, as well as multiple different timings.
19. The encoding device according to claim 14, characterized in that, The first period or the second period is dynamically determined based on the subject.
20. The encoding device according to any one of claims 1 to 4, characterized in that, The circuit saves the generated first three-dimensional data model in the memory. The circuit generates a second three-dimensional data generation model based on the first three-dimensional data generation model stored in the memory.
21. The encoding device according to any one of claims 1 to 4, characterized in that, The circuit saves the generated first 3D data generation model and the second 3D data generation model in the memory. The circuit generates an initial model based on the first three-dimensional data generation model and the second three-dimensional data generation model stored in the memory. The circuit generates a third three-dimensional data generation model corresponding to the third time step based on the initial model.
22. A decoding device, characterized in that, have: Circuits; and The memory is connected to the circuit. The circuit, during operation, Obtain the bit stream. The first three-dimensional data generation model corresponding to the first time step and the second three-dimensional data generation model corresponding to the second time step are derived from the bitstream decoding. The first 3D data generation model and the second 3D data generation model output a 2D image of the subject as viewed from the viewpoint and the viewing direction when they are input with viewpoint information including the viewpoint and the viewing direction.
23. The decoding device according to claim 22, characterized in that, The first 3D data generation model and the second 3D data generation model are both learning models that use neural networks.
24. The decoding device according to claim 22, characterized in that, The bitstream contains first time information representing the first time and second time information representing the second time.
25. The decoding apparatus according to claim 24, characterized in that, The bitstream includes a first frame number corresponding to the first time point and a second frame number corresponding to the second time point.
26. The decoding apparatus according to any one of claims 22 to 25, characterized in that, The bitstream contains frame rate information related to the frame rate of the multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model. The multiple learning images are two-dimensional images obtained by taking pictures at multiple different time points.
27. The decoding apparatus according to any one of claims 22 to 25, characterized in that, The bitstream contains viewpoint information, which includes the viewpoint and line-of-sight direction of multiple learning images used in the generation of the first 3D data generation model and the second 3D data generation model.
28. The decoding apparatus according to claim 27, characterized in that, The multiple learning images are two-dimensional images obtained by photographing the subject from different viewpoints and line-of-sight directions. The viewpoint information includes the different viewpoints and line-of-sight directions.
29. The decoding apparatus according to any one of claims 22 to 25, characterized in that, The bitstream contains difference information representing the difference between the first three-dimensional data generation model and the second three-dimensional data generation model.
30. The decoding apparatus according to claim 29, characterized in that, The difference includes the difference between the weight parameters corresponding to the nodes contained in the first 3D data generation model and the second 3D data generation model.
31. The decoding device according to claim 29, characterized in that, The bitstream contains reference target information, which indicates that the difference information is calculated with reference to the first three-dimensional data generation model.
32. The decoding apparatus according to any one of claims 22 to 25, characterized in that, The first moment corresponds to a random access point. The first 3D data generation model is decoded either by intra-frame prediction or by inter-frame prediction with a prediction value of 0.
33. The decoding device according to claim 32, characterized in that, The first 3D data generation model and the second 3D data generation model are contained in one of a plurality of sets. The first three-dimensional data generation model is the first in the data order among the multiple three-dimensional data generation models contained in the set.
34. The decoding apparatus according to claim 33, characterized in that, The bitstream contains permission information indicating whether the 3D data generation model is permitted to reference other 3D data generation models in the decoding of each of the plurality of 3D data generation models.
35. The decoding apparatus according to any one of claims 22 to 25, characterized in that, The first three-dimensional data generation model corresponds to the first period including the first moment. The second three-dimensional data generation model corresponds to the second period that includes the second time point.
36. The decoding apparatus according to claim 35, characterized in that, The multiple first learning images used in the generation of the first three-dimensional data generation model are two-dimensional images obtained by taking pictures at multiple different timings during the first period.
37. The decoding apparatus according to claim 35, characterized in that, When the first three-dimensional data generation model is input with respect to the time included in the first period, it outputs a two-dimensional image of the subject at the input time.
38. The decoding apparatus according to claim 35, characterized in that, The bitstream contains information indicating the upper limit of the number of images that the first 3D data generation model can generate.
39. The decoding apparatus according to claim 36, characterized in that, The bitstream contains first information related to the plurality of first learning images. The first information includes multiple viewpoints and multiple line-of-sight directions corresponding to the multiple first learning images, as well as multiple different timings.
40. The decoding apparatus according to claim 35, characterized in that, The first period or the second period is dynamically determined based on the subject.
41. The decoding apparatus according to any one of claims 22 to 25, characterized in that, The circuit saves the generated first three-dimensional data model in the memory. The circuit generates a second three-dimensional data generation model based on the first three-dimensional data generation model stored in the memory.
42. The decoding apparatus according to any one of claims 22 to 25, characterized in that, The circuit saves the generated first 3D data generation model and the second 3D data generation model in the memory. The circuit generates an initial model based on the first three-dimensional data generation model and the second three-dimensional data generation model stored in the memory. The circuit generates a third three-dimensional data generation model corresponding to the third time step based on the initial model.
43. An encoding method, characterized in that, Obtain the first three-dimensional data generation model corresponding to the first time step and the second three-dimensional data generation model corresponding to the second time step. A bitstream is generated by encoding the first three-dimensional data generation model and the second three-dimensional data generation model. When the first 3D data generation model and the second 3D data generation model are input with viewpoint information including viewpoint and viewing direction, they output a 2D image of the subject as viewed from the viewpoint and the viewing direction.
44. A decoding method, characterized in that, Obtain the bit stream. From the bitstream, the first three-dimensional data generation model corresponding to the first time step and the second three-dimensional data generation model corresponding to the second time step are decoded. When the first 3D data generation model and the second 3D data generation model are input with viewpoint information including viewpoint and viewing direction, they output a 2D image of the subject as viewed from the viewpoint and the viewing direction.
Citation Information
Patent Citations
Map display device
WO2014020663A1