Encoding device, decoding device, encoding method, and decoding method

The encoding and decoding devices efficiently arrange and process features with a time dimension on images to enhance coding efficiency, improve image quality, and reduce circuit size, addressing the challenges of existing three-dimensional data encoding methods.

WO2025204846A1PCT designated stage Publication Date: 2025-10-02PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/008985
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2025-03-11
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing methods for encoding three-dimensional data face challenges in improving encoding efficiency, image quality, reducing processing volume, and circuit size, as well as selecting appropriate elements or operations such as filters, block sizes, motion vectors, and reference pictures.

Method used

The proposed solution involves an encoding device and decoding device that utilize a circuit and memory to acquire and encode/decode information indicating a plurality of features with a time dimension, arranged on images to generate a bitstream, allowing for improved coding efficiency, reduced processing volume, and efficient data arrangement.

Benefits of technology

This approach enhances coding efficiency, improves image quality, reduces circuit size, and optimizes processing speed by effectively arranging and decoding features with a time dimension, facilitating efficient encoding and decoding processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025008985_02102025_PF_FP_ABST
    Figure JP2025008985_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An encoding device according to the present invention comprises a circuit and a memory that is connected to the circuit, wherein, in operation, the circuit acquires first information that is included in a model and that indicates a plurality of feature amounts having a time dimension (S301), disposes the first information on one or more images (S302), and generates a bitstream by encoding the one or more images (S303). For example, the first information may include a plurality of pieces of data, and the bitstream may include second information that indicates the order of disposition of the plurality of data.
Need to check novelty before this filing date? Find Prior Art

Description

Encoding device, decoding device, encoding method, and decoding method

[0001] The present disclosure relates to an encoding device, a decoding device, an encoding method, and a decoding method.

[0002] Various techniques are known for encoding three-dimensional data. For example, there is a method for encoding a model representing three-dimensional data. In addition, a technique for expressing a dynamic model is known (see, for example, Non-Patent Document 1).

[0003] “K-Planes: Explicit Radiance Fields in Space, Time, and Appearance”, https: / / arxiv.org / pdf / 2301.10241

[0004] With regard to the above-mentioned encoding methods, it is desirable to propose new methods to improve encoding efficiency, improve image quality, reduce the amount of processing, reduce the circuit scale, or appropriately select elements or operations such as filters, block sizes, motion vectors, reference pictures or reference blocks.

[0005] The present disclosure provides a configuration or method that can contribute to one or more of, for example, improved coding efficiency, improved image quality, reduced processing amount, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, etc. Note that the present disclosure may include a configuration or method that can contribute to benefits other than those described above.

[0006] For example, an encoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit, and in operation, the circuit acquires first information included in a model, the first information indicating a plurality of features having a time dimension, places the first information on one or more images, and generates a bitstream by encoding the one or more images.

[0007] A decoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit, and in operation, the circuit decodes one or more images from a bitstream generated by encoding one or more images in which first information included in a model, the first information indicating a plurality of features having a temporal dimension, is arranged, and obtains the first information from the one or more images.

[0008] Each embodiment of the present disclosure, or a partial configuration or method thereof, enables at least one of, for example, improved coding efficiency, improved image quality, reduced encoding / decoding processing volume, reduced circuit size, or improved encoding / decoding processing speed. Alternatively, each embodiment of the present disclosure, or a partial configuration or method thereof, enables appropriate selection of components / operations such as filters, block sizes, motion vectors, reference pictures, and reference blocks in encoding and decoding. Note that the present disclosure also includes disclosure of configurations or methods that may provide benefits other than those described above. For example, a configuration or method that improves coding efficiency while suppressing an increase in processing volume.

[0009] Further advantages and benefits of certain aspects of the present disclosure will become apparent from the specification and drawings. While such advantages and / or benefits may be obtained by several embodiments and features described in the specification and drawings, not all of them necessarily need to be provided to obtain one or more advantages and / or benefits.

[0010] These general or specific aspects may be realized by a system, an integrated circuit, a computer program, or a recording medium such as a computer-readable CD-ROM, or may be realized by any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0011] A configuration or method according to an aspect of the present disclosure may contribute to, for example, one or more of improved coding efficiency, improved image quality, reduced processing amount, reduced circuit size, improved processing speed, and appropriate selection of elements or operations, etc. Note that a configuration or method according to an aspect of the present disclosure may also contribute to benefits other than those described above.

[0012] FIG. 1 is a schematic diagram illustrating an example of a configuration of a transmission system according to an embodiment. FIG. 2 is a diagram illustrating an example of a hierarchical structure of data in a stream. FIG. 3 is a diagram illustrating an example of a slice configuration. FIG. 4 is a diagram illustrating an example of a tile configuration. FIG. 5 is a diagram illustrating an example of an encoding structure for scalable encoding. FIG. 6 is a diagram illustrating an example of an encoding structure for scalable encoding. FIG. 7 is a block diagram illustrating an example of a configuration of an encoding device according to an embodiment. FIG. 8 is a block diagram illustrating an example of an implementation of an encoding device. FIG. 9 is a flowchart illustrating an example of an overall encoding process by an encoding device. FIG. 10 is a block diagram illustrating a configuration of a decoding device according to an embodiment. FIG. 11 is a block diagram illustrating an example of an implementation of a decoding device. FIG. 12 is a flowchart illustrating an example of an overall decoding process by a decoding device. FIG. 13 is a diagram illustrating an example of resolution of a dynamic 3D scene along the X, Y, and Z axes according to an embodiment. FIG. 14 is a diagram illustrating an example of multiple scales according to an embodiment. FIG. 15 is a diagram illustrating an example of a dynamic 3D scene with multiple frames according to an embodiment. FIG. 16 is a diagram illustrating an example of a 3D video sequence with multiple dynamic 3D scenes according to an embodiment. FIG. 17 is a diagram illustrating an example of a feature plane according to an embodiment. FIG. 18 is a diagram illustrating an example of multiple feature transformations in a dynamic 3D scene according to an embodiment. FIG. 19 is a diagram illustrating an example of a set of time-independent feature planes according to an embodiment. FIG. 20 is a diagram illustrating an example of a set of time-dependent feature planes according to an embodiment. FIG. 21 is a diagram illustrating an example of multiple scales of time-independent feature planes according to an embodiment. FIG. 22 is a diagram illustrating an example of multiple scales of time-dependent feature planes according to an embodiment. FIG. 23 is a diagram illustrating an example of a model of a 3D scene according to an embodiment. FIG. 24 is a diagram illustrating another example of a model of a 3D scene according to an embodiment. FIG. 25 is a diagram illustrating a rendering process by a rendering unit according to an embodiment. FIG. 26 is a diagram illustrating an example of images used in a training process according to an embodiment. FIG. 27 is a diagram illustrating an example of a model training process according to an embodiment. FIG. 28 is a block diagram illustrating a configuration of an encoding device according to an embodiment. FIG. 29 is a block diagram illustrating a configuration of a decoding device according to an embodiment.FIG. 30 is a block diagram showing another example of the configuration of an encoding device according to an embodiment. FIG. 31 is a block diagram showing another example of the configuration of a decoding device according to an embodiment. FIG. 32 is a flowchart of decoding processing by a decoding device according to the first aspect. FIG. 33 is a diagram showing an example of packing feature planes into multiple images according to an embodiment. FIG. 34 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 35 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 36 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 37 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 38 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 39 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 40 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 41 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 42 is a diagram showing another example of packing feature planes into multiple images according to an embodiment. FIG. 43 is a diagram illustrating another example of packing feature planes into multiple images according to an embodiment. FIG. 44 is a diagram illustrating an example of unpacking time-independent feature plane data from a feature map image into a feature plane according to an embodiment. FIG. 45 is a diagram illustrating an example of unpacking segments of time-dependent feature plane data into a feature plane according to an embodiment. FIG. 46 is a diagram illustrating an example of updating a model of a dynamic 3D representation according to an embodiment. FIG. 47 is a diagram illustrating an example of updating a model using time-dependent feature plane data by a decoding device according to an embodiment. FIG. 48 is a diagram illustrating another example of packing feature planes into multiple images according to an embodiment. FIG. 49 is a flowchart of encoding processing by an encoding device according to the first aspect. FIG. 50 is a diagram illustrating an example of packing a feature plane into a feature map image according to an embodiment. FIG. 51 is a diagram illustrating an example of dividing a time-dependent feature plane into segments according to an embodiment. FIG. 52 is a flowchart of encoding processing by an encoding device according to an embodiment. FIG. 53 is a flowchart of decoding processing by a decoding device according to an embodiment.FIG. 54 is a flowchart of a decoding process by a decoding device according to a second aspect. FIG. 55 is a block diagram showing the configuration of an encoding device according to an embodiment. FIG. 56 is a block diagram showing the configuration of a decoding device according to an embodiment. FIG. 57 is a diagram showing an example of packing feature planes and neural network weights into an image according to an embodiment. FIG. 58 is a diagram showing another example of packing feature planes and neural network weights into an image according to an embodiment. FIG. 59 is a diagram showing another example of packing feature planes and neural network weights into an image according to an embodiment. FIG. 60 is a diagram showing another example of packing feature planes and neural network weights into an image according to an embodiment. FIG. 61 is a diagram showing another example of packing feature planes and neural network weights into an image according to an embodiment. FIG. 62 is a diagram showing an example of updating a model of dynamic 3D representation according to an embodiment. FIG. 63 is a diagram showing an example of updating a model of dynamic 3D representation according to an embodiment with neural network weights. FIG. 64 is a diagram showing an example of updating a model of dynamic 3D representation according to an embodiment using voxel grid information. Fig. 65 is a flowchart of encoding processing by an encoding device according to the second aspect. Fig. 66 is a flowchart of encoding processing by an encoding device according to an embodiment. Fig. 67 is a flowchart of decoding processing by a decoding device according to an embodiment. Fig. 68 is a diagram showing the overall configuration of a content supply system that realizes a content distribution service. Fig. 69 is a diagram showing an example of a display screen of a web page. Fig. 70 is a diagram showing an example of a display screen of a web page. Fig. 71 is a diagram showing an example of a smartphone. Fig. 72 is a block diagram showing an example of the configuration of a smartphone.

[0013] [Introduction] An encoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit. The circuit, in operation, acquires first information included in a model, the first information indicating a plurality of features having a time dimension, places the first information on one or more images, and encodes the one or more images to generate a bitstream. This allows the encoding device to encode the first information including the plurality of features having a time dimension included in the model using image encoding. This facilitates encoding of the information included in the model.

[0014] For example, the first information may include a plurality of pieces of data, and the bitstream may include second information indicating an arrangement order of the plurality of pieces of data, thereby allowing the decoding device to appropriately obtain the plurality of pieces of data from the image using the second information.

[0015] For example, the first information may include a plurality of pieces of data, and the bitstream may include third information indicating a start position of at least one of the pieces of data in the one or more images, allowing the decoding device to appropriately obtain the pieces of data from the images using the third information.

[0016] For example, the first information may include a plurality of feature planes each representing a plurality of feature quantities arranged two-dimensionally. This allows the encoding device to encode the information of the plurality of feature planes included in the model using image encoding. This facilitates the encoding process of the information included in the model.

[0017] For example, the first information may include first data including a plurality of feature quantities without the time dimension and a plurality of second data including a plurality of feature quantities with the time dimension, each of which corresponds to a different time interval, the one or more images including a plurality of first images and a plurality of second images, the first data being divided and arranged among the plurality of first images, and the plurality of second data being arranged among the plurality of second images, respectively. This reduces padding data included in the one or more images, thereby improving encoding efficiency.

[0018] For example, the first information may include first data including a plurality of feature amounts without the time dimension and a plurality of second data including a plurality of feature amounts with the time dimension, each of the first data corresponding to a different time interval, the one or more images including a first image and a second image, the first data being arranged in the first image, and the plurality of second data being arranged together in the second image. This allows the number of images to be encoded to be reduced, thereby improving encoding efficiency.

[0019] For example, the first information may include first data and second data having a higher resolution than the first data, the one or more images may include a first image and a second image that is encoded after the first image, the first data may be arranged in the first image, and the second data may be arranged in the second image. In this way, the decoding device can first decode the first data and generate a model using the first data.

[0020] For example, the first information may include first data and second data having a higher resolution than the first data, each of the one or more images may include a first component and second and third components each smaller than the first component, the first data may be arranged in at least one of the second component and the third component, and the second data may be arranged in the first component. This allows multiple data to be arranged efficiently using multiple components.

[0021] A decoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit, and the circuit, in operation, decodes one or more images from a bitstream generated by encoding one or more images in which first information included in a model, the first information indicating a plurality of features having a time dimension, is arranged, and obtains the first information from the one or more images. This allows the decoding device to decode the first information indicating a plurality of features having a time dimension included in the model using a decoding process based on image encoding. This makes it possible to easily realize the decoding process of the information included in the model.

[0022] For example, the first information may include a plurality of pieces of data, and the bitstream may include second information indicating an arrangement order of the plurality of pieces of data, thereby allowing the decoding device to appropriately obtain the plurality of pieces of data from the image using the second information.

[0023] For example, the first information may include a plurality of pieces of data, and the bitstream may include third information indicating a start position of at least one of the pieces of data in the one or more images, allowing the decoding device to appropriately obtain the pieces of data from the images using the third information.

[0024] For example, the first information may include a plurality of feature planes each representing a plurality of feature quantities arranged two-dimensionally. This allows the decoding device to decode the information of the plurality of feature planes included in the model using a decoding process based on image coding. This makes it possible to easily realize the decoding process of the information included in the model.

[0025] For example, the first information may include first data including a plurality of feature quantities without the time dimension and a plurality of second data including a plurality of feature quantities with the time dimension, each of which corresponds to a different time interval, the one or more images including a plurality of first images and a plurality of second images, the first data being divided and arranged among the plurality of first images, and the plurality of second data being arranged among the plurality of second images, respectively. This reduces padding data included in the one or more images, thereby improving encoding efficiency.

[0026] For example, the first information may include first data including a plurality of feature amounts without the time dimension and a plurality of second data including a plurality of feature amounts with the time dimension, each of the first data corresponding to a different time interval, the one or more images including a first image and a second image, the first data being arranged in the first image, and the plurality of second data being arranged together in the second image. This allows the number of images to be encoded to be reduced, thereby improving encoding efficiency.

[0027] For example, the first information may include first data and second data having a higher resolution than the first data, the one or more images may include a first image and a second image that is encoded after the first image, the first data may be arranged in the first image, and the second data may be arranged in the second image. In this way, the decoding device can first decode the first data and generate a model using the first data.

[0028] For example, the first information may include first data and second data having a higher resolution than the first data, each of the one or more images may include a first component and second and third components each smaller than the first component, the first data may be arranged in at least one of the second component and the third component, and the second data may be arranged in the first component. This allows multiple data to be arranged efficiently using multiple components.

[0029] An encoding method according to one aspect of the present disclosure acquires first information included in a model, the first information indicating a plurality of feature quantities having a time dimension, places the first information on one or more images, and encodes the one or more images to generate a bitstream. According to this encoding method, the first information indicating a plurality of feature quantities having a time dimension included in the model can be encoded using image encoding. This facilitates the encoding process of the information included in the model.

[0030] A decoding method according to one aspect of the present disclosure includes decoding one or more images from a bitstream generated by encoding one or more images in which first information included in a model, the first information indicating a plurality of feature quantities having a time dimension, is arranged, and acquiring the first information from the one or more images. According to this decoding method, the first information indicating a plurality of feature quantities having a time dimension included in the model can be decoded using a decoding process based on image encoding. This makes it possible to easily realize the decoding process of the information included in the model.

[0031] According to one aspect of the present disclosure, there is provided an encoding device including a circuit and a memory connected to the circuit, the circuit operating to acquire a plurality of neural network weights included in a model, arrange the plurality of neural network weights on one or more images, and generate a bitstream by encoding the one or more images. This allows the encoding device to encode the plurality of neural network weights included in the model using image encoding, thereby facilitating the encoding process of information included in the model.

[0032] For example, the bitstream may include first information indicating an arrangement order of the plurality of neural network weights, whereby the decoding device can appropriately obtain the plurality of neural network weights from the image using the first information.

[0033] For example, the bitstream may include metadata, and the metadata may include the first information and a plurality of parameters related to the model, thereby allowing the decoding device to generate a model using the plurality of parameters.

[0034] For example, the circuit may further acquire second information included in the model, the second information indicating a plurality of feature quantities having a time dimension, and the arrangement may arrange the plurality of neural network weights and the second information on the one or more images. This allows the encoding device to encode the second information indicating a plurality of feature quantities having a time dimension included in the model using image encoding. This facilitates the encoding process of the information included in the model.

[0035] For example, the plurality of neural network weights and the second information may be arranged in the same image included in the one or more images, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0036] For example, the plurality of neural network weights and the second information may be arranged in different images included in the one or more images, thereby allowing the decoding device to independently decode the plurality of neural network weights and the second information.

[0037] For example, each of the one or more images may include a plurality of components, and the plurality of neural network weights and the second information may be arranged in different components among the plurality of components, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0038] For example, each of the one or more images may include a plurality of layers, and the plurality of neural network weights and the second information may be arranged in different layers among the plurality of layers, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0039] A decoding device according to one aspect of the present disclosure includes a circuit and a memory connected to the circuit. The circuit, in operation, decodes one or more images from a bitstream generated by encoding the one or more images to which a plurality of neural network weights included in a model are assigned, and obtains the plurality of neural network weights from the one or more images. This allows the decoding device to decode the plurality of neural network weights included in the model using a decoding process based on image encoding. This facilitates the decoding process of information included in the model.

[0040] For example, the bitstream may include first information indicating an arrangement order of the plurality of neural network weights, whereby the decoding device can appropriately obtain the plurality of neural network weights from the image using the first information.

[0041] For example, the bitstream may include metadata, and the metadata may include the first information and a plurality of parameters related to the model, thereby allowing the decoding device to generate a model using the plurality of parameters.

[0042] For example, the one or more images may include the plurality of neural network weights and second information included in the model, the second information indicating a plurality of feature quantities having a time dimension, and the acquiring step may include acquiring the plurality of neural network weights and the second information from the one or more images. This allows the decoding device to decode the second information indicating a plurality of feature quantities having a time dimension included in the model using a decoding process based on image coding. This facilitates the decoding process of the information included in the model.

[0043] For example, the plurality of neural network weights and the second information may be arranged in the same image included in the one or more images, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0044] For example, the plurality of neural network weights and the second information may be arranged in different images included in the one or more images, thereby allowing the decoding device to independently decode the plurality of neural network weights and the second information.

[0045] For example, each of the one or more images may include a plurality of components, and the plurality of neural network weights and the second information may be arranged in different components among the plurality of components, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0046] For example, each of the one or more images may include a plurality of layers, and the plurality of neural network weights and the second information may be arranged in different layers among the plurality of layers, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0047] An encoding method according to one aspect of the present disclosure acquires a plurality of neural network weights included in a model, arranges the plurality of neural network weights on one or more images, and encodes the one or more images to generate a bitstream. According to this encoding method, the plurality of neural network weights included in the model can be encoded using image encoding, thereby easily realizing the encoding process of information included in the model.

[0048] A decoding method according to one aspect of the present disclosure decodes one or more images from a bitstream generated by encoding the one or more images to which multiple neural network weights included in a model are assigned, and obtains the multiple neural network weights from the one or more images. According to this decoding method, the multiple neural network weights included in the model can be decoded using a decoding process based on image encoding. This facilitates the decoding process of information included in the model.

[0049] These comprehensive or specific aspects may be realized as a system, a method, an integrated circuit, a computer program, or a computer-readable recording medium such as a CD-ROM, or may be realized as any combination of a system, a method, an integrated circuit, a computer program, and a recording medium.

[0050] [Definition of Terms] As an example, each term may be defined as follows.

[0051] (1) Image: A unit of data made up of a set of pixels, consisting of pictures or blocks smaller than pictures, and includes both moving images and still images.

[0052] (2) Picture: A processing unit of an image composed of a set of pixels, and is sometimes called a frame or field.

[0053] (3) Block: A processing unit for a set containing a specific number of pixels, and can be named anything, as shown in the following examples. It can also be shaped anything, including, for example, a rectangle made up of M×N pixels, a square made up of M×M pixels, a triangle, a circle, or any other shape.

[0054] (Examples of blocks) Slice / tile / brick CTU / superblock / basic division unit VPDU / hardware processing division unit CU / processing block unit / prediction block unit (PU) / orthogonal transform block unit (TU) / unit Sub-block

[0055] (4) Pixel / Sample A pixel / sample is a minimum unit point that constitutes an image, and includes not only pixels at integer positions but also pixels at decimal positions generated based on pixels at integer positions.

[0056] (5) Pixel Value / Sample Value: A value inherent to a pixel, including not only brightness value, color difference value, and RGB gradation, but also depth value or binary values ​​of 0 and 1.

[0057] (6) Flags: In addition to one bit, flags may be multi-bit, for example, parameters or indexes of two or more bits. In addition, flags may be multi-valued using other bases as well as two values ​​using binary numbers.

[0058] (7) Signal: A signal that is symbolized or coded to transmit information, including discrete digital signals as well as analog signals that take continuous values.

[0059] (8) Stream / Bitstream: A digital data string or flow. A stream / bitstream may consist of a single stream or multiple streams divided into multiple layers. It also includes transmission by serial communication over a single transmission path as well as transmission by packet communication over multiple transmission paths.

[0060] (9) Difference / Difference In the case of scalar quantities, in addition to simple difference (x-y), it is sufficient to include difference calculations, including absolute value of difference (|x-y|), squared difference (x^2-y^2), square root of difference (√(x-y)), weighted difference (ax-by: a, b are constants), and offset difference (x-y+a: a is an offset).

[0061] (10) Sum In the case of a scalar quantity, in addition to simple sum (x + y), it is sufficient if a sum operation is included, including the absolute value of the sum (|x + y|), sum of squares (x^2 + y^2), square root of the sum (√(x + y)), weighted sum (ax + by: a, b are constants), and offset sum (x + y + a: a is an offset).

[0062] (11) Based on: This includes cases where factors other than the one being based on are taken into consideration. It also includes cases where a result is obtained directly or via an intermediate result.

[0063] (12) Using (used, using) This includes cases where elements other than the target of use are taken into account. It also includes cases where a result is obtained directly or via an intermediate result.

[0064] (13) Prohibit (forbid) This can be rephrased as not being allowed. Also, not prohibiting or being allowed does not necessarily mean obligation.

[0065] (14) Limit (restriction / restrict / restricted) This can be rephrased as not being permitted. Also, not prohibiting something or being permitted does not necessarily mean that it is an obligation. Furthermore, it is sufficient if something is partially prohibited in terms of quantity or quality, and it also includes cases where something is completely prohibited.

[0066] (15) Chroma: An adjective, denoted by the symbols Cb and Cr, that specifies that a sample array or a single sample represents one of two color difference signals associated with a primary color. Instead of the term chroma, the term chrominance can also be used.

[0067] (16) Luma: An adjective, denoted by the symbol or subscript Y or L, that specifies that a sample array or a single sample represents a monochrome signal associated with a primary color. Instead of the term luma, the term luminance may also be used.

[0068] [Description] In the drawings, the same reference numerals refer to the same or similar elements, and the sizes and relative positions of the elements in the drawings are not necessarily drawn to scale.

[0069] Hereinafter, embodiments will be described in detail with reference to the drawings. Note that the embodiments described below are all comprehensive or specific examples. The numerical values, shapes, materials, components, the arrangement and connection of the components, steps, and the relationship and order of the steps shown in the following embodiments are merely examples and are not intended to limit the scope of the claims.

[0070] Below, embodiments of an encoding device and a decoding device will be described. The embodiments are examples of encoding devices and decoding devices to which the processes and / or configurations described in each aspect of the present disclosure can be applied. The processes and / or configurations can also be implemented in encoding devices and decoding devices different from the embodiments. For example, with regard to the processes and / or configurations applied to the embodiments, any of the following may be implemented.

[0071] (1) Any of the multiple components of the encoding device or decoding device of the embodiments described in each aspect of the present disclosure may be replaced or combined with other components described in any of the aspects of the present disclosure.

[0072] (2) In the encoding device or decoding device according to the embodiment, the functions or processes performed by some of the components of the encoding device or decoding device may be changed in any way, such as by adding, replacing, or deleting a function or process. For example, any function or process may be replaced with or combined with another function or process described in any of the aspects of the present disclosure.

[0073] (3) In the method implemented by the encoding device or decoding device according to the embodiment, some of the processes included in the method may be arbitrarily modified, such as by addition, replacement, deletion, etc. For example, any process in the method may be replaced with or combined with another process described in any of the aspects of the present disclosure.

[0074] (4) Some of the components constituting the encoding device or decoding device of the embodiment may be combined with components described in any of the aspects of the present disclosure, or may be combined with components having some of the functions described in any of the aspects of the present disclosure, or may be combined with components that perform some of the processing performed by the components described in each aspect of the present disclosure.

[0075] (5) A component having part of the functionality of the encoding device or decoding device of an embodiment, or a component that performs part of the processing of the encoding device or decoding device of an embodiment, may be combined or replaced with a component described in any of the aspects of the present disclosure, a component having part of the functionality described in any of the aspects of the present disclosure, or a component that performs part of the processing described in any of the aspects of the present disclosure.

[0076] (6) In the method implemented by the encoding device or decoding device of the embodiment, any of the multiple processes included in the method may be replaced or combined with the process described in any of the aspects of the present disclosure or any similar process.

[0077] (7) Some of the processes included in the method implemented by the encoding device or decoding device of the embodiment may be combined with the processes described in any of the aspects of the present disclosure.

[0078] (8) The implementation of the processes and / or configurations described in each aspect of the present disclosure is not limited to the encoding device or decoding device of the embodiments. For example, the processes and / or configurations may be implemented in a device used for a purpose other than video encoding or video decoding disclosed in the embodiments.

[0079] [System Configuration] FIG. 1 is a schematic diagram showing an example of the configuration of a transmission system according to this embodiment.

[0080] The transmission system Trs is a system that transmits a stream generated by encoding an image and decodes the transmitted stream. Such a transmission system Trs includes, for example, an encoding device 100, a network Nw, and a decoding device 200, as shown in FIG.

[0081] An image is input to the encoding device 100. The encoding device 100 generates a stream by encoding the input image and outputs the stream to the network Nw. The stream includes, for example, the encoded image and control information for decoding the encoded image. The image is compressed by this encoding.

[0082] Note that the original image before encoding that is input to the encoding device 100 is also called an original image, an original signal, or an original sample. The image may be a moving image or a still image. The image is a broader concept than sequences, pictures, and blocks, and is not limited in spatial or temporal domain unless otherwise specified. The image is composed of an array of pixels or pixel values, and the signal representing the image or the pixel values ​​is also called a sample. The stream may also be called a bitstream, coded bitstream, compressed bitstream, or coded signal. The encoding device may also be called an image encoding device or a moving image encoding device, and the encoding method used by the encoding device 100 may also be called an encoding method, an image coding method, or a moving image coding method.

[0083] The network Nw transmits the stream generated by the encoding device 100 to the decoding device 200. The network Nw may be the Internet, a wide area network (WAN), a local area network (LAN), or a combination of these. The network Nw is not necessarily limited to a bidirectional communication network, but may also be a unidirectional communication network that transmits broadcast waves such as terrestrial digital broadcasting or satellite broadcasting. The network Nw may also be replaced by a storage medium on which a stream is recorded, such as a DVD (Digital Versatile Disc) or a BD (Blu-Ray Disc (registered trademark)).

[0084] The decoding device 200 generates a decoded image, which is, for example, an uncompressed image, by decoding the stream transmitted over the network Nw. For example, the decoding device decodes the stream according to a decoding method corresponding to the encoding method used by the encoding device 100.

[0085] The decoding device may be called an image decoding device or a video decoding device, and the decoding method performed by the decoding device 200 may be called a decoding method, an image decoding method, or a video decoding method.

[0086] [Data Structure] Figure 2 is a diagram showing an example of a hierarchical structure of data in a stream. The stream includes, for example, a video sequence. This video sequence includes, for example, a video parameter set (VPS), a sequence parameter set (SPS), a picture parameter set (PPS), supplemental enhancement information (SEI), and multiple pictures, as shown in Figure 2(a).

[0087] In a video composed of multiple layers, the VPS includes coding parameters common to multiple layers, and coding parameters related to multiple layers included in the video or to each individual layer.

[0088] The SPS includes parameters used for the sequence, i.e., encoding parameters that the decoding device 200 refers to in order to decode the sequence. For example, the encoding parameters may indicate the width or height of a picture. Note that there may be multiple SPSs.

[0089] The PPS includes parameters used for a picture, i.e., encoding parameters referenced by the decoding device 200 to decode each picture in a sequence. For example, the encoding parameters may include a reference value of the quantization width used in decoding the picture and a flag indicating the application of weighted prediction. Note that there may be multiple PPSs. Furthermore, the SPS and PPS may be simply referred to as parameter sets.

[0090] A picture may include a picture header and one or more slices, as shown in Fig. 2(b), where the picture header includes coding parameters that the decoding device 200 references to decode the one or more slices.

[0091] As shown in (c) of Fig. 2, a slice includes a slice header and one or more bricks. The slice header includes coding parameters that are referenced by the decoding device 200 to decode the one or more bricks.

[0092] A brick includes one or more coding tree units (CTUs), as shown in FIG. 2(d).

[0093] Note that a picture may not contain slices, but may instead contain tile groups, where a tile group contains one or more tiles, and a brick may contain slices.

[0094] A CTU is also called a superblock or a basic division unit. As shown in (e) of Fig. 2, such a CTU includes a CTU header and one or more coding units (CUs). The CTU header includes coding parameters that the decoding device 200 references to decode the one or more CUs.

[0095] A CU may be divided into multiple small CUs. Furthermore, as shown in (f) of FIG. 2, a CU includes a CU header, prediction information, and residual coefficient information. The prediction information is information for predicting the CU, and the residual coefficient information is information indicating a prediction residual, which will be described later. A CU is basically the same as a PU (Prediction Unit) and a TU (Transform Unit), but may include multiple TUs smaller than the CU, for example, in an SBT, which will be described later. A CU may also be processed for each VPDU (Virtual Pipeline Decoding Unit) that constitutes the CU. A VPDU is a fixed unit that can be processed in one stage, for example, when performing pipeline processing in hardware.

[0096] Note that the stream may not have some of the layers shown in FIG. 2 . The order of these layers may be changed, or some layers may be replaced with other layers. A picture currently being processed by a device such as the encoding device 100 or the decoding device 200 is referred to as a current picture. If the processing is encoding, the current picture is synonymous with a picture to be encoded, and if the processing is decoding, the current picture is synonymous with a picture to be decoded. A block, such as a CU or CU, currently being processed by a device such as the encoding device 100 or the decoding device 200 is referred to as a current block. If the processing is encoding, the current block is synonymous with a block to be encoded, and if the processing is decoding, the current block is synonymous with a block to be decoded.

[0097] [Picture Structure: Slices / Tiles] In order to decode pictures in parallel, pictures may be structured in slices or tiles.

[0098] A slice is a basic coding unit that constitutes a picture. A picture is made up of, for example, one or more slices. A slice is made up of one or more consecutive CTUs.

[0099] FIG. 3 illustrates an example of a slice configuration. For example, a picture includes 11×8 CTUs and is divided into four slices (slices 1-4). Slice 1 may include, for example, 16 CTUs, slice 2 may include, for example, 21 CTUs, slice 3 may include, for example, 29 CTUs, and slice 4 may include, for example, 22 CTUs. Each CTU in a picture belongs to one of the slices. The shape of a slice is determined by dividing the picture horizontally. The slice boundaries do not necessarily have to be at the edges of the screen, but may be anywhere within the boundaries of the CTUs within the screen. The processing order (encoding order or decoding order) of the CTUs in a slice is, for example, raster scan order. Each slice also includes a slice header and coded data. The slice header may describe the characteristics of the slice, such as the address of the first CTU in the slice and the slice type.

[0100] A tile is a unit of rectangular area that makes up a picture. A number called TileId may be assigned to each tile in raster scan order.

[0101] FIG. 4 is a diagram illustrating an example of a tile configuration. For example, a picture includes 11×8 CTUs and is divided into four rectangular tiles (tiles 1-4). When tiles are used, the processing order of the CTUs is changed compared to when tiles are not used. When tiles are not used, multiple CTUs in a picture are processed, for example, in raster scan order. When tiles are used, at least one CTU in each of multiple tiles is processed, for example, in raster scan order. For example, as shown in FIG. 4, the processing order of the multiple CTUs included in tile 1 is from the left end of the first column of tile 1 to the right end of the first column of tile 1, and then from the left end of the second column of tile 1 to the right end of the second column of tile 1.

[0102] It should be noted that one tile may include one or more slices, and one slice may include one or more tiles.

[0103] Note that a picture may be composed of tile sets. A tile set may include one or more tile groups or one or more tiles. A picture may be composed of only one of tile sets, tile groups, and tiles. For example, the order in which multiple tiles for each tile set are scanned in raster order is defined as the basic coding order of the tiles. A collection of one or more tiles in each tile set that follow the basic coding order is defined as a tile group. Such a picture may be composed by the dividing unit 102 (see FIG. 7 ), which will be described later.

[0104] [Scalable Coding] FIGS. 5 and 6 are diagrams showing an example of the structure of a scalable stream.

[0105] As shown in FIG. 5 , the encoding device 100 may generate a temporally / spatially scalable stream by encoding each of a plurality of pictures into one of a plurality of layers. For example, the encoding device 100 may achieve scalability by encoding pictures layer by layer, where an enhancement layer exists above a base layer. This coding of each picture is called scalable coding. This allows the decoding device 200 to switch the image quality of the image displayed by decoding the stream. That is, the decoding device 200 determines up to which layer to decode based on internal factors such as its own performance and external factors such as the state of the communication bandwidth. As a result, the decoding device 200 can freely switch between low-resolution content and high-resolution content and decode the same content. For example, a user of the stream may watch a video stream partway through using a smartphone while on the move, and then watch the rest of the video using a device such as an Internet TV after returning home. Note that the above-mentioned smartphone and device each incorporate a decoding device 200 with the same or different performance. In this case, if the device decodes the upper layers of the stream, the user can view high-quality video after returning home. This eliminates the need for the encoding device 100 to generate multiple streams with the same content but different image qualities, thereby reducing the processing load.

[0106] Furthermore, the enhancement layer may include meta-information based on image statistical information, etc. The decoding device 200 may generate high-quality moving images by super-resolving pictures in the base layer based on the meta-information. Super-resolution may be either an improvement in the signal-to-noise (SN) ratio at the same resolution or an increase in resolution. The meta-information may include information for specifying linear or nonlinear filter coefficients used in the super-resolution process, or information for specifying parameter values ​​in the filter process, machine learning, or least-squares calculation used in the super-resolution process.

[0107] Alternatively, a picture may be divided into tiles or the like according to the meaning of each object in the picture. In this case, the decoding device 200 may decode only a portion of the picture by selecting tiles to be decoded. Furthermore, attributes of objects (such as a person, a car, or a ball) and their positions within the picture (such as coordinate positions within the same picture) may be stored as meta information. In this case, the decoding device 200 can identify the position of a desired object based on the meta information and determine the tile containing the object. For example, as shown in FIG. 6 , the meta information is stored using a data storage structure different from that of image data, such as SEI in HEVC. This meta information indicates, for example, the position, size, or color of the main object.

[0108] Furthermore, the meta information may be stored in units consisting of multiple pictures, such as streams, sequences, or random access units, etc. This allows the decoding device 200 to obtain the time at which a specific person appears in a video, and by using the time and the information in units of pictures, it is possible to identify the picture in which the object exists and the position of the object within that picture.

[0109] [Encoding Device] Next, a description will be given of an encoding device 100 according to an embodiment. Fig. 7 is a block diagram showing an example of the configuration of the encoding device 100 according to an embodiment. The encoding device 100 encodes an image in units of blocks.

[0110] 7, the encoding device 100 is a device that encodes an image in units of blocks, and includes a division unit 102, a subtraction unit 104, a transformation unit 106, a quantization unit 108, an entropy encoding unit 110, an inverse quantization unit 112, an inverse transformation unit 114, an addition unit 116, a block memory 118, a loop filter unit 120, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, a prediction control unit 128, and a prediction parameter generation unit 130. Note that the intra prediction unit 124 and the inter prediction unit 126 are each configured as part of a prediction processing unit.

[0111] [Implementation Example of Encoding Device] Fig. 8 is a block diagram showing an implementation example of the encoding device 100. The encoding device 100 includes a processor a1 and a memory a2. For example, several components of the encoding device 100 shown in Fig. 7 are implemented by the processor a1 and the memory a2 shown in Fig. 8.

[0112] The processor a1 is a circuit that performs information processing and is a circuit that can access the memory a2. For example, the processor a1 is a dedicated or general-purpose electronic circuit that encodes images. The processor a1 may be a processor such as a CPU. The processor a1 may also be a collection of multiple electronic circuits. For example, the processor a1 may fulfill the roles of multiple components of the encoding device 100 shown in FIG. 7 , excluding the components for storing information.

[0113] The memory a2 is a dedicated or general-purpose memory that stores information used by the processor a1 to encode images. The memory a2 may be an electronic circuit and may be connected to the processor a1. The memory a2 may also be included in the processor a1. The memory a2 may also be a collection of multiple electronic circuits. The memory a2 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as storage, a recording medium, or the like. The memory a2 may also be a non-volatile memory or a volatile memory.

[0114] For example, the memory a2 may store an image to be encoded, or a stream corresponding to the encoded image, or may store a program for the processor a1 to encode the image.

[0115] Furthermore, for example, the memory a2 may serve as a component for storing information among the multiple components of the encoding device 100 shown in Fig. 7. Specifically, the memory a2 may serve as the block memory 118 and the frame memory 122 shown in Fig. 7. More specifically, the memory a2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).

[0116] 7 may not be implemented in the encoding device 100, and all of the above-described processes may not be performed. Some of the components shown in Fig. 7 may be included in another device, and some of the above-described processes may be performed by another device.

[0117] Below, the overall processing flow of the encoding device 100 will be explained, and then each component included in the encoding device 100 will be explained.

[0118] [Overall Flow of Encoding Process] FIG. 9 is a flowchart showing an example of the overall encoding process performed by the encoding device 100.

[0119] First, the division unit 102 of the encoding device 100 divides a picture included in an original image into a plurality of fixed-size blocks (128 x 128 pixels) (step Sa_1). Then, the division unit 102 selects a division pattern for the fixed-size blocks (step Sa_2). That is, the division unit 102 further divides the fixed-size block into a plurality of blocks that constitute the selected division pattern. Then, the encoding device 100 performs the processes of steps Sa_3 to Sa_9 for each of the plurality of blocks.

[0120] The prediction processing unit, which is made up of the intra prediction unit 124 and the inter prediction unit 126, and the prediction control unit 128 generate a predicted image of the current block (step Sa_3). Note that the predicted image is also called a predicted signal, a predicted block, or a predicted sample.

[0121] Next, the subtraction unit 104 generates a difference between the current block and the predicted image as a prediction residual (step Sa_4). Note that the prediction residual is also called a prediction error.

[0122] Next, the transform unit 106 and the quantization unit 108 perform transform and quantization on the predicted image to generate a plurality of quantized coefficients (step Sa_5).

[0123] Next, the entropy coding unit 110 generates a stream by performing coding (specifically, entropy coding) on ​​the plurality of quantized coefficients and prediction parameters related to generation of a predicted image (step Sa_6).

[0124] Next, the inverse quantization unit 112 and the inverse transform unit 114 perform inverse quantization and inverse transform on the plurality of quantized coefficients to reconstruct the prediction residuals (step Sa_7).

[0125] Next, the adder 116 reconstructs the current block by adding the predicted image to the restored prediction residual (step Sa_8). This generates a reconstructed image. Note that the reconstructed image is also called a reconstructed block, and in particular, the reconstructed image generated by the encoding device 100 is also called a locally decoded block or a locally decoded image.

[0126] Once this reconstructed image is generated, the loop filter unit 120 performs filtering on the reconstructed image as needed (step Sa_9).

[0127] Then, the encoding device 100 determines whether or not encoding of the entire picture has been completed (step Sa_10), and if it determines that encoding has not been completed (No in step Sa_10), it repeats the processing from step Sa_2.

[0128] In the above example, the encoding device 100 selects one division pattern for fixed-size blocks and encodes each block according to that division pattern, but it may also encode each block according to each of a plurality of division patterns. In this case, the encoding device 100 may evaluate the cost for each of the plurality of division patterns and select, for example, the stream obtained by encoding according to the division pattern with the smallest cost as the stream to be finally output.

[0129] Furthermore, the processing of steps Sa_1 to Sa_10 may be performed sequentially by the encoding device 100, or some of the processing may be performed in parallel, or the order of the processing may be changed.

[0130] The coding process performed by the coding device 100 is hybrid coding that uses predictive coding and transform coding. The predictive coding is performed by a coding loop that includes a subtraction unit 104, a transform unit 106, a quantization unit 108, an inverse quantization unit 112, an inverse transform unit 114, an addition unit 116, a loop filter unit 120, a block memory 118, a frame memory 122, an intra prediction unit 124, an inter prediction unit 126, and a prediction control unit 128. In other words, the prediction processing unit that includes the intra prediction unit 124 and the inter prediction unit 126 forms part of the coding loop.

[0131] [Decoding Device] Next, a description will be given of a decoding device 200 capable of decoding the stream output from the above-described encoding device 100. Fig. 10 is a block diagram showing an example of the configuration of the decoding device 200 according to an embodiment. The decoding device 200 is a device that decodes a stream, which is an encoded image, in units of blocks.

[0132] 10 , the decoding device 200 includes an entropy decoding unit 202, an inverse quantization unit 204, an inverse transform unit 206, an adder 208, a block memory 210, a loop filter unit 212, a frame memory 214, an intra prediction unit 216, an inter prediction unit 218, a prediction control unit 220, a prediction parameter generation unit 222, and a partition determination unit 224. Note that the intra prediction unit 216 and the inter prediction unit 218 are each configured as part of a prediction processing unit.

[0133] [Implementation Example of Decoding Device] Fig. 11 is a block diagram showing an implementation example of a decoding device 200. The decoding device 200 includes a processor b1 and a memory b2. For example, multiple components of the decoding device 200 shown in Fig. 10 are implemented by the processor b1 and the memory b2 shown in Fig. 11.

[0134] The processor b1 is a circuit that performs information processing and is a circuit that can access the memory b2. For example, the processor b1 is a dedicated or general-purpose electronic circuit that decodes a stream. The processor b1 may be a processor such as a CPU. The processor b1 may also be a collection of multiple electronic circuits. For example, the processor b1 may fulfill the roles of multiple components of the decoding device 200 shown in FIG. 10 and the like, excluding the component for storing information.

[0135] The memory b2 is a dedicated or general-purpose memory that stores information for the processor b1 to decode the stream. The memory b2 may be an electronic circuit and may be connected to the processor b1. The memory b2 may also be included in the processor b1. The memory b2 may also be a collection of multiple electronic circuits. The memory b2 may also be a magnetic disk, an optical disk, or the like, and may also be expressed as storage, a recording medium, or the like. The memory b2 may also be a non-volatile memory or a volatile memory.

[0136] For example, the memory b2 may store an image or a stream, or may store a program for the processor b1 to decode the stream.

[0137] Furthermore, for example, the memory b2 may serve as a component for storing information among the multiple components of the decoding device 200 shown in Fig. 10 etc. Specifically, the memory b2 may serve as the block memory 210 and the frame memory 214 shown in Fig. 10. More specifically, the memory b2 may store a reconstructed image (specifically, a reconstructed block or a reconstructed picture, etc.).

[0138] Note that not all of the components shown in Fig. 10 etc. may be implemented, and not all of the above-described processes may be performed, in the decoding device 200. Some of the components shown in Fig. 10 etc. may be included in another device, and some of the above-described processes may be executed by another device.

[0139] Below, the overall processing flow of the decoding device 200 will be described, followed by a description of each component included in the decoding device 200. Note that detailed description of the components included in the decoding device 200 that perform the same processing as the components included in the encoding device 100 will be omitted. For example, the inverse quantization unit 204, inverse transform unit 206, adder 208, block memory 210, frame memory 214, intra prediction unit 216, inter prediction unit 218, prediction control unit 220, and loop filter unit 212 included in the decoding device 200 perform the same processing as the inverse quantization unit 112, inverse transform unit 114, adder 116, block memory 118, frame memory 122, intra prediction unit 124, inter prediction unit 126, prediction control unit 128, and loop filter unit 120 included in the encoding device 100, respectively.

[0140] [Overall Flow of Decoding Process] FIG. 12 is a flowchart showing an example of the overall decoding process performed by the decoding device 200.

[0141] First, the partition determination unit 224 of the decoding device 200 determines a partition pattern for each of a plurality of fixed-size blocks (128 × 128 pixels) included in a picture based on parameters input from the entropy decoding unit 202 (step Sp_1). This partition pattern is the partition pattern selected by the encoding device 100. Then, the decoding device 200 performs the processes of steps Sp_2 to Sp_6 on each of the plurality of blocks that make up that partition pattern.

[0142] The entropy decoding unit 202 decodes (specifically, entropy decodes) the coded quantized coefficients and prediction parameters of the current block (step Sp_2).

[0143] Next, the inverse quantization unit 204 and the inverse transform unit 206 perform inverse quantization and inverse transform on the plurality of quantized coefficients to reconstruct the prediction residuals of the current block (step Sp_3).

[0144] Next, the prediction processing unit, which is made up of the intra prediction unit 216, the inter prediction unit 218, and the prediction control unit 220, generates a predicted image of the current block (step Sp_4).

[0145] Next, the adder 208 reconstructs the current block into a reconstructed image (also called a decoded image block) by adding the predicted image to the prediction residual (step Sp_5).

[0146] Then, when this reconstructed image is generated, the loop filter unit 212 performs filtering on the reconstructed image (step Sp_6).

[0147] Then, the decoding device 200 determines whether or not the decoding of the entire picture is completed (step Sp_7), and if it determines that the decoding is not completed (No in step Sp_7), it repeats the process from step Sp_1.

[0148] The processes of steps Sp_1 to Sp_7 may be performed sequentially by the decoding device 200, or some of the processes may be performed in parallel, or the order of the processes may be changed.

[0149] 3D Video: We describe the field of immersive video streaming, also known as 3D (three-dimensional) video or volumetric video. Immersive video streaming allows viewers to experience six degrees of freedom (6DoF) within a video scene, allowing them to move and rotate in any direction. This technology has been used in interactive web viewers or virtual reality (VR) goggles.

[0150] To achieve a fully immersive effect, scenes must be rendered quickly, ideally in real time, based on the viewer's position and viewing direction. However, immersive video streaming presents several challenges, including the requirement for significant bandwidth and computational resources.

[0151] For example, virtual reality experiences require higher resolutions and frame rates to create realistic and immersive environments, and immersive video requires encoding information for multiple viewpoints, resulting in significantly larger file sizes.

[0152] Dynamic streaming is also essential to adjust video quality based on viewer orientation and focus and adapt to varying network conditions, so efficient compression and streaming techniques are needed to reduce throughput while improving the quality of the viewer experience.

[0153] Recent advances have introduced many methods for efficiently representing 3D scenes using feature planes and neural networks trained specifically for each scene. These techniques can effectively encode 3D video data in small file sizes while maintaining high quality of the rendered image.

[0154] Examples of immersive viewing experiences include sports, concerts, or reality shows. In sports viewing, viewers can experience the action from different perspectives on the field or court, immersing themselves in the live sporting event. In concert viewing, concertgoers can virtually be in the middle of the performance, enjoying an immersive experience. In reality show viewing, viewers can explore the environment from different perspectives, enhancing the viewer's experience of the reality show.

[0155] Immersive viewing experiences can also be applied to games or interactive entertainment, such as immersive gaming experiences where players can move through three-dimensional environments and interact with three-dimensional characters in a virtual world, and interactive storytelling experiences where viewers can influence the progression of the plot by making choices that affect the story.

[0156] Immersive viewing experiences can also be applied in education, design, or training. Immersive learning environments can provide students with interactive experiences that enhance comprehension and retention of material. In the design field, architects, interior designers, or engineers can use immersive technology to visualize and manipulate three-dimensional models of their projects. In training, immersive simulations can provide realistic training scenarios for various professions, such as medical, military, and aviation.

[0157] Immersive viewing experiences can also be applied to architecture or heritage. Immersive architectural visualization allows clients to experience proposed designs before construction begins, explore virtual buildings, evaluate spatial relationships, and provide feedback to architects or developers. Immersive experiences also allow users to explore historical sites or artifacts in detail, preserving cultural heritage and promoting virtual tourism.

[0158] Another approach for streaming immersive video is to encode and decode only the spatial information within the feature plane of each frame of 3D video, treating it as a static scene. However, this approach fails to exploit the temporal redundancy inherent in 3D video, hindering efficient data compression.

[0159] In this embodiment, feature planes containing temporal information are encoded as images using a conventional video codec (image coding standard) and transmitted. We also describe an optimal packing method for feature planes to minimize delay while maintaining high quality. We also describe additional information to be sent to a decoding device for 3D scene reconstruction. This allows us to provide a versatile decoding device capable of rendering a variety of 3D scenes.

[0160] 3D Video Resolution Figure 13 is a diagram illustrating an example of the resolution of a dynamic 3D scene along the X, Y, and Z axes. A bounding box defines the boundary of the 3D region of space that contains the scene of interest that is captured over the entire duration of the dynamic 3D scene. Specifically, the bounding box is the minimum coordinate point (e.g., in meters) of the bounding box of the 3D scene shown in Figure 13 (X min , Y min , Z min ), and the maximum coordinate point (e.g., in meters) of the bounding box of the 3D scene (X max , Y max , Z max ) is defined by

[0161] Also, W is the resolution of the 3D scene along the X axis (e.g., in pixels). X and W, the resolution of the 3D scene along the Y axis (e.g., in pixels). Y and W, the resolution of the 3D scene along the Z axis (e.g., in pixels). Z is set.

[0162] Here, the above information of the bounding box is used for efficient reconstruction at the decoder. This information is needed for both object-centric and forward-looking scenes. In object-centric scenes, the camera is typically focused on a particular object or point of interest in the scene and often rotates around it to show different perspectives of the object.

[0163] Camera pose in forward-facing scenes is characterized by a roughly consistent view direction relative to the environment, with primarily only translation of the camera pose.

[0164] 14 is a diagram illustrating an example of multiple scales of resolution for capturing different levels of detail in a dynamic 3D scene. The scale is defined by a parameter indicating the scale, the scale resolution multiplier, which is a one-dimensional array of integers (e.g., [1, 2, ..., N]). Note that hereinafter, the scale when the scale resolution multiplier = N is also referred to as "scale N."

[0165] For example, the scale resolution multiplier is a power of 2, e.g., [1, 2, 4, 8]. Note that the scale resolution multiplier is not limited to a power of 2. For example, the scale resolution multiplier may be a floating-point value, etc. In another example, there may be only one scale, i.e., scale resolution multiplier=1.

[0166] 15 shows an example of a dynamic 3D scene with multiple frames. The temporal aspect (denoted by S) of a video depicting a dynamic 3D scene is scaled by a parameter W that represents its temporal resolution. T In this representation, the initial frame of a dynamic 3D scene is aligned to the value 0, while the initial frame of a video sequence is aligned to the value 0. S Is W T Corresponds to L S denotes the total number of frames in scene S, and L S By W T The value of is determined. Note that a 3D scene refers to a series of continuous actions in a particular environment, and includes a number of consecutive frames.

[0167] Figure 16 shows an example of a 3D video sequence with multiple dynamic 3D scenes, where each frame can be identified by a combination of a scene number and a frame number within the scene. X , W Y , W Z , W T ) is common throughout the entire period (all frames) of the scene. In other words, the resolution is set for each scene. The resolution may be common to multiple scenes. Also, the examples shown in FIGS. 13 to 16 may be combined.

[0168] [Feature Plane] The encoding device transforms a 3D scene into a feature plane and encodes the transformed feature plane. FIG. 17 is a diagram showing an example of a feature plane. The feature plane, for example, is composed of multiple two-dimensional layers. Each layer is called a channel and is composed of multiple pixels arranged two-dimensionally. Each pixel in the feature plane represents multi-channel feature data and is a one-dimensional array of floating-point data. For example, an array of pixel values ​​obtained from all channels represents one feature.

[0169] For example, the feature value indicates whether an object exists at the corresponding spatial position, an attribute value of the object, or a combination thereof. Here, the attribute value includes at least one of color, transmittance, and reflectance, for example.

[0170] The height resolution of the feature plane W 2 and width resolution W 1 is obtained from the resolution information of the dynamic 3D scene. The number of channels C is determined based on, for example, the specific 3D scene representation method used. That is, the number of channels C is common to multiple scenes, regardless of the dynamic 3D scene.

[0171] For example, the value stored in each cell of a feature plane containing multiple channels is a 32-bit or 64-bit floating-point value. As an example, the dimensions of a feature plane may be 32(C) x 64(W). 1 ) x 64 (W 2 ) Note that the values ​​stored in each cell of the feature plane channels may be 32-bit or 64-bit fixed-point values, or values ​​greater than 64-bit floating-point or fixed-point values, or values ​​smaller than 32-bit floating-point or fixed-point values.

[0172] Figure 18 shows an example of multiple feature transformations in a dynamic 3D scene. It shows multiple feature planes corresponding to multiple base planes (XY, YZ, XZ, XT, YT, ZT) of a single frame at a single scale (scale=1). Such a representation is generated for each scale of the dynamic 3D scene, resulting in six feature planes at each scale.

[0173] Examples of mapping information from 4D (four-dimensional) points (x, y, z, t) of a 3D scene to grid nodes of a feature plane include simple projection and interpolation. A time-independent feature plane (XY, YZ, XZ), which is a feature plane that does not depend on time, mainly contains information about static parts of a 3D scene. For example, the feature value of each pixel in the time-independent feature plane is obtained by projecting an object (moving object) onto the corresponding plane (XY, YZ, XZ), and indicates whether an object exists at the corresponding position, the attribute value of the object, or a combination of these.

[0174] Furthermore, the time-dependent feature plane (XT, YT, ZT), which is a feature plane that depends on time, contains dynamic information of the scene. For example, each pixel in the time-dependent feature plane indicates a feature amount projected onto the corresponding plane (XT, YT, ZT). This feature amount indicates whether or not an object is moving, the amount of movement, the direction of movement, the difference, or a combination of these.

[0175] 19 is a diagram showing an example of a set of time-independent feature planes. As shown in FIG. 19, the dimensions of one time-independent feature plane are C×W 1 ×W 2 Here, W 1 and W 2 is the product of the dimensions in each axis of the bounding box of the 3D scene at scale 1 and the scale resolution multiplier. For example, at scale 1 in the feature plane of XY, W 1 =W x , W 2 =W Y At the scale 2 of the feature plane of XY, W 1 = 2 x W x , W 2 = 2 x W Y At the scale N of the feature plane in XY, W 1 = N × W x , W 2 = N × W Y is.

[0176] 19, the set of time-independent feature planes includes an XY feature plane set, a YZ feature plane set, and an XZ feature plane set. Each of the XY, YZ, and XZ feature plane sets includes feature planes at multiple scales (1 to N).

[0177] 20 is a diagram showing an example of a set of time-dependent feature planes. As shown in FIG. 20, the dimensions of one time-dependent feature plane are C×W T ×W 1 Here, W T corresponds to the time dimension and represents the total duration T of the 3D scene. 1 is the product of the dimensions of each axis of the bounding box of the 3D scene at scale 1 and the scale resolution multiplier. For example, at scale 1 of the feature plane of XT, W 1 =W x At scale 2 of the feature plane of XT, W 1 = 2 x W x At the scale N of the feature plane of XT, W 1 = N × W x is.

[0178] 20, the set of time-dependent feature planes includes a set of XT feature planes, a set of YT feature planes, and a set of ZT feature planes. Each of the XT, YT, and ZT feature plane sets includes feature planes of multiple scales (1 to N).

[0179] Although only one scale of temporal resolution is shown in Fig. 20, multiple scales may be applied to the temporal resolution. Adding a scale in the time dimension improves the accuracy of the temporal resolution and also improves the accessibility of the video.

[0180] Figure 21 shows an example of multiple scales of time-independent feature planes. Figure 22 shows an example of multiple scales of time-dependent feature planes. As shown in Figures 21 and 22, changes in a 3D region of space (a 3D scene) are recorded as features on multiple feature planes with different resolutions across multiple scales. Figures 21 and 22 show an example where M scales are used, with M x 3 time-independent feature planes and M x 3 time-dependent feature planes. Note that the scale multiplier for scale M is N.

[0181] An example of a model of a 3D scene is shown in Figure 23. As shown in Figure 23, the features of the feature plane and the weights of the neural network are trained (learned) for a specific dynamic 3D scene to generate a model representing the 3D scene.

[0182] A dynamic 3D scene with multiple frames is represented by a model. The model includes time-independent feature planes and neural network weights as static information. The model also includes time-dependent feature planes as dynamic information. The time-independent feature planes are time-independent (spatial) feature planes, and the time-dependent feature planes are time-dependent (spatiotemporal) feature planes. The neural network weights are the weights of a trained neural network. Each model is associated with 3D scene parameters that are used to provide additional information to the model.

[0183] Although an example in which both feature planes (time-independent feature planes and time-dependent feature planes) and neural network weights are used has been shown here, a model may be generated using only feature planes without using neural network weights.

[0184] FIG. 24 is a diagram illustrating another example of a model of a 3D scene. As illustrated in FIG. 24, the model may further include a voxel grid as static information. The voxel grid is a 3D voxel grid with a resolution smaller than that of the feature plane. For example, the voxel grid includes information about the density of the 3D scene. For example, the voxel grid represents the object space using an octree structure, and each leaf includes information indicating whether or not the leaf includes a 3D point, or information indicating the number of 3D points included in the leaf. Using the voxel grid can speed up the reconstruction process. Note that the voxel grid may also include information about motion within the 3D scene.

[0185] FIG. 25 is a diagram illustrating the rendering process performed by the rendering unit. The rendering unit generates a rendered image by performing the rendering process using a model representing a dynamic 3D scene. Specifically, the rendering unit acquires the model (feature plane and neural network weights), 3D scene parameters, an input viewpoint, and an input time t. The viewpoint indicates the position and orientation of the camera. The rendering unit uses the features of the feature plane included in the model and the neural network weights to render (reconstruct) a rendered image viewed from the input viewpoint at the input time t. In this way, a rendered image viewed from an input arbitrary viewpoint (free viewpoint) is generated using the model.

[0186] More specifically, the four-dimensional coordinates (x, y, z, t) of the 3D model corresponding to the coordinates (x', y', t) on the rendering image plane are determined based on the input viewpoint. Feature quantities corresponding to these coordinates (x, y, z, t) are determined from the feature plane. Color information and other information is generated from these feature quantities and a neural network using neural network weights, and is reflected on the rendering image plane. The rendering image is generated by performing these processes on each pixel (coordinate) on the rendering image plane.

[0187] The 3D scene parameters include parameters for providing basic information required for rendering the 3D scene, such as at least one of a parameter identifying a 3D video sequence, a parameter identifying a scene, a start frame number, an end frame number, dimensions of a 3D scene bounding box, a feature plane grid resolution, a scale resolution multiplier, a feature plane quantization parameter, a start time, an end time, and a duration, etc. The model and 3D scene parameters are set during a training process, which will be described later.

[0188] 26 is a diagram showing an example of images used in the training process, in which images of a dynamic 3D scene are captured from multiple viewpoints (camera positions and orientations).

[0189] Figure 27 shows an example of the model training process (learning process). The features of the feature plane and the weights of the neural network are optimized based on the photometric loss, which is the difference in pixel values ​​between the ground truth image and the rendered image.

[0190] The rendering unit generates a rendered image using a viewpoint (camera position and orientation) from training data (teacher data), 3D scene parameters, and information from the model (e.g., color). The 3D scene parameters are parameters generated during the training process of a specific dynamic 3D scene. The model is initialized based on the 3D scene parameters. The difference between the rendered image and a ground truth image from the training data is calculated, and the calculated difference is fed back to the model. This optimizes the feature plane and neural network weights included in the model.

[0191] 28 is a block diagram showing the configuration of an encoding device 300 according to an embodiment. A model of a 3D scene representation method trained on one of the 3D scenes in a 3D video sequence is input to the encoding device 300. The encoding device 300 includes a conversion unit 301, an NNC encoding unit 302, a metadata conversion unit 303, and an image encoding unit 304.

[0192] The transform unit 301 extracts the feature planes, neural network weights, and associated 3D scene parameters from the model, and outputs the neural network weights to the NNC encoder 302.

[0193] The NNC (Neural Network Codec) encoding unit 302 compresses (encodes) the neural network weights to generate compressed neural network weights, and outputs the generated compressed neural network weights to the image encoding unit 304. For example, the NNC encoding unit 302 compresses the neural network weights using NNC (Neural Network Coding) of the MPEG standard. This allows for highly efficient encoding of data related to the neural network. However, the method for compressing the neural network weights is not limited to this.

[0194] The conversion unit 301 also rearranges (packs) the extracted feature planes into an image and outputs the generated image to the image encoding unit 304. The conversion unit 301 also outputs packing information related to the rearrangement to the metadata conversion unit 303. The conversion unit 301 updates 3D scene parameters required for reconstructing a 3D scene in the decoding device and outputs the updated 3D scene parameters to the metadata conversion unit 303.

[0195] The metadata conversion unit 303 converts the 3D scene parameters and packing information into metadata (parameter set) to be transmitted together with the bitstream. That is, the metadata conversion unit 303 generates metadata including the 3D scene parameters and packing information.

[0196] The image encoding unit 304 generates a bitstream by encoding the image, the compressed neural network weights, and the metadata. That is, the bitstream includes the image (encoded image), the compressed neural network weights, and the metadata. For example, the neural network weights are signaled as SEI (Supplemental Enhancement Information) or VUI (Video Usability Information) data.

[0197] For example, the image encoding unit 304 encodes the image using a known video codec. The configuration of the image encoding unit 304 may be the same as that of the encoding device 100 described above.

[0198] 29 is a block diagram showing the configuration of a decoding device 400 according to an embodiment. A bitstream is input to the decoding device 400. For example, the bitstream is generated by the encoding device 300 described above.

[0199] The decoding device 400 includes an image decoding unit 401 , a metadata inverse conversion unit 402 , an NNC decoding unit 403 , an inverse conversion unit 404 , and a rendering unit 405 .

[0200] The image decoding unit 401 decodes the bitstream to obtain images, compressed neural network weights, and metadata. For example, the image decoding unit 401 decodes images using a known video codec. The configuration of the image decoding unit 401 may be the same as that of the decoding device 200 described above.

[0201] The metadata inverse conversion unit 402 converts the metadata into 3D scene parameters and packing information. That is, the metadata inverse conversion unit 402 acquires the 3D scene parameters and packing information from the metadata.

[0202] The NNC decoder 403 generates neural network weights by expanding (decoding) the compressed neural network weights. The inverse transformer 404 transforms the image into a feature plane using the packing information.

[0203] The renderer 405 uses the neural network weights, feature planes, and 3D scene parameters to generate or update a model, and uses the model to generate a rendered image, which is the image seen from an input viewpoint requested by the software or user.

[0204] Fig. 30 is a block diagram showing another example of the configuration of an encoding device. The encoding device 310 shown in Fig. 30 differs from the encoding device 300 shown in Fig. 28 in that compressed neural network weights are stored in metadata.

[0205] The encoding device 310 receives a model of a 3D scene representation method trained on one of the 3D scenes in the 3D video sequence, and includes a transform unit 311, an NNC encoding unit 312, a metadata transform unit 313, and an image encoding unit 314.

[0206] The conversion unit 311 extracts feature planes, neural network weights, and 3D scene parameters from the model. The conversion unit 311 outputs the neural network weights to the NNC encoding unit 312. The conversion unit 311 also rearranges (packs) the extracted feature planes into an image and outputs the generated image to the image encoding unit 314. The conversion unit 311 also outputs packing information related to the rearrangement to the metadata conversion unit 313. The conversion unit 311 updates 3D scene parameters required for reconstructing a 3D scene in the decoding device and outputs the updated 3D scene parameters to the metadata conversion unit 313.

[0207] The NNC encoding unit 312 generates compressed neural network weights by compressing the neural network weights, and outputs the generated compressed neural network weights to the metadata conversion unit 313.

[0208] The metadata conversion unit 313 converts the 3D scene parameters, the compressed neural network weights, and the packing information into metadata. For example, the compressed neural network weights are packed and stored in a URL, and URL information indicating the URL is included in the metadata.

[0209] The image encoding unit 314 generates a bitstream by encoding images and metadata. That is, the bitstream includes images (encoded images) and metadata. For example, the image encoding unit 314 encodes images using a known video codec. The configuration of the image encoding unit 314 may be the same as that of the encoding device 100 described above.

[0210] 31 is a block diagram showing another example of the configuration of a decoding device. A bitstream is input to the decoding device 410. For example, the bitstream is generated by the encoding device 310 described above.

[0211] The decoding device 410 includes an image decoding unit 411 , a metadata inverse conversion unit 412 , an NNC decoding unit 413 , an inverse conversion unit 414 , and a rendering unit 415 .

[0212] The image decoding unit 411 decodes the bitstream to obtain images and metadata. The metadata inverse conversion unit 412 converts the metadata into 3D scene parameters, packing information, and compressed neural network weights. For example, the image decoding unit 411 decodes images using a known video codec. The configuration of the image decoding unit 411 may be the same as that of the decoding device 200 described above.

[0213] The NNC decoder 413 generates neural network weights by expanding (decoding) the compressed neural network weights. The inverse transformer 414 transforms the image into a feature plane using the packing information.

[0214] The renderer 415 uses the neural network weights, feature planes, and 3D scene parameters to generate or update a model, and uses the model to generate a rendered image, which is the image seen from an input viewpoint requested by the software or user.

[0215] If the metadata includes URL information indicating a storage device storing encoded data such as neural network weights, the decoding device 410 acquires the encoded data such as compressed neural network weights from the URL indicated by the URL information. The NNC decoding unit 413 generates neural network weights by decoding the acquired compressed neural network weights. In this way, the size of the bitstream can be reduced by transmitting the encoded data of the neural network weights (compressed neural network weights) to the decoding device via a separate means, such as via a URL, without adding them to the bitstream.

[0216] [First Aspect] The decoding process by the decoding device according to the first aspect will be described below. Fig. 32 is a flowchart of the decoding process by the decoding device according to the first aspect.

[0217] First, the decoding device decodes images and metadata from a bitstream (S101). The decoding device also obtains packing information and 3D scene parameters from the metadata.

[0218] Next, the decoding device converts the image into a feature map image (S102). The feature map image has a higher bit depth than the decoded image. For example, the decoding device generates high-bit image data by inverse quantizing low-bit image data. For example, low-bit 10-bit integer values ​​are inverse quantized to generate high-bit 32-bit floating-point values. In another example, 12-bit or 8-bit integer values ​​are inverse quantized to generate 32-bit floating-point values. The inverse quantization process is, for example, a left-shift operation to increase precision. The 3D scene parameters also include a feature plane quantization parameter for determining the inverse quantization step (quantization step). The decoding device performs inverse quantization using the step indicated by the feature plane quantization parameter.

[0219] Next, the decoding device unpacks multiple segments of the feature plane from the feature map image using the packing information (S103).

[0220] Next, the decoding device updates the model using the segments of the feature plane (S104). The model includes the previously acquired feature planes, and the new feature plane is added by the update.

[0221] Next, the decoding device generates a rendering image for the specified viewpoint and time using the updated model (S105).

[0222] Although an example in which the image included in the bitstream is quantized has been shown here, the image included in the bitstream may not be quantized. In this case, the decoding device does not perform step S102, and in step S103, the decoding device unpacks multiple segments of the feature plane from the image obtained from the bitstream.

[0223] Figure 33 shows an example of packing (rearranging) feature planes into multiple images. In this example, the image resolution is determined based on the size of the time-dependent feature plane data. Specifically, time-independent feature plane data, which has a large amount of data, is divided into multiple frames and packed.

[0224] In the example of Fig. 33, there is one piece of time-independent feature plane data and x pieces of time-dependent feature plane data (time-dependent feature plane data 1 to x). The multiple pieces of time-dependent feature plane data correspond to different times or time intervals.

[0225] Here, the time-independent feature planar data is typically larger than the time-dependent feature planar data. Therefore, in the upper part of Figure 33, the image resolution (size) w1 x h1 is set so that the time-dependent feature planar data can be arranged in one frame (image), and the time-independent feature planar data is transmitted as one large frame. Also, x pieces of time-dependent feature planar data are transmitted as separate frames. In this case, each frame of the time-dependent feature planar data includes large padding data.

[0226] On the other hand, in the lower part of Figure 33, the image resolution w2 x h2 is set so that one piece of time-dependent feature planar data can be placed. In this case, the time-independent feature planar data is divided into y frames. Also, x pieces of time-dependent feature planar data are transmitted as separate frames. This reduces padding data included in frames of time-dependent feature planar data. In this case, y + x frames, from frame m to frame m + y + x - 1, are included in the bitstream.

[0227] Note that the bitstream may include information indicating whether the information included in each frame is time-independent feature plane data or time-dependent feature plane data. For example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard. This allows a decoding device to identify whether a decoded frame contains time-independent feature plane data or time-dependent feature plane data. Therefore, the decoding device can appropriately update a model using the decoded time-independent feature plane data or time-dependent feature plane data.

[0228] Figure 34 shows another example of packing feature planes into multiple images. In this example, there is one time-independent feature plane data and x time-dependent feature plane data (time-dependent feature plane data 1 to x), just like in the example shown in Figure 33.

[0229] In the example shown in FIG. 34, the image resolution is set to w1×h1 so that the time-dependent feature planar data can be placed in one frame, and the time-independent feature planar data is transmitted as one large frame.

[0230] Furthermore, x time-dependent feature planes are combined and packed into one frame, which has the advantage of reducing the number of frames compared to the case shown in Figure 33. Also, frame m and frame m+1 are included in the bitstream.

[0231] Note that the bitstream may include information indicating whether the information included in each frame is time-independent feature plane data or time-dependent feature plane data, for example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard.

[0232] Furthermore, the bitstream may include information indicating the number of pieces of time-dependent feature planar data included in one frame and information indicating the mapping start position within the frame of each piece of time-dependent feature planar data (e.g., the start position of data 1 to x shown in FIG. 34 ). For example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard. In this way, when a decoded frame includes multiple pieces of time-dependent feature planar data, the decoding device can appropriately acquire the multiple pieces of time-dependent feature planar data by extracting each piece of time-dependent feature planar data using the information indicating the mapping start position.

[0233] In addition, if the starting position of the first data (e.g., data 1 in Figure 34) is a predetermined position such as (0, 0) in the image, information indicating the starting position of the first data does not need to be included in the bitstream.

[0234] Figure 35 shows another example of packing feature planes into multiple images. In this example, there is one time-independent feature plane data and x time-dependent feature plane data (time-dependent feature plane data 1 to x), just like in the example shown in Figure 33.

[0235] In the example shown in Figure 35, the time-dependent feature planar data includes time-dependent feature planar data at multiple scales (1 to N). Similarly to the example shown in Figure 33, the image resolution w2 x h2 is set so that one piece of time-dependent feature planar data can be arranged. The time-independent feature planar data is divided into y+1 frames. The x pieces of time-dependent feature planar data are transmitted as separate frames.

[0236] Furthermore, in the example shown in Figure 35, the low-scale time-independent feature planar data is encoded first, i.e., placed in the first of the y+1 frames in which time-independent feature planar data will be placed, and the high-scale time-independent feature planar data is split and placed in the last y frames in the y+1 frames in which time-independent feature planar data will be placed.

[0237] Here, the low-scale time-independent feature planar data refers to one or more scales on the lower side of scales 1 to N. The high-scale time-independent feature planar data refers to one or more scales on the higher side of scales 1 to N, for example, scales 1 to N other than the scales included in the low-scale time-independent feature planar data. For example, the low-scale time-independent feature planar data includes scales 1 and 2, and the high-case time-dependent feature planar data includes the remaining scales.

[0238] This has the advantage that the low-scale data contained in the first frame of the 3D scene can be reconstructed quickly, and frames m through m+y+x are included in the bitstream.

[0239] Note that the bitstream may include information indicating whether the information included in each frame is time-independent feature plane data or time-dependent feature plane data, for example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard.

[0240] Furthermore, information indicating the scale of the time-independent feature plane data included in each frame may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. This allows a decoding device, when a decoded frame is time-independent feature plane data, to identify the scale of the time-independent feature plane data of the decoded frame. Thus, the decoding device can appropriately acquire multi-scale time-independent feature plane data.

[0241] Figure 36 shows another example of packing feature planes into multiple images. In the example shown in Figure 36, similar to the example shown in Figure 35, there is one piece of time-independent feature plane data and x pieces of time-dependent feature plane data (time-dependent feature plane data 1 to x). The time-dependent feature plane data includes time-dependent feature plane data at multiple scales (1 to N). Furthermore, the image resolution w2 × h2 is set to be smaller than the resolution w1 × h1.

[0242] In the example shown in Figure 36, the time-independent feature planar data and the time-dependent feature planar data are divided into multiple frames. Furthermore, the low-scale time-independent feature planar data and the time-dependent feature planar data are encoded first. Specifically, the low-scale time-independent feature planar data and the time-dependent feature planar data are divided into u frames. For example, the low-scale time-independent feature planar data is divided into u pieces, and the time-dependent feature planar data 1 to x is also divided into u pieces. Each frame stores one of the u pieces of low-scale time-independent feature planar data and the u pieces of time-dependent feature planar data. Furthermore, the high-scale time-independent feature planar data is divided into v pieces and stored in v frames.

[0243] This allows for fast playback of dynamic scenes by arranging low-scale time-independent feature plane data and time-dependent feature plane data in the previous frame.

[0244] It should be noted that the bitstream may also include information indicating whether the information contained in a frame is time-independent feature plane data, time-dependent feature plane data, or both, for example, this information may be added to metadata, or to the SEI, VUI, or header information of a video coding standard.

[0245] Furthermore, when both types of data are included in one frame, the bitstream may include information indicating the number of pieces of time-independent feature planar data or time-dependent feature planar data included in the frame, and information indicating the mapping start position within the frame for each piece of data (e.g., start position A and start position B shown in FIG. 36 ). For example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard. This allows a decoding device to appropriately acquire multiple pieces of data by using the information indicating the mapping start position when the decoded frame includes multiple pieces of data.

[0246] In addition, if the start position of the first data (for example, start position A in Figure 36) is a predetermined position such as (0, 0) in the image, information indicating the start position of the first data does not need to be included in the bitstream.

[0247] Figure 37 shows another example of packing feature planes into multiple images. In the example shown in Figure 37, two scales 1 and 2 of a time-independent feature plane are relocated to an image in YUV420 format. The feature plane has C channels, and the C channels are located in two frames m and m+1. Here, the time-independent feature plane may be data corresponding to a single feature plane (e.g., any of XY, YZ, and XZ) or data corresponding to multiple feature planes (e.g., two or more of XY, YZ, and XZ).

[0248] Specifically, channels 1 to k of scale 2 are arranged in the Y component of frame m, and channels k+1 to C of scale 2 are arranged in the Y component of frame m+1. Also, channels 1 to k of scale 1 are arranged in the U component of frame m, and channels k+1 to C of scale 1 are arranged in the U component of frame m+1.

[0249] In the YUV420 format, 2x2 pixels in the Y component correspond to one pixel in the U or V component. This is the same as the size relationship between two scales when the scale resolution multiplier is a power of 2. Therefore, the above arrangement allows efficient rearrangement of data at two scales. In addition, multiple channels of a feature plane are arranged adjacent to each other to form a rectangle. This allows parallel decoding using tiles, slices, or subpictures.

[0250] Furthermore, by dividing multiple channels included in a feature plane into multiple frames, the frame resolution can be reduced, which allows, for example, the feature plane to be coded at a resolution lower than that supported by existing image coding standards, thereby enabling the coding of the feature plane to be implemented using existing image coding standards.

[0251] Note that the bitstream may include information indicating the number of dimensions of the channels included in the frame (e.g., information indicating k in the example shown in FIG. 37 ), information indicating the scale of feature plane data included in the frame, or mapping information indicating the scale of data included in each YUV component. For example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard. This allows the decoding device to appropriately acquire feature planes from multiple decoded frames using the mapping information.

[0252] Although FIG. 37 shows an example in which two scales are arranged, the number of scales may be three or more.

[0253] Figure 38 shows another example of packing feature planes into multiple images. In the example shown in Figure 38, four scales 1 to 4 of a time-independent feature plane are rearranged into an image in YUV420 format. The feature plane has C channels, and the C channels are arranged in two frames m and m+1. In this case, the same effect as above can be achieved.

[0254] When five or more scales are arranged, for example, the largest scale is arranged in the Y component, the second largest scale is arranged in one of the U or V components, and the remaining scales are arranged in the other of the U or V components.

[0255] Also, although an example in which C channels are divided and arranged in two frames has been shown here, the C channels may be arranged in one frame, or may be divided and arranged in three or more frames.

[0256] Figure 39 is a diagram showing another example of packing feature planes into multiple images. In the example shown in Figure 39, two scales 1 and 2 of time-independent feature planes are rearranged in an image in YUV420 format. The time-independent feature planes include an XY feature plane, a YZ feature plane, and an XZ feature plane. The XY feature plane, the YZ feature plane, and the XZ feature plane are arranged in different frames.

[0257] In this example, all channels of the feature planes are packed into one frame. Specifically, all channels of scale 2 of the XY feature plane (channels 1 through C) are placed in the Y component of frame m, and all channels of scale 1 of the XY feature plane are placed in the U component of frame m. All channels of scale 2 of the YZ feature plane are placed in the Y component of frame m+1, and all channels of scale 1 of the YZ feature plane are placed in the U component of frame m+1. All channels of scale 2 of the XZ feature plane are placed in the Y component of frame m+2, and all channels of scale 1 of the XZ feature plane are placed in the U component of frame m+2.

[0258] Additionally, multiple channels of a feature plane are arranged adjacent to each other to form a rectangle, allowing for parallel decoding using tiles, slices, or sub-pictures.

[0259] In this way, by storing multiple channels included in one feature plane (XY, YZ, or XZ) in one frame, the number of frames to be coded can be reduced, which allows, for example, the feature plane to be coded at a frame rate lower than that supported by existing image coding standards, thereby enabling coding of the feature plane using existing image coding standards.

[0260] Note that information indicating whether the information included in a frame is an XY feature plane, a YZ feature plane, or an XZ feature plane may be included in the bitstream. For example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard. This allows a decoding device to identify whether the decoded frame is an XY feature plane, a YZ feature plane, or an XZ feature plane. Therefore, the decoding device can appropriately update the model using, for example, the decoded XY feature plane, YZ feature plane, or XZ feature plane.

[0261] Although FIG. 39 shows an example in which two scales are arranged, the number of scales may be three or more.

[0262] Figure 40 is a diagram showing another example of packing feature planes into multiple images. In the example shown in Figure 40, four scales 1 to 4 of time-independent feature planes are rearranged in an image in YUV420 format. The time-independent feature planes include an XY feature plane, a YZ feature plane, and an XZ feature plane. The XY feature plane, the YZ feature plane, and the XZ feature plane are arranged in different frames.

[0263] Specifically, all channels of scale 4 on the XY feature plane (channels 1 to C) are placed in the Y component of frame m, all channels of scale 3 on the XY feature plane are placed in the U component of frame m, and all channels of scales 2 and 1 on the XY feature plane are placed in the V component of frame m. All channels of scale 4 on the YZ feature plane are placed in the Y component of frame m+1, all channels of scale 3 on the YZ feature plane are placed in the U component of frame m+1, and all channels of scales 2 and 1 on the YZ feature plane are placed in the V component of frame m+1. All channels of scale 4 on the XZ feature plane are placed in the Y component of frame m+2, all channels of scale 3 on the XZ feature plane are placed in the U component of frame m+2, and all channels of scales 2 and 1 on the XZ feature plane are placed in the V component of frame m+2. In this case, the same effect as above can be achieved.

[0264] When five or more scales are arranged, for example, the largest scale is arranged in the Y component, the second largest scale is arranged in one of the U or V components, and the remaining scales are arranged in the other of the U or V components.

[0265] Figure 41 shows another example of packing feature planes into multiple images. In the example shown in Figure 41, two time-independent feature planes, scales 1 and 2, are rearranged into an image in YUV420 format. The time-independent feature planes include an XY feature plane, a YZ feature plane, and an XZ feature plane. Each feature plane has C channels.

[0266] In this example, the XY feature plane, the YZ feature plane, and the XZ feature plane are stored in one frame. Specifically, the scale 2 of the XY feature plane, the YZ feature plane, and the XZ feature plane is stored in the Y component, and the scale 1 of the XY feature plane, the YZ feature plane, and the XZ feature plane is stored in the U component.

[0267] Although an example has been described in which all feature planes of the time-independent feature planes are stored in one frame, similarly, all feature planes of the time-dependent feature planes may be stored in one frame.

[0268] Additionally, multiple channels of a feature plane are arranged adjacently to form a rectangle, allowing for parallel decoding using tiles, slices, or subpictures. Additionally, in this example, the feature plane data is placed on the left side of the image, speeding up decoding.

[0269] Note that information indicating whether the information included in a frame is XY, YZ, XZ, or a combination thereof may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. Furthermore, when a frame includes XY, YZ, and XZ, information indicating the mapping start position within the frame of each feature plane data (e.g., start positions A to F shown in FIG. 41) may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. In this way, when a decoded frame includes XY, YZ, and XZ, the decoding device can appropriately acquire data of the XY feature plane, the YZ feature plane, and the XZ feature plane using the information indicating the mapping start position.

[0270] In addition, if the start position of the first data (for example, start position A and start position D in Figure 41) is a predetermined position such as (0, 0) in the image, information indicating the start position of the first data does not need to be included in the bitstream.

[0271] Note that the method of mapping XY, YZ, and XZ is not limited to the above. For example, mapping may be performed in a different order, such as YZ, XZ, and XY, starting from the top of the image. Furthermore, the start position does not necessarily have to be a pixel at the edge of the frame, and may be set to any position. This allows the encoding device to flexibly map multiple feature planes. For example, the encoding device can reduce the amount of data by performing mapping so as to increase the encoding efficiency of image encoding.

[0272] Figure 42 shows another example of packing feature planes into multiple images. In the example shown in Figure 42, time-dependent feature planes are rearranged into one frame in YUV420 format. The time-dependent feature planes include an XT feature plane, a YT feature plane, and a ZT feature plane. The XT feature plane, the YT feature plane, and the ZT feature plane are stored in the Y component. Each feature plane also includes N scales, from scale 1 to scale N.

[0273] The number of channels C is expressed as p×q. In this example, p=2, q=8, and the number of channels C is p×q=16. In this case, the height of the arranged scale N is k×q, and the width is N×W. x ×p, where k corresponds to the temporal resolution of one channel and the number of frames included in the 3D scene. N is a scale resolution multiplier corresponding to the maximum scale N. Note that if further lower scales are included, they may be placed in the U or V component.

[0274] Note that information indicating whether the information included in a frame is XT, YT, XT, or a combination thereof may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. Furthermore, when a frame includes XT, YT, and XT, information indicating the mapping start position within the frame of each feature plane (e.g., start positions A to C shown in FIG. 42) may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. In this way, when a decoded frame includes XT, YT, and XT, the decoding device can appropriately acquire data of the XT feature plane, the YT feature plane, and the XT feature plane using the information indicating the mapping start position.

[0275] Note that the method of mapping XT, YT, and XT is not limited to the above. For example, mapping may be performed in a different order, such as YT, ZT, and XT, starting from the top of the image. Furthermore, the start position does not necessarily have to be a pixel at the edge of the frame, and may be set to any position. This allows the encoding device to flexibly map multiple feature planes. For example, the encoding device can reduce the amount of data by performing mapping so as to increase the encoding efficiency of image encoding.

[0276] Figure 43 shows another example of packing feature planes into multiple images. In the example shown in Figure 43, a time-independent feature plane and a time-dependent feature plane are rearranged in one frame of the YUV420 format. The time-independent feature plane includes two scales, 1 and 2. The time-independent feature plane includes an XY feature plane, a YZ feature plane, and an XZ feature plane. The time-dependent feature plane includes an XT feature plane, a YT feature plane, and a ZT feature plane. The time-independent feature plane of scale 2 and the time-independent feature plane are stored in the Y component. The time-independent feature plane of scale 1 is stored in the Y component.

[0277] Although FIG. 43 shows an example in which a time-independent feature plane including two scales is arranged, the number of scales may be three or more.

[0278] Furthermore, information indicating whether the information included in a frame is XY, YZ, XZ, XT, YT, XT, or a combination thereof may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. Furthermore, when a frame includes XY, YZ, XZ, XT, YT, and XT, information indicating mapping start position information within the frame for each feature plane may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. In this way, when a decoded frame includes XY, YZ, XZ, XT, YT, and XT, the decoding device can appropriately acquire data of the feature planes of XY, YZ, XZ, XT, YT, and ZT using the information indicating the mapping start position.

[0279] Note that the method of mapping XY, YZ, XZ, XT, YT, and XT is not limited to the above. For example, mapping may be performed in a different order, such as YZ, YT, XZ, ZT, XY, and XT, starting from the top of the image. Furthermore, the start position does not necessarily have to be a pixel at the edge of the frame, and may be set to any position. This allows the encoding device to flexibly map multiple feature planes. For example, the encoding device can reduce the amount of data by performing mapping so as to increase the encoding efficiency of image encoding.

[0280] Figure 44 shows an example of unpacking time-independent feature plane data from a feature map image into a feature plane in step S103 shown in Figure 32. Multiple rectangular segments of the feature plane are stacked to form a feature plane of C channels. This process is performed for each scale. Figure 44 shows an example for scale 1 and scale N.

[0281] The rectangular segments are received over one or more frames. In this example, the received C channels of feature plane data are received as p columns and q rows of rectangular image segments, where p×q=C.

[0282] Figure 45 shows an example of unpacking segments of time-dependent feature plane data into feature planes in step S103 shown in Figure 32. In this example, the time-dependent feature plane includes scales 1 to N. Each scale includes channels 1 to C. Multiple rectangular segments of the feature plane are stacked to form a feature plane with C channels.

[0283] Here, k in Figure 45 corresponds to the number of frames included in the 3D scene. That is, a small k has the advantage that the 3D scene can be decoded and played back quickly, while a large k allows for encoding a 3D scene of longer duration.

[0284] Figure 46 is a diagram showing an example of updating a model of a dynamic 3D representation in step S104 shown in Figure 32. As shown in Figure 46, the received time-independent feature plane data is used to update the time-independent feature planes included in the model. The received time-dependent feature plane data is used to update the time-dependent feature planes included in the model. The received neural network weights are used to update the neural network weights included in the model.

[0285] Note that the model update may be performed across multiple scales at once or scale by scale, e.g., lower scales may be updated first, allowing for faster reconstruction of the first frame of the 3D scene. For example, this model update may be performed by the rendering unit 405 or 415.

[0286] Figure 47 is a diagram showing an example of model updating by the decoding device using time-dependent feature plane data in step S104 shown in Figure 32. As shown in Figure 47, the model is updated by adding newly received time-dependent feature plane data to previously received time-dependent feature plane data. In this way, the model is updated in stages by the rendering unit. This has the advantage of updating the initial frame first, allowing for faster playback of the dynamic 3D scene.

[0287] Note that the models based on model data previously received by the decoding device may include models initialized with default values.

[0288] Figure 48 shows another example of packing feature planes into multiple images. In this example, multiple segments of the feature plane are classified according to their interest level. Then, in the initial frame, segments of high interest (or foreground) regions are transmitted with priority over segments of low interest (or background). In this example, information about high interest regions is stored in the lower left corner of the feature plane.

[0289] Specifically, the segments with the highest level of interest are arranged in the first frame M, the segments with the second highest level of interest are arranged in the subsequent frame M+1, the segments with the third highest level of interest are arranged in the subsequent frames M+2 and M+3, and the segments with the lowest level of interest are arranged in the subsequent frames M+4 and M+5. In this way, the segments with each level of interest are arranged together in one or more frames. Furthermore, the frame in which the segment with the highest level of interest is arranged is encoded (transmitted) first. This has the advantage that the decoding device can more quickly reconstruct, for example, a specific area of ​​a 3D scene that a user is interested in seeing.

[0290] The following describes a coding device that generates a bitstream to be decoded by the decoding device according to the first aspect. Fig. 49 is a flowchart of the coding process performed by the coding device according to the first aspect.

[0291] First, the encoding device extracts a plurality of segments of feature planes from a model including a plurality of feature planes (S201).

[0292] Next, the encoding device generates a feature map image by packing (rearranging) the segments of the feature plane. The encoding device also stores the rearrangement order and start addresses of the segments of the feature plane in metadata (S202). At this time, the encoding device, for example, places padding data in areas of the image other than the feature plane data.

[0293] Next, the encoding device converts the high-bit feature map image into a low-bit-depth image by quantizing it (S203). The low-bit-depth image has a lower bit depth than the feature map image.

[0294] For example, the encoding device may quantize high-bit 32-bit floating-point data to obtain 10-bit integer values. Alternatively, the encoding device may quantize 32-bit floating-point values ​​to obtain 12-bit or 8-bit integer values. The quantization process may include, for example, a process of reducing precision by a right-shift operation. The 3D scene parameters may also include a feature plane quantization parameter for determining a quantization step.

[0295] Next, the encoding device encodes the low bit-depth image and the metadata into a bitstream (S204), that is, the bitstream includes the low bit-depth image and the metadata.

[0296] Although an example in which the feature map image is quantized has been described above, the feature map image does not have to be quantized. In this case, the encoding device does not perform step S203, and instead encodes the feature map image in step S204.

[0297] Fig. 50 is a diagram showing an example of packing a feature plane into a feature map image. Multiple channels 1 to C of the feature plane shown in Fig. 50 are arranged in one low bit-depth image. The encoding device also generates packing information related to the packing and stores the generated packing information in the bitstream.

[0298] As shown in Fig. 50, the packing information includes an ID (XY_N) that identifies a feature plane. This ID indicates the type of feature plane (XY, YZ, XZ, XT, YT, ZT) and the scale. For example, (XY_N) indicates that the data in the frame is data of scale N on the XY feature plane.

[0299] Furthermore, (u, v) indicates the start position of the data within the frame, and in this example, indicates the pixel position (coordinates) as the start position. In this example, start position P11 shown in Figure 50 is indicated.

[0300] (1, C) indicates the number of channels included in the data. In this example, it is indicated that the data includes channel 1 to channel C.

[0301] (p, q) indicates the arrangement of data for multiple channels, specifically, the multiple channels are arranged in p rows and q columns.

[0302] (0) indicates the arrangement order of multiple channels, and in this example, indicates that the arrangement order is the raster scan order.

[0303] Fig. 51 is a diagram showing an example of dividing a time-dependent feature plane into segments in step S202 of Fig. 49. As shown in Fig. 51, the time-dependent feature plane is divided into a plurality of segments. Note that, although an example is shown in which the shapes of the plurality of segments after division are the same, the shapes of the plurality of segments after division may be different.

[0304] As described above, the encoding device according to the first aspect transmits time-dependent feature plane data in one frame, and divides large feature plane data into multiple segments before transmitting them. The decoding device quickly reconstructs a 3D scene by updating a model using the received segments. The encoding device also prioritizes transmitting an initial set of time-dependent feature planes and feature plane segments belonging to lower scales of time-independent feature planes. The encoding device also combines multiple scales of time-dependent or time-independent feature planes into one frame.

[0305] This improves the robustness of the reconstruction of the 3D scene and reduces the amount of processing required by the decoder. Furthermore, the low latency allows the decoder to quickly reconstruct the first frame of the 3D scene. Furthermore, the low latency allows the decoder to start playing the 3D scene earlier, thereby providing a higher quality immersive experience.

[0306] This aspect may be implemented by combining at least a part of other aspects of the present disclosure. Furthermore, this aspect may be implemented by combining a part of the process, a part of the configuration of any device, a part of the syntax, etc., described in any of the aspects with other aspects.

[0307] Note that some or all of the processes in the decoding device may be similarly executed in the encoding device. Also, not all of the components described in this aspect are necessarily required, and the device may include only some of the components of the first aspect.

[0308] Although the above description has been given of an example in which a feature plane is placed on an image, the data to be placed is not limited to this. Any information contained in the model other than the feature plane may be similarly placed on the image. For example, this information may be information indicating a feature quantity having a time dimension.

[0309] In addition, although the above description has been given with reference to an example of YUV components as the multiple components contained in an image, the multiple components are not limited to this. For example, the multiple components may be YCbCr, YPbPr, RGB, or the like.

[0310] As described above, the encoding device according to the embodiment performs the processing shown in Fig. 52. The encoding device acquires first information (e.g., multiple feature planes) included in a model, the first information indicating multiple feature quantities having a time dimension (S301), arranges the first information on one or more images (S302), and generates a bitstream by encoding the one or more images (S303). This allows the encoding device to encode the first information, which indicates multiple feature quantities having a time dimension and is included in the model, using image encoding. This makes it easy to encode the information included in the model.

[0311] For example, the first information includes a plurality of data (e.g., time-independent feature plane data and a plurality of time-dependent feature plane data), and the bitstream includes second information indicating the arrangement order of the plurality of data, so that the decoding device can appropriately obtain the plurality of data from the image using the second information.

[0312] For example, the first information includes a plurality of data (e.g., time-independent feature plane data and a plurality of time-dependent feature plane data), and the bitstream includes third information indicating a starting position of at least one of the plurality of data in one or more images, such that a decoding device can use the third information to appropriately retrieve the plurality of data from the images.

[0313] For example, the first information includes a plurality of feature planes each representing a plurality of feature quantities arranged two-dimensionally. This allows the encoding device to encode the information of the plurality of feature planes included in the model using image encoding. This facilitates the encoding process of the information included in the model.

[0314] For example, as shown in Fig. 33, the first information includes first data (e.g., time-independent feature plane data) including multiple feature quantities without a time dimension, and multiple second data (e.g., time-dependent feature plane data) including multiple feature quantities with a time dimension, each corresponding to a different time interval. The one or more images include multiple first images and multiple second images, and the first data is divided and arranged among the multiple first images, and the multiple second data is respectively arranged among the multiple second images. This reduces padding data included in the one or more images, thereby improving encoding efficiency.

[0315] For example, as shown in Fig. 34, the first information includes first data (e.g., time-independent feature plane data) including multiple feature quantities without a time dimension, and multiple second data (e.g., time-dependent feature plane data) including multiple feature quantities with a time dimension, each corresponding to a different time interval. The one or more images include a first image and a second image, where the first data is arranged in the first image and the multiple second data are arranged together in the second image. This reduces the number of images to be encoded, thereby improving encoding efficiency.

[0316] For example, as shown in Figures 35 and 36, the first information includes first data (low-scale time-independent feature plane data) and second data (high-scale time-independent feature plane data) having a higher resolution than the first data. The one or more images include the first image and a second image that is encoded after the first image. The first data is arranged in the first image, and the second data is arranged in the second image. This allows the decoding device to first decode the first data and generate a model using the first data.

[0317] For example, as shown in Figures 37 to 41, the first information includes first data (low-scale time-independent feature planar data) and second data (high-scale time-independent feature planar data) having a higher resolution than the first data. Each of the one or more images includes a first component (e.g., a Y component) and second and third components (e.g., a U component and a V component) that are smaller than the first component. The first data is arranged in at least one of the second and third components, and the second data is arranged in the first component. This allows multiple data to be arranged efficiently using multiple components.

[0318] The configuration of the encoding device is the same as that of the encoding device shown in Fig. 8. The encoding device includes a processor (or circuit) and a memory, and the processor (or circuit) performs the above-mentioned processing using the memory during operation.

[0319] Furthermore, a decoding device according to an embodiment performs the process shown in Fig. 53. The decoding device decodes one or more images from a bitstream generated by encoding one or more images in which first information (e.g., multiple feature planes) included in a model, the first information indicating multiple feature quantities having a time dimension, is arranged (S401), and acquires the first information from the one or more images (S402). This allows the decoding device to decode the first information included in the model and indicating multiple feature quantities having a time dimension using a decoding process based on image encoding. This makes it possible to easily realize the decoding process of the information included in the model.

[0320] For example, the first information includes a plurality of data (e.g., time-independent feature plane data and a plurality of time-dependent feature plane data), and the bitstream includes second information indicating the arrangement order of the plurality of data, so that the decoding device can appropriately obtain the plurality of data from the image using the second information.

[0321] For example, the first information includes a plurality of data (e.g., time-independent feature plane data and a plurality of time-dependent feature plane data), and the bitstream includes third information indicating a starting position of at least one of the plurality of data in one or more images, such that a decoding device can use the third information to appropriately retrieve the plurality of data from the images.

[0322] For example, the first information includes a plurality of feature planes each representing a plurality of feature quantities arranged two-dimensionally. This allows the decoding device to decode the information of the plurality of feature planes included in the model using a decoding process based on image coding. This makes it possible to easily realize the decoding process of the information included in the model.

[0323] For example, as shown in Fig. 33, the first information includes first data (e.g., time-independent feature plane data) including multiple feature quantities without a time dimension, and multiple second data (e.g., time-dependent feature plane data) including multiple feature quantities with a time dimension, each corresponding to a different time interval. The one or more images include multiple first images and multiple second images, and the first data is divided and arranged among the multiple first images, and the multiple second data is respectively arranged among the multiple second images. This reduces padding data included in the one or more images, thereby improving encoding efficiency.

[0324] For example, as shown in Fig. 34, the first information includes first data (e.g., time-independent feature plane data) including multiple feature quantities without a time dimension, and multiple second data (e.g., time-dependent feature plane data) including multiple feature quantities with a time dimension, each corresponding to a different time interval. The one or more images include a first image and a second image, where the first data is arranged in the first image and the multiple second data are arranged together in the second image. This reduces the number of images to be encoded, thereby improving encoding efficiency.

[0325] For example, as shown in Figures 35 and 36, the first information includes first data (low-scale time-independent feature plane data) and second data (high-scale time-independent feature plane data) having a higher resolution than the first data. The one or more images include the first image and a second image that is encoded after the first image. The first data is arranged in the first image, and the second data is arranged in the second image. This allows the decoding device to first decode the first data and generate a model using the first data.

[0326] For example, as shown in Figures 37 to 41, the first information includes first data (low-scale time-independent feature planar data) and second data (high-scale time-independent feature planar data) having a higher resolution than the first data. Each of the one or more images includes a first component (e.g., a Y component) and second and third components (e.g., a U component and a V component) that are smaller than the first component. The first data is arranged in at least one of the second and third components, and the second data is arranged in the first component. This allows multiple data to be arranged efficiently using multiple components.

[0327] The configuration of the decoding device is, for example, the same as that of the decoding device 200 shown in Fig. 11. The decoding device includes a processor (or circuit) and a memory, and the processor (or circuit) performs the above processing using the memory during operation.

[0328] [Second Aspect] The following describes the decoding process by the decoding device according to the second aspect. Fig. 54 is a flowchart of the decoding process by the decoding device according to the second aspect.

[0329] First, the decoding device decodes images and metadata from a bitstream (S501). The decoding device also obtains packing information and 3D scene parameters from the metadata.

[0330] Next, the decoding device converts the image into a feature map image (S502). The feature map image has a higher bit depth than the decoded image. For example, the decoding device generates high-bit image data by inverse quantizing low-bit image data. For example, low-bit 10-bit integer values ​​are inverse quantized to generate high-bit 32-bit floating-point values. In another example, 12-bit or 8-bit integer values ​​are inverse quantized to generate 32-bit floating-point values. The inverse quantization process is, for example, a left-shift operation to increase precision. The 3D scene parameters also include a feature plane quantization parameter for determining the inverse quantization step (quantization step). The decoding device performs inverse quantization using the step indicated by the feature plane quantization parameter.

[0331] Next, the decoding device unpacks the plurality of segments of the feature plane and the plurality of neural network weights from the feature map image using the packing information (S503).

[0332] Next, the decoder updates the model using the segments of the feature planes and the neural network weights (S504). The model includes the previously obtained feature planes, and new feature planes are added by the update.

[0333] Next, the decoding device generates a rendering image for the specified viewpoint and time using the updated model (S505).

[0334] Although the example shown here shows a case where the image included in the bitstream is quantized, the image included in the bitstream does not have to be quantized. In this case, the decoding device does not perform step S502, and in step S503, unpacks multiple segments of the feature plane and multiple neural network weights from the image obtained from the bitstream.

[0335] 55 is a block diagram showing the configuration of an encoding device according to an embodiment. The encoding device 500 receives as input a model of a 3D scene representation method trained for one of the 3D scenes in a 3D video sequence. The model includes a feature plane and neural network weights.

[0336] The encoding device 500 comprises a transform unit 501, a metadata transform unit 502, and an image encoding unit 503. The transform unit 501 extracts feature planes, neural network weights, and associated 3D scene parameters from the model.

[0337] The conversion unit 501 also rearranges (packs) the extracted feature planes and neural network weights into an image, and outputs the generated image to the image encoding unit 503. The conversion unit 501 also outputs packing information related to the rearrangement to the metadata conversion unit 502.

[0338] The packing information includes information for unpacking the neural network weights from the frame, such as at least one of a starting address of the neural network weights, a number of neural network weights, a number of neural network layers, a type of neural network layer, a number of neurons per layer, and a neural network identifier.

[0339] The conversion unit 501 also updates the 3D scene parameters required for reconstructing the 3D scene in the decoding device, and outputs the updated 3D scene parameters to the metadata conversion unit 502 .

[0340] The metadata conversion unit 502 converts the 3D scene parameters and packing information into metadata (parameter sets) that are transmitted together with the bitstream.

[0341] The image encoding unit 503 generates a bitstream by encoding images and metadata. That is, the bitstream includes images (encoded images) and metadata. For example, the image encoding unit 503 encodes images using a known video codec. The configuration of the image encoding unit 503 may be the same as that of the encoding device 100 described above.

[0342] 56 is a block diagram showing the configuration of a decoding device according to an embodiment. A bitstream is input to the decoding device 600. For example, the bitstream is generated by the encoding device 500 described above.

[0343] The decoding device 600 includes an image decoding unit 601 , a metadata inverse conversion unit 602 , an inverse conversion unit 603 , and a rendering unit 604 .

[0344] The image decoding unit 601 decodes the bitstream to obtain images and metadata. The metadata inverse conversion unit 602 converts the metadata into 3D scene parameters and packing information. For example, the image decoding unit 601 decodes images using a known video codec. The configuration of the image decoding unit 601 may be the same as that of the decoding device 200 described above, for example.

[0345] The inverse transform unit 603 uses the packing information to transform the image into a feature plane and neural network weights.

[0346] The renderer 604 uses the neural network weights, feature planes, and 3D scene parameters to generate or update a model, and uses the model to generate a rendered image, which is the image seen from an input viewpoint requested by the software or user.

[0347] Figure 57 shows an example of packing (rearranging) feature planes and neural network weights into an image. Frame m is an image decoded from the bitstream.

[0348] In step S503 shown in Fig. 54, the decoding device unpacks the neural network weights from the frame. The decoding device generates a neural network using the packing information and the neural network weights included in the metadata.

[0349] Note that the mapping method is not limited to the example shown in FIG. 57 . For example, the neural network weights may be mapped to the upper left pixel in the image. Furthermore, the starting position does not necessarily have to be a pixel at the edge of the frame, and the starting position may be set to any position. This allows the encoding device to flexibly map multiple feature planes. For example, the encoding device can reduce the amount of data by performing mapping so as to improve the encoding efficiency of image encoding.

[0350] This allows the encoding device to transmit additional data using unused regions within a frame. Because high accuracy of the neural network weights is required, the encoding device may apply lossless or low quantization (low QP) to the frame, slice, tile, or coding unit containing the neural network weights. Note that when a high QP is used, the encoding device may ensure high accuracy by adding redundancy to the neural network weights. For example, forward error correction may be used as a method for adding redundancy. For example, the neural network weights may be stored three times, with the accuracy being determined based on the similarity of at least two of the three iterations in the decoding device.

[0351] Note that information indicating whether a frame includes neural network weights may be included in the bitstream. For example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard. Furthermore, if a frame includes neural network weights, information indicating the number of neural network weights included in one frame and information indicating a mapping start position within the frame of the neural network weights (e.g., start position A shown in FIG. 57 ) may be included in the bitstream. For example, this information may be added to metadata, or SEI, VUI, or header information of an image coding standard. In this way, when a decoded frame includes one or more neural network weights, the decoding device can appropriately acquire one or more neural network weights using the information indicating the mapping start position.

[0352] Figure 58 shows another example of packing feature planes and neural network weights into an image. In the example shown in Figure 58, frame m includes slice 1 and slice 2. Slice 1 contains feature plane data, and slice 2 contains neural network weights. In this way, the feature plane and the neural network weights may be stored in different slices included in a single frame. Note that tiles may be used instead of slices. This allows the feature plane and the neural network weights to be coded and decoded in parallel.

[0353] Figure 59 shows another example of packing feature planes and neural network weights into an image, where the feature planes and neural network weights are stored in different frames.

[0354] Specifically, frame m is an image decoded from the bitstream and contains only neural network weight data. Frame m+1 is an image decoded from the bitstream and contains only feature plane data (e.g., time-independent feature plane data and time-dependent feature plane data). In step S503 shown in Figure 54, the decoding device obtains neural network weights from the frame (feature map image) and generates a neural network using the obtained neural network weights.

[0355] Note that information indicating the area of ​​the padding data may be included in the bitstream. For example, this information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. This allows the decoding device to identify which areas in the decoded frame are padding data, and thus to appropriately obtain neural network weights or feature plane data.

[0356] In this way, the feature plane and the neural network weights are stored in separate frames. This allows the encoding device to apply appropriate processing to the feature plane and the neural network weights on a frame-by-frame basis. The encoding device can, for example, apply different quantization processing to each frame. For example, the encoding device can apply lossless quantization or low quantization (low QP) to the neural network weight frames and high quantization (high QP) to the frames containing the feature plane data.

[0357] The bitstream may include information indicating whether the information included in the frame is neural network weights or feature plane data. The information may indicate whether the information included in the frame is time-independent feature plane data, time-dependent feature plane data, or both. For example, the information may be added to metadata, or SEI, VUI, or header information of an image coding standard. This allows the decoding device to identify whether the information included in the decoded frame is neural network weights or feature plane data, and appropriately acquire the neural network weights or feature plane data.

[0358] Figure 60 shows another example of packing feature planes and neural network weights into an image. In the example shown in Figure 60, the feature planes and neural network weights are stored in an image in YUV420 format.

[0359] Specifically, feature plane data is stored in the Y component of a frame, and neural network weights are stored in the U component. Note that the neural network weights may be placed in the V component, or in both the U and V components. In step S503 shown in FIG. 54, the decoding device unpacks the neural network weights from the frame and uses the unpacked neural network weights to generate a neural network to be included in the model.

[0360] The mapping method is not limited to the example shown in FIG. 60 . For example, neural network weights may be placed in the Y component, and feature plane data may be placed in the U or V component. This allows the encoding device to flexibly map each piece of information. For example, the encoding device can reduce the amount of data by performing mapping so as to increase the encoding efficiency of image encoding.

[0361] Furthermore, this method can efficiently use the U or V component contained in a frame, allowing more data to be transmitted simultaneously.

[0362] The bitstream may include information indicating the information (feature plane or neural network weights) contained in each of the Y, U, and V components of a frame. The information may indicate whether the information contained in the component is time-independent feature plane data, time-dependent feature plane data, or both. For example, the information may be added to metadata, or to the SEI, VUI, or header information of an image coding standard. For example, in the example shown in FIG. 60 , the bitstream may include information indicating that the Y component contains feature plane data (time-independent feature plane data and time-dependent feature plane data) and the U component contains neural network weights. This allows the decoding device to identify which components in the decoded frame contain neural network weights or feature plane data, thereby enabling appropriate acquisition of the neural network weights or feature plane data.

[0363] Figure 61 is a diagram showing another example of packing feature planes and neural network weights into an image. In the example shown in Figure 61, multiple frames included in a bitstream are classified into an enhancement layer and a base layer. For example, the base layer can be decoded without reference to other layers, while the base layer must be referenced to decode the enhancement layer.

[0364] In the example shown in Figure 61, the base layer frames contain only feature plane data (time-independent feature plane data and time-dependent feature plane data), and the enhancement layer (spatial layer) frames contain only neural network weight data. Furthermore, the base layer frames containing feature plane data and the enhancement layer frames containing neural network weights are frames at the same time t.

[0365] In step S503 shown in FIG. 54, the decoding device unpacks the neural network weights from the frame and generates a neural network using the neural network weights.

[0366] Note that the mapping method is not limited to the example shown in FIG. 61 . For example, neural network weights may be stored in the base layer, and feature plane data may be stored in the enhancement layer. This allows the encoding device to flexibly map each piece of information. For example, the encoding device can reduce the amount of data by performing mapping so as to improve the encoding efficiency of image encoding.

[0367] This method also allows multiple layers of video to be used to transmit both feature planes and neural network weights.

[0368] Note that information indicating information (feature plane data or neural network weights) included in the base layer or enhancement layer may be included in the bitstream. Note that the information may indicate whether the information included in the layer is time-independent feature plane data, time-dependent feature plane data, or both. For example, the information may be added to metadata, or SEI, VUI, or header information of an image coding standard. For example, in the example shown in FIG. 61 , information indicating that the base layer includes feature plane data (time-independent feature plane data and dependent feature plane data) and the enhancement layer includes neural network weights may be included in the bitstream. This allows the decoding device to identify whether each layer in a decoded frame includes neural network weights or feature plane data, thereby enabling appropriate acquisition of neural network weights or feature plane data.

[0369] Although the above describes an example in which neural network weights are placed on an image, at least one of other types of information, such as voxel grid data, 3D Gaussian distribution data, or point cloud data, may be placed on an image in a similar manner instead of or in addition to the neural network weights. Furthermore, a combination of these different types of information may be placed on the same frame.

[0370] Figure 62 is a diagram showing an example of updating a model of a dynamic 3D representation in step S504 shown in Figure 54. As shown in Figure 62, the time-independent feature plane included in the model is updated using the received time-independent feature plane data. The time-dependent feature plane included in the model is updated using the received time-dependent feature plane data. The neural network weights included in the model are updated using the received neural network weights. The voxel grid included in the model is updated using the received voxel grid information. For example, this model update is performed by the rendering unit 604.

[0371] FIG. 63 is a diagram showing an example of updating the model of dynamic 3D expression with neural network weights in step S504.

[0372] The neural network is updated using the received neural network weights, for example, by a rendering unit.

[0373] The received neural network weights are: (1) a list of neural network weights [w 11 , w 12 . . . w 1N , b 11 ,... b L2 ], (2) the number of layers L in the neural network, (3) a list of the number of neurons in each layer [1, N..., 2], (4) the type of neural network layer (e.g., convolutional layer or fully connected layer, etc.), and (5) a neural network identifier (e.g., NN_1, etc.).

[0374] In the example shown in Figure 63, the neural network includes L layers: Layer 1 includes one neuron, Layer 2 includes N neurons, and Layer L includes two neurons.

[0375] The received neural network weight information is copied to construct the neural network, and the layer type (e.g., convolutional layer or fully connected layer) is signaled so that the decoding device can understand the characteristics of the layer to be decoded, for example, among multiple layers each having different functions.

[0376] FIG. 64 is a diagram showing an example of the process of updating the dynamic 3D representation model in step S504 shown in FIG. 54 using voxel grid information.

[0377] The voxel grid information includes (1) a list of features of the voxel grid (e.g., the data value of each node) and (2) the dimensions of the voxel grid along the X, Y, and Z directions (e.g., [w x , w y , w z ]), and (3) a voxel grid identifier (e.g., VG_1).

[0378] The received voxel grid information is used to update the nodes of the voxel grid. For example, the features of each node included in the voxel grid are trained feature arrays composed of floating-point values. The features of the node can be used to calculate detailed information about the 3D scene in the region of the voxel grid, such as density or color. The voxel grid identifier is used to identify the representation type (such as density or color) of the voxel grid. That is, the voxel grid identifier indicates the representation type of the voxel grid.

[0379] The following describes a coding device that generates a bitstream to be decoded by the decoding device according to the second aspect. Figure 65 is a flowchart of the coding process performed by the coding device according to the second aspect.

[0380] First, the encoding device extracts a plurality of segments of feature planes from a model including a plurality of feature planes (S601).

[0381] Next, the encoding device generates a feature map image by packing (rearranging) the multiple segments of the feature plane and the neural network weights. The encoding device also stores (i) the rearrangement order and start addresses of the multiple segments of the feature plane and (ii) the rearrangement order and start addresses of the neural network weights in metadata (S602). At this time, the encoding device, for example, places padding data in areas of the image other than the feature plane data and the neural network weights.

[0382] Next, the encoding device converts the high-bit feature map image into a low-bit-depth image by quantizing it (S603). The low-bit-depth image has a lower bit depth than the feature map image.

[0383] For example, the encoding device may quantize high-bit 32-bit floating-point data to obtain 10-bit integer values. Alternatively, the encoding device may quantize 32-bit floating-point values ​​to obtain 12-bit or 8-bit integer values. The quantization process may include, for example, a process of reducing precision by a right-shift operation. The 3D scene parameters may also include a feature plane quantization parameter for determining a quantization step.

[0384] Note that because the neural network weights need to be highly accurate, the encoder may apply lossless or low quantization (low QP) to frames containing the neural network weights, and if a high QP is used, the encoder may ensure high accuracy by incorporating redundancy into the neural network weights.

[0385] Next, the encoding device encodes the low bit-depth image and the metadata into a bitstream (S604), that is, the bitstream includes the low bit-depth image and the metadata.

[0386] Although an example in which the feature map image is quantized has been shown here, the feature map image does not have to be quantized. In this case, the encoding device does not perform step S603, and instead encodes the feature map image in step S604.

[0387] As described above, the encoding device according to the second aspect can efficiently encode neural network weight data using an existing video codec without requiring any other codec-related techniques or devices. The encoding device also replaces padding data with useful data, such as neural network weight data, within a frame. The encoding device can also encode both feature planes and neural network weights using a single video codec. This allows for efficient encoding and decoding processes.

[0388] This aspect may be implemented by combining at least a part of other aspects of the present disclosure. Furthermore, this aspect may be implemented by combining a part of the process, a part of the configuration of any device, a part of the syntax, etc., described in any of the aspects with other aspects.

[0389] Note that some or all of the processes in the decoding device may be similarly executed in the encoding device. Also, not all of the components described in this aspect are necessarily required, and the device may include only some of the components of the first aspect.

[0390] Although the above description has been given of an example in which a feature plane is placed on an image, the data to be placed is not limited to this. Any information contained in the model other than the feature plane may be similarly placed on the image. For example, this information may be information indicating a feature quantity having a time dimension.

[0391] In the above description, an example was given in which both the feature plane and the neural network weights are arranged in the image, but the feature plane may not be arranged in the image and may be encoded using a different method.

[0392] In addition, although the above description has been given with reference to an example of YUV components as the multiple components contained in an image, the multiple components are not limited to this. For example, the multiple components may be YCbCr, YPbPr, RGB, or the like.

[0393] As described above, the encoding device according to the embodiment performs the process shown in Fig. 66. The encoding device acquires multiple neural network weights included in a model (S701), places the multiple neural network weights on one or more images (S702), and generates a bitstream by encoding the one or more images (S703). This allows the encoding device to encode the multiple neural network weights included in the model using image encoding. This makes it easy to encode the information included in the model.

[0394] For example, the bitstream includes first information indicating the arrangement order of the neural network weights, allowing the decoding device to appropriately obtain the neural network weights from the image using the first information.

[0395] For example, the bitstream includes metadata, and the metadata includes the first information and a plurality of parameters (e.g., 3D scene parameters) related to the model, so that the decoding device can generate the model using the plurality of parameters.

[0396] For example, the encoding device further acquires second information (e.g., a feature plane) included in the model, the second information indicating a plurality of feature quantities having a time dimension, and in the arrangement, arranges the plurality of neural network weights and the second information on one or more images. In this way, the encoding device can encode the second information indicating the plurality of feature quantities having a time dimension included in the model using image encoding. This makes it easy to encode the information included in the model.

[0397] For example, as shown in Figures 57 and 58, the neural network weights and the second information are arranged in the same image included in one or more images, which allows the neural network weights and the second information to be arranged efficiently.

[0398] For example, as shown in Fig. 59, the plurality of neural network weights and the second information are arranged in different images included in one or more images, so that the decoding device can independently decode the plurality of neural network weights and the second information.

[0399] For example, as shown in Fig. 60, each of the one or more images includes a plurality of components (e.g., a Y component, a U component, and a V component), and the plurality of neural network weights and the second information are arranged on different components among the plurality of components, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0400] For example, as shown in Fig. 61, each of the one or more images includes multiple layers (e.g., an enhancement layer and a base layer), and the multiple neural network weights and the second information are arranged in different layers among the multiple layers, thereby enabling the multiple neural network weights and the second information to be arranged efficiently.

[0401] The configuration of the encoding device is the same as that of the encoding device shown in Fig. 8. The encoding device includes a processor (or circuit) and a memory, and the processor (or circuit) performs the above-mentioned processing using the memory during operation.

[0402] Furthermore, a decoding device according to an embodiment performs the process shown in Fig. 67. The decoding device decodes one or more images from a bitstream generated by encoding one or more images to which multiple neural network weights included in a model are assigned (S801), and obtains multiple neural network weights from the one or more images (S802). This allows the decoding device to decode the multiple neural network weights included in the model using a decoding process based on image encoding. This makes it easy to decode information included in the model.

[0403] For example, the bitstream includes first information indicating the arrangement order of the neural network weights, allowing the decoding device to appropriately obtain the neural network weights from the image using the first information.

[0404] For example, the bitstream includes metadata, and the metadata includes the first information and a plurality of parameters (e.g., 3D scene parameters) related to the model, so that the decoding device can generate the model using the plurality of parameters.

[0405] For example, one or more images may include multiple neural network weights and second information (e.g., a feature plane) included in the model, the second information indicating multiple feature quantities having a time dimension. The acquisition step involves acquiring the multiple neural network weights and the second information from the one or more images. This allows the decoding device to decode the second information indicating the multiple feature quantities having a time dimension included in the model using a decoding process based on image encoding. This facilitates the decoding process of the information included in the model.

[0406] For example, as shown in Figures 57 and 58, the neural network weights and the second information are arranged in the same image included in one or more images, which allows the neural network weights and the second information to be arranged efficiently.

[0407] For example, as shown in Fig. 59, the plurality of neural network weights and the second information are arranged in different images included in one or more images, so that the decoding device can independently decode the plurality of neural network weights and the second information.

[0408] For example, as shown in Fig. 60, each of the one or more images includes a plurality of components (e.g., a Y component, a U component, and a V component), and the plurality of neural network weights and the second information are arranged on different components among the plurality of components, thereby enabling the plurality of neural network weights and the second information to be arranged efficiently.

[0409] For example, as shown in Fig. 61, each of the one or more images includes multiple layers (e.g., an enhancement layer and a base layer), and the multiple neural network weights and the second information are arranged in different layers among the multiple layers, thereby enabling the multiple neural network weights and the second information to be arranged efficiently.

[0410] The configuration of the decoding device is, for example, the same as that of the decoding device 200 shown in Fig. 11. The decoding device includes a processor (or circuit) and a memory, and the processor (or circuit) performs the above processing using the memory during operation.

[0411] One or more aspects disclosed herein may be implemented in combination with at least a part of other aspects of the present disclosure. Also, some processes shown in the flowcharts of one or more aspects disclosed herein, some configurations of devices, some syntax, etc. may be implemented in combination with other aspects.

[0412] [Implementation and Application] In each of the above embodiments, each of the functional or operational blocks can typically be realized by an MPU (micro processing unit), memory, etc. Furthermore, the processing by each of the functional blocks may be realized as a program execution unit such as a processor that reads and executes software (programs) recorded on a recording medium such as a ROM. The software may be distributed. The software may be recorded on various recording media such as semiconductor memory. It is also possible to realize each functional block by hardware (dedicated circuitry).

[0413] The processing described in each embodiment may be realized by centralized processing using a single device (system), or may be realized by distributed processing using multiple devices. Furthermore, the processor that executes the program may be a single processor or multiple processors. That is, centralized processing or distributed processing may be performed.

[0414] The aspects of the present disclosure are not limited to the above examples, and various modifications are possible, and these modifications are also included within the scope of the aspects of the present disclosure.

[0415] Furthermore, application examples of the video coding method (image coding method) or video decoding method (image decoding method) shown in each of the above embodiments and various systems for implementing the application examples will be described below. Such systems may be characterized by having an image coding device using the image coding method, an image decoding device using the image decoding method, or an image coding / decoding device that includes both. Other configurations of such systems can be appropriately changed depending on the situation.

[0416] [Example of Use] Fig. 68 shows the overall configuration of an appropriate content supply system ex100 that realizes a content distribution service. The area where communication services are provided is divided into cells of a desired size, and base stations ex106, ex107, ex108, ex109, and ex110, which are fixed wireless stations in the illustrated example, are installed in each cell.

[0417] In this content supply system ex100, devices such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, and a smartphone ex115 are connected to the Internet ex101 via an Internet service provider ex102 or a communication network ex104 and base stations ex106 to ex110. The content supply system ex100 may connect a combination of any of the above devices. In various implementations, the devices may be connected to each other directly or indirectly via a telephone network, short-range wireless communication, or the like, without going through the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be connected to devices such as the computer ex111, the game console ex112, the camera ex113, the home appliance ex114, and the smartphone ex115 via the Internet ex101, etc. The streaming server ex103 may also be connected to a terminal in a hotspot on an airplane ex117 via a satellite ex116.

[0418] Note that wireless access points, hot spots, etc. may be used instead of the base stations ex106 to ex110. Furthermore, the streaming server ex103 may be directly connected to the communication network ex104 without going through the Internet ex101 or the Internet service provider ex102, or may be directly connected to the airplane ex117 without going through the satellite ex116.

[0419] The camera ex113 is a device such as a digital camera that can take still images and videos. The smartphone ex115 is a smartphone, mobile phone, or PHS (Personal Handyphone System) that supports mobile communication systems such as 2G, 3G, 3.9G, 4G, and the upcoming 5G.

[0420] The home appliance ex114 is a refrigerator, or an appliance included in a home fuel cell cogeneration system, or the like.

[0421] In the content supply system ex100, a terminal having a photographing function is connected to a streaming server ex103 via a base station ex106 or the like, thereby enabling live streaming and the like. In live streaming, a terminal (such as a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, or a terminal on an airplane ex117) may perform the encoding process described in each of the above embodiments on still image or video content captured by a user using the terminal, may multiplex the video data obtained by encoding with audio data obtained by encoding audio corresponding to the video, and may transmit the obtained data to the streaming server ex103. In other words, each terminal functions as an image encoding device according to one aspect of the present disclosure.

[0422] Meanwhile, the streaming server ex103 streams the transmitted content data to the requesting client. The client is a computer ex111, a game console ex112, a camera ex113, a home appliance ex114, a smartphone ex115, a terminal on an airplane ex117, or the like, which is capable of decoding the encoded data. Each device that receives the distributed data decodes and plays back the received data. That is, each device may function as an image decoding device according to one aspect of the present disclosure.

[0423] [Distributed Processing] The streaming server ex103 may also be multiple servers or multiple computers that process, record, and distribute data in a distributed manner. For example, the streaming server ex103 may be implemented using a CDN (Content Delivery Network), where content distribution is achieved through a network connecting numerous edge servers distributed around the world. In a CDN, a physically nearby edge server is dynamically assigned depending on the client. Content is then cached and distributed to that edge server, thereby reducing delays. Furthermore, when certain types of errors occur or communication conditions change due to increased traffic, processing can be distributed among multiple edge servers, the distribution entity can be switched to another edge server, or distribution can be continued by bypassing the failed portion of the network, thereby achieving high-speed and stable distribution.

[0424] In addition to the distributed processing of the distribution itself, the encoding of captured data may be performed by each terminal, by the server, or by multiple terminals. For example, encoding generally involves two processing loops. The first loop detects the complexity of the image or the amount of code for each frame or scene. The second loop maintains image quality while improving encoding efficiency. For example, a terminal may perform the first encoding process, and the server that receives the content may perform the second encoding process, thereby improving content quality and efficiency while reducing the processing load on each terminal. In this case, if there is a request to receive and decode the data in near real time, the data encoded by a terminal can be received and played back by another terminal, enabling more flexible real-time distribution.

[0425] As another example, the camera ex113 or the like extracts features from an image, compresses the data related to the features as metadata, and transmits the compressed data to the server. The server performs compression according to the meaning (or importance of the content) of the image, for example, by determining the importance of an object from the features and switching the quantization precision accordingly. The feature data is particularly effective in improving the accuracy and efficiency of motion vector prediction when the server re-compresses the image. Alternatively, the terminal may perform simple encoding such as VLC (variable length coding), and the server may perform encoding with a high processing load such as CABAC (context-adaptive binary arithmetic coding).

[0426] As another example, in a stadium, shopping mall, factory, or the like, there may be multiple pieces of video data that have been shot by multiple terminals of almost the same scene. In this case, using the multiple terminals that shot the video and, as necessary, other terminals and servers that did not shoot the video, encoding processes are assigned to each of them, for example, in units of GOPs (Group of Pictures), pictures, or tiles obtained by dividing a picture, for distributed processing. This reduces delays and achieves better real-time performance.

[0427] Since multiple video data are of almost the same scene, the server may manage and / or instruct the video data shot by each terminal to be mutually referential. The server may also receive encoded data from each terminal and change the reference relationships between multiple data, or correct or replace the pictures themselves and re-encode them. This allows for the generation of streams with improved quality and efficiency for each piece of data.

[0428] Furthermore, the server may perform transcoding to change the encoding method of the video data before distributing it. For example, the server may convert an MPEG-based encoding method into a VP-based encoding method (e.g., VP9), or convert H.264 to H.265.

[0429] In this way, the encoding process can be performed by a terminal or one or more servers. Therefore, although the following uses terms such as "server" or "terminal" to refer to the entity performing the process, some or all of the processing performed by the server may be performed by the terminal, and some or all of the processing performed by the terminal may be performed by the server. The same applies to the decoding process.

[0430] [3D, Multi-Angle] Images or videos of different scenes or the same scene taken from different angles using multiple devices such as cameras ex113 and / or smartphones ex115 that are approximately synchronized with each other are increasingly being integrated and used. The videos taken by each device are integrated based on the relative positional relationship between the devices obtained separately, or on areas where feature points in the videos match.

[0431] The server may not only encode two-dimensional video, but also encode still images automatically or at a time specified by the user based on scene analysis of the video and transmit them to the receiving terminal. Furthermore, if the server can acquire the relative positional relationship between the capturing terminals, it can generate a three-dimensional shape of the scene based on not only two-dimensional video but also video of the same scene captured from different angles. The server may separately encode three-dimensional data generated by a point cloud or the like, or may select or reconstruct the video to be transmitted to the receiving terminal from video captured by multiple terminals based on the results of recognizing or tracking people or objects using the three-dimensional data.

[0432] In this way, the user can enjoy a scene by arbitrarily selecting each video corresponding to each shooting terminal, or can enjoy content in which a video from a selected viewpoint is cut out from 3D data reconstructed using multiple images or videos. Furthermore, together with the video, sound may also be collected from multiple different angles, and the server may multiplex the sound from a specific angle or space with the corresponding video and transmit the multiplexed video and sound.

[0433] In recent years, content that associates the real world with a virtual world, such as Virtual Reality (VR) and Augmented Reality (AR), has also become popular. In the case of VR images, the server creates viewpoint images for the right eye and left eye, respectively, and may perform encoding that allows reference between each viewpoint video using Multi-View Coding (MVC) or the like, or may encode them as separate streams without referencing each other. When decoding the separate streams, it is preferable to play them in synchronization with each other so that a virtual three-dimensional space is reproduced according to the user's viewpoint.

[0434] In the case of AR images, the server superimposes virtual object information in the virtual space onto camera information in the real space based on the three-dimensional position or the movement of the user's viewpoint. The decoding device may acquire or store virtual object information and three-dimensional data, generate a two-dimensional image according to the movement of the user's viewpoint, and smoothly connect the two-dimensional image to create superimposed data. Alternatively, the decoding device may send the movement of the user's viewpoint to the server in addition to a request for virtual object information. The server may create superimposed data according to the movement of the viewpoint received from the three-dimensional data stored on the server, encode the superimposed data, and distribute it to the decoding device. Note that the superimposed data may have an α value indicating transparency in addition to RGB, and the server may set the α value of parts other than the object created from the three-dimensional data to 0, etc., to encode the parts in a transparent state. Alternatively, the server may generate data by setting a predetermined RGB value as the background, like a chromakey, and using the background color for parts other than the object.

[0435] Similarly, the decoding of distributed data may be performed by each client terminal, by the server, or by multiple terminals. For example, one terminal may first send a reception request to the server, and then other terminals may receive and decode content according to the request, after which the decoded signal is transmitted to a device having a display. By distributing the processing and selecting appropriate content regardless of the capabilities of the communication terminals themselves, high-quality data can be reproduced. As another example, large-sized image data may be received on a TV or other device, and only a portion of the picture, such as a tile into which the picture is divided, may be decoded and displayed on the viewer's personal device. This allows the viewer to share the overall picture while checking their own area of ​​responsibility or an area of ​​interest in more detail.

[0436] In situations where multiple short-, medium-, or long-range wireless communications are available indoors and outdoors, seamless content reception may be possible using distribution system standards such as MPEG-DASH (Dynamic Adaptive Streaming over HTTP). Users may freely select and switch in real time between decoding devices or display devices, such as their own terminals and indoor / outdoor displays. Furthermore, decoding can be performed while switching between decoding and display devices using their own location information, etc. This allows information to be mapped and displayed on a part of the wall or ground of a neighboring building with an embedded display device while the user is traveling to their destination. It is also possible to switch the bit rate of received data based on the accessibility of the encoded data on the network, such as when the encoded data is cached on a server that can be quickly accessed from the receiving terminal or copied to an edge server in a content delivery service.

[0437] [Web Page Optimization] FIG. 69 is a diagram showing an example of a display screen of a web page on a computer ex111 or the like. FIG. 70 is a diagram showing an example of a display screen of a web page on a smartphone ex115 or the like. As shown in FIGS. 69 and 70 , a web page may include multiple link images that are links to image content, and the appearance of the link images may differ depending on the device used to view the page. When multiple link images are visible on the screen, the display device (decoding device) may display a still image or I-picture contained in each content as a link image until the user explicitly selects the link image, or until the link image approaches the center of the screen or the entire link image is within the screen. Alternatively, the display device (decoding device) may display a video such as a GIF animation using multiple still images or I-pictures, or may receive only the base layer and decode and display the video.

[0438] When a link image is selected by a user, the display device performs decoding while giving top priority to the base layer. Note that if the HTML (HyperText Markup Language) constituting the web page contains information indicating that the content is scalable, the display device may decode up to the enhancement layer. Furthermore, to ensure real-time performance, before selection or when the communication bandwidth is very limited, the display device decodes and displays only forward-reference pictures (I pictures, P pictures, and forward-reference-only B pictures), thereby reducing the delay between the decoding time of the first picture and the display time (the delay from the start of content decoding to the start of display). Furthermore, the display device may intentionally ignore the picture reference relationships and roughly decode all B pictures and P pictures using forward reference, and then perform normal decoding as the number of received pictures increases over time.

[0439] [Autonomous Driving] When transmitting and receiving still image or video data such as two-dimensional or three-dimensional map information for automatic driving or driving assistance of a vehicle, the receiving terminal may receive weather or construction information as meta information in addition to image data belonging to one or more layers, and may associate and decode these. Note that the meta information may belong to a layer, or may simply be multiplexed with the image data.

[0440] In this case, since a vehicle, drone, airplane, or the like including the receiving terminal is moving, the receiving terminal can transmit location information of the receiving terminal, thereby realizing seamless reception and decoding while switching between base stations ex106 to ex110. Furthermore, the receiving terminal can dynamically switch how much meta information to receive or how much to update map information depending on the user's selection, the user's situation, and / or the state of the communication bandwidth.

[0441] In the content supply system ex100, the client can receive, decode, and play back encoded information sent by a user in real time.

[0442] [Distribution of Personal Content] The content supply system ex100 also allows for unicast or multicast distribution of not only high-quality, long-duration content from video distribution companies, but also low-quality, short-duration content from individuals. It is expected that such personal content will continue to increase in the future. To improve the quality of personal content, the server may perform editing before encoding. This can be achieved, for example, using the following configuration.

[0443] During shooting, either in real time or after accumulating and shooting, the server performs recognition processing such as detecting shooting errors, scene search, semantic analysis, and object detection from the original image data or encoded data. Based on the recognition results, the server manually or automatically corrects out-of-focus or camera shake, deletes less important scenes such as scenes with lower brightness or out-of-focus compared to other pictures, emphasizes object edges, changes color, and performs other editing. The server then encodes the edited data based on the editing results. It is also known that viewing rates decrease if the shooting time is too long. Therefore, the server may automatically clip not only less important scenes as described above but also scenes with little movement, based on the image processing results, so that the content falls within a specific time range depending on the shooting time. Alternatively, the server may generate and encode a digest based on the results of the semantic analysis of the scenes.

[0444] Personal content may contain content that, if left as is, violates copyright, moral rights, or portrait rights, and may cause the scope of sharing to exceed the intended scope, resulting in inconvenience to individuals. Therefore, for example, the server may intentionally defocus images of people's faces on the periphery of the screen or the interior of a house before encoding. Furthermore, the server may recognize whether the image to be encoded contains the face of a person other than a pre-registered person, and if so, perform processing such as blurring the face. Alternatively, as pre- or post-processing before encoding, the user may specify a person or background area they wish to modify in the image for copyright or other reasons. The server may replace the specified area with another image or blur the focus. If the image contains a person, the server may track the person in the video and replace the image of the person's face.

[0445] Because viewing personal content with small data volumes requires high real-time performance, the decoding device first receives the base layer as a top priority, and then decodes and plays it back, depending on the bandwidth. The decoding device may also receive an enhancement layer during this time, and if the content is played back more than once, such as when playback is looped, it may play back high-quality video including the enhancement layer. A stream that has undergone scalable encoding in this way can provide an experience in which the video appears rough when not selected or when viewing begins, but gradually becomes smarter and the image quality improves. In addition to scalable encoding, a similar experience can also be provided when a rough stream played the first time and a second stream that is encoded with reference to the first video are configured as a single stream.

[0446] [Other Application Examples] Furthermore, these encoding or decoding processes are generally performed by the LSIex500 possessed by each terminal. The LSI (large scale integration circuitry) ex500 (see FIG. 68) may be a single-chip or multi-chip configuration. Furthermore, video encoding or decoding software may be embedded in some kind of recording medium (such as a CD-ROM, flexible disk, or hard disk) readable by the computer ex111, and the encoding or decoding process may be performed using that software. Furthermore, if the smartphone ex115 is equipped with a camera, video data captured by the camera may be transmitted. This video data is data encoded and processed by the LSIex500 possessed by the smartphone ex115.

[0447] The LSIex500 may be configured to download and activate application software. In this case, the terminal first determines whether it supports the content encoding method or has the capability to execute a specific service. If the terminal does not support the content encoding method or does not have the capability to execute a specific service, the terminal downloads a codec or application software, and then acquires and plays the content.

[0448] Furthermore, at least one of the moving image encoding device (image encoding device) or moving image decoding device (image decoding device) of each of the above embodiments can be incorporated into a digital broadcasting system, not limited to the content supply system ex100 via the Internet ex101. Since multiplexed data in which video and audio are multiplexed is transmitted and received over broadcast radio waves using a satellite or the like, there is a difference in that it is more suited to multicast than the content supply system ex100, which has a configuration that is easy to use for unicast, but similar applications are possible with regard to encoding and decoding processes.

[0449] [Hardware Configuration] Fig. 71 is a diagram showing further details of the smartphone ex115 shown in Fig. 68. Fig. 72 is a diagram showing an example configuration of the smartphone ex115. The smartphone ex115 includes an antenna ex450 for transmitting and receiving radio waves to and from the base station ex110, a camera unit ex465 capable of capturing video and still images, and a display unit ex458 for displaying video captured by the camera unit ex465 and decoded data of the video and the like received by the antenna ex450. The smartphone ex115 further includes an operation unit ex466 such as a touch panel, an audio output unit ex457 such as a speaker for outputting voice or sound, an audio input unit ex456 such as a microphone for inputting voice, a memory unit ex467 capable of storing captured video or still images, recorded voice, received video or still images, encoded data such as email, or decoded data, and a slot unit ex464 that is an interface with a SIM (Subscriber Identity Module) ex468 for identifying a user and authenticating access to various data including the network. Note that an external memory may be used instead of the memory unit ex467.

[0450] A main control unit ex460 that comprehensively controls the display unit ex458 and operation unit ex466, etc., is connected to a power supply circuit unit ex461, an operation input control unit ex462, a video signal processing unit ex455, a camera interface unit ex463, a display control unit ex459, a modulation / demodulation unit ex452, a multiplexing / separation unit ex453, an audio signal processing unit ex454, a slot unit ex464, and a memory unit ex467 via a synchronization bus ex470.

[0451] When the power key is turned on by a user's operation, the power supply circuit unit ex461 starts up the smartphone ex115 to an operable state and supplies power to each unit from the battery pack.

[0452] The smartphone ex115 performs processes such as telephone calls and data communications under the control of a main control unit ex460 having a CPU, ROM, RAM, etc. During a call, an audio signal collected by an audio input unit ex456 is converted into a digital audio signal by an audio signal processing unit ex454, subjected to spectrum spread processing by a modulation / demodulation unit ex452, subjected to digital-to-analog conversion and frequency conversion processing by a transmission / reception unit ex451, and the resulting signal is transmitted via an antenna ex450. The received data is also amplified and subjected to frequency conversion and analog-to-digital conversion processing, subjected to spectrum despreading processing by a modulation / demodulation unit ex452, and converted into an analog audio signal by an audio signal processing unit ex454, which is then output from an audio output unit ex457. During data communication mode, text, still images, or video data is sent to the main control unit ex460 via an operation input control unit ex462 based on operations on the main unit's operation unit ex466, etc. Similar transmission and reception processing is performed. When transmitting video, still images, or video and audio in the data communication mode, the video signal processing unit ex455 compresses and encodes the video signal stored in the memory unit ex467 or the video signal input from the camera unit ex465 using the moving image encoding method described in each of the above embodiments, and sends the encoded video data to the multiplexing / demultiplexing unit ex453. The audio signal processing unit ex454 encodes the audio signal collected by the audio input unit ex456 while the camera unit ex465 is capturing video or still images, and sends the encoded audio data to the multiplexing / demultiplexing unit ex453. The multiplexing / demultiplexing unit ex453 multiplexes the encoded video data and encoded audio data using a predetermined method, and modulates and converts the multiplexed video data and audio data in the modulation / demodulation unit (modulation / demodulation circuit unit) ex452 and the transmission / reception unit ex451, before transmitting the multiplexed video data and audio data via the antenna ex450.

[0453] In order to decode the multiplexed data received via the antenna ex450, such as when receiving video attached to an email or chat, or video linked to a web page, the multiplexing / separation unit ex453 separates the multiplexed data into a video data bit stream and an audio data bit stream, and supplies the encoded video data to the video signal processing unit ex455 and the encoded audio data to the audio signal processing unit ex454 via the synchronization bus ex470. The video signal processing unit ex455 decodes the video signal using a video decoding method corresponding to the video encoding method described in each of the above embodiments, and the video or still image contained in the linked video file is displayed on the display unit ex458 via the display control unit ex459. The audio signal processing unit ex454 decodes the audio signal, and audio is output from the audio output unit ex457. As real-time streaming becomes increasingly common, audio playback may be socially inappropriate depending on the user's situation. Therefore, it is preferable that the initial setting be a configuration in which only the video data is played without playing the audio signal, and audio may be played in sync only when the user performs an operation such as clicking on the video data.

[0454] Although the smartphone ex115 has been used as an example, three other implementation formats are possible: a transmitting / receiving terminal having both an encoder and a decoder, a transmitting terminal having only an encoder, and a receiving terminal having only a decoder. In the digital broadcasting system, multiplexed data in which audio data is multiplexed with video data is received or transmitted. However, in addition to audio data, text data related to the video may also be multiplexed into the multiplexed data. Furthermore, the video data itself may be received or transmitted instead of the multiplexed data.

[0455] Although the main control unit ex460 including a CPU has been described as controlling the encoding or decoding process, various terminals often include a GPU (Graphics Processing Unit). Therefore, a configuration may be adopted in which a memory shared by the CPU and GPU, or a memory whose addresses are managed for common use, is used to take advantage of the GPU's performance to process a large area in a batch. This shortens the encoding time, ensures real-time performance, and achieves low latency. It is particularly efficient to perform motion estimation, deblocking filtering, SAO (Sample Adaptive Offset), and transformation / quantization processes in a batch, such as by picture, by the GPU rather than by the CPU.

[0456] The present disclosure is applicable to, for example, encoding devices and decoding devices.

[0457] 100 Encoding device 102 Splitting unit 104 Subtraction unit 106 Transformation unit 108 Quantization unit 110 Entropy encoding unit 112, 204 Inverse quantization unit 114, 206 Inverse transformation unit 116, 208 Addition unit 118, 210 Block memory 120, 212 Loop filter unit 122, 214 Frame memory 124, 216 Intra prediction unit 126, 218 Inter prediction unit 128, 220 Prediction control unit 130, 222 Prediction parameter generation unit 200 Decoding device 202 Entropy decoding unit 224 Splitting decision unit 300, 310, 500 Encoding device 301, 311, 501 Transformation unit 302, 312 NNC encoding unit 303, 313, 502 Metadata conversion unit 304, 314, 503 Image encoding unit 400, 410, 600 Decoding device 401, 411, 601 Image decoding unit 402, 412, 602 Metadata inverse conversion unit 403, 413 NNC decoding unit 404, 414, 603 Inverse conversion unit 405, 415, 604 Rendering unit a1, b1 Processor a2, b2 Memory

Claims

1. An encoding device comprising: a circuit; and a memory connected to the circuit, wherein the circuit, in operation, acquires first information included in a model, the first information indicating a plurality of features having a time dimension; places the first information on one or more images; and generates a bitstream by encoding the one or more images.

2. The encoding device according to claim 1, wherein the first information includes a plurality of data, and the bit stream includes second information indicating the arrangement order of the plurality of data.

3. The encoding device according to claim 1, wherein the first information includes a plurality of data, and the bitstream includes third information indicating a starting position of at least one of the plurality of data in the one or more images.

4. The encoding device according to claim 1, wherein the first information includes a plurality of feature planes each representing a plurality of feature quantities arranged two-dimensionally.

5. The encoding device of claim 1, wherein the first information includes first data including a plurality of features not having the time dimension and a plurality of second data including a plurality of features having the time dimension, each of which corresponds to a different time interval; the one or more images include a plurality of first images and a plurality of second images; the first data is divided and arranged in the plurality of first images; and the plurality of second data is respectively arranged in the plurality of second images.

6. The encoding device of claim 1, wherein the first information includes first data including a plurality of features not having the time dimension and a plurality of second data including a plurality of features having the time dimension, each of which corresponds to a different time interval; the one or more images include a first image and a second image; the first data is arranged in the first image; and the plurality of second data are arranged together in the second image.

7. The encoding device according to claim 1, wherein the first information includes first data and second data having a higher resolution than the first data, the one or more images include a first image and a second image that is encoded after the first image, the first data is arranged in the first image, and the second data is arranged in the second image.

8. The encoding device according to claim 1, wherein the first information includes first data and second data having a higher resolution than the first data, each of the one or more images includes a first component and a second and third components each smaller than the first component, the first data being arranged in at least one of the second and third components, and the second data being arranged in the first component.

9. A decoding device comprising: a circuit; and a memory connected to the circuit, wherein the circuit, in operation, decodes one or more images from a bit stream generated by encoding one or more images in which first information included in a model, the first information indicating a plurality of features having a time dimension, is arranged, and obtains the first information from the one or more images.

10. The decoding device according to claim 9, wherein the first information includes a plurality of data, and the bit stream includes second information indicating the arrangement order of the plurality of data.

11. The decoding device according to claim 9, wherein the first information includes a plurality of data, and the bitstream includes third information indicating a starting position of at least one of the plurality of data in the one or more images.

12. The decoding device according to claim 9, wherein the first information includes a plurality of feature planes each representing a plurality of feature quantities arranged two-dimensionally.

13. The decoding device of claim 9, wherein the first information includes first data including a plurality of features not having the time dimension and a plurality of second data including a plurality of features having the time dimension, each of which corresponds to a different time interval; the one or more images include a plurality of first images and a plurality of second images; the first data is divided and arranged in the plurality of first images; and the plurality of second data is respectively arranged in the plurality of second images.

14. The decoding device of claim 9, wherein the first information includes first data including a plurality of features not having the time dimension and a plurality of second data including a plurality of features having the time dimension, each of which corresponds to a different time interval; the one or more images include a first image and a second image; the first data is arranged in the first image; and the plurality of second data are arranged together in the second image.

15. A decoding device as described in claim 9, wherein the first information includes first data and second data having a higher resolution than the first data, the one or more images include a first image and a second image that is encoded after the first image, the first data is arranged in the first image, and the second data is arranged in the second image.

16. A decoding device as described in claim 9, wherein the first information includes first data and second data having a higher resolution than the first data, each of the one or more images includes a first component and second and third components each smaller than the first component, the first data is arranged in at least one of the second and third components, and the second data is arranged in the first component.

17. An encoding method comprising: acquiring first information included in a model, the first information indicating a plurality of feature quantities having a time dimension; arranging the first information on one or more images; and encoding the one or more images to generate a bitstream.

18. A decoding method comprising: decoding one or more images from a bitstream generated by encoding one or more images in which first information included in a model, the first information indicating a plurality of features having a time dimension, is arranged; and obtaining the first information from the one or more images.

Citation Information

Patent Citations

  • Image decoding method, image coding method, image decoder, and image encoder

    WO2022225025A1

  • Video encoding device, video decoding device, video encoding method and video decoding method

    WO2023112879A1

  • Encoding device, decoding device, encoding method, and decoding method

    WO2024195509A1