Image processing methods and equipment

VN126388APending Publication Date: 2026-06-15SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
VN · VN
Patent Type
Applications
Current Assignee / Owner
SONY GROUP CORP
Filing Date
2024-07-22
Publication Date
2026-06-15

AI Technical Summary

Technical Problem

Existing methods for rendering 3D data composed of multiple layer images face challenges in maintaining image quality due to constraints on the characteristics of the MPI layer, such as fixed layer number ratios and resolutions, which can lead to reduced image quality and resolution.

Method used

The proposed solution involves generating and encoding interlayer relative information regarding the relative relationship between layers of 3D data, which includes information about the number of layers and resolution ratios between color and alpha layers. This information is added to the bitstream to optimize the layer relationships based on content.

Benefits of technology

By transmitting interlayer relative information, the decoder can optimize the layer relationships, thereby suppressing the reduction in image quality and allowing for more flexible layer arrangements without increasing processing time or data amounts unnecessarily.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure VN1202603823_0
    Figure VN1202603823_0
Patent Text Reader

Abstract

The invention relates to an image processing device and method, and more specifically, to an image processing device and method capable of preventing degradation of acquired images by rendering 3D data comprising multiple image layers. The creation of interlayer relative information relating to the relative relationship between layers of 3D data comprising color component image layers and alpha component image layers, and the creation of a bitstream by encoding 3D data and adding interlayer relative information to the bitstream is performed. Furthermore, the creation of 3D data by decoding the bitstream, the analysis of interlayer relative information extracted from the bitstream, and the rendering of 3D data based on interlayer relative information and the creation of a viewport image observable from any viewport are performed. The inventive technique may be applied to, for example, an image processing device, an electronic device, an image processing method, or a program.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device and method

[0001] The present disclosure relates to an image processing device and method, and more particularly to an image processing device and method that can suppress degradation in quality of an image rendered from 3D data composed of multiple layer images.

[0002] Conventionally, one method for representing the three-dimensional shape of an object is Representing Scenes as Neural Radiance Fields for View Synthesis (NeRF), which generates a radiance field corresponding to the space containing the object and approximates the radiance field using a neural network (see, for example, Non-Patent Documents 1 to 3). NeRF represents the spatial information of a 3D object using a Multi-Layer Perceptron (MLP).

[0003] There is also a method called MPI (Multi Plane Image) that divides a three-dimensional space into multiple layers (2D images) to represent it. To speed up the rendering process of NeRF, a method called NeX was devised that converts MLP to this MPI for handling (see, for example, Non-Patent Document 4). MPI is a group of images in which multiple 2D images are arranged in layers, and 3D data can be handled as a 2D image.

[0004] Also, a method has been proposed for extending the MPI specification by using SEI (Supplemental Enhancement Information) for a bitstream in which MPI is coded (see, for example, Non-Patent Document 5).

[0005] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, Ren Ng, "NeRF: Representing scenes as neural radiance fields for view synthesis", ECCV 2020, 2020 / 3 / 19Thomas Muller, Alex Evans, Christoph Schied, Alexander Keller, "Instant neural graphics primitives with a multiresolution hash encoding", arXiv preprint arXiv:2201.05989, 2022 / 1 / 16Ricardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron, Alexey Dosovitskiy, Daniel Duckworth, "NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections", https: / / arxiv.org / abs / 2008.02268, 2020 / 8 / 5Suttisak Wizadwongsa, Pakkapon Phongthawee, Jiraphon Yenphraphai, "NeX: Real-time View Synthesis with Neural Basis Expansion", arXiv:2103.05606v2 [cs.CV] 12 Apr 2021Taoran Lu, Peng Yin, Guan-Ming Su, Dae Yeol Lee, Tsung-wei Huang, Sean McCarthy, Walt Husak, Gary J. Sullivan: AHG9: Multiplane Image Information SEI, JVET-AE0066-v5, version 5, 2023-07-15

[0006] However, even if the extension method described in Non-Patent Document 5 is applied, there is a risk that the quality of the image rendered from 3D data may be reduced due to restrictions on the characteristics of MPI layers.

[0007] The present disclosure has been made in view of such circumstances, and makes it possible to suppress a decrease in the quality of an image obtained by rendering 3D data configured from a plurality of layer images.

[0008] An image processing device according to one aspect of the present technology is an image processing device including: an inter-layer relative information generation unit that generates inter-layer relative information regarding a relative relationship between layers of 3D data configured from layer images of color components and layer images of alpha components; and an encoding unit that encodes the 3D data to generate a bitstream and adds the inter-layer relative information to the bitstream.

[0009] An image processing method according to one aspect of the present technology is an image processing method that includes generating inter-layer relative information regarding a relative relationship between layers of 3D data configured from a layer image of a color component and a layer image of an alpha component, encoding the 3D data to generate a bitstream, and adding the inter-layer relative information to the bitstream.

[0010] An image processing device according to another aspect of the present technology is an image processing device including: a decoding unit that decodes a bitstream and generates 3D data composed of a layer image of a color component and a layer image of an alpha component; an inter-layer relative information analysis unit that analyzes inter-layer relative information regarding a relative relationship between layers of the 3D data extracted from the bitstream; and a viewpoint image generation unit that renders the 3D data based on the inter-layer relative information and generates a viewpoint image viewed from an arbitrary viewpoint.

[0011] An image processing method according to another aspect of the present technology is an image processing method that includes: decoding a bitstream to generate 3D data composed of a layer image of a color component and a layer image of an alpha component; analyzing inter-layer relative information regarding a relative relationship between layers of the 3D data extracted from the bitstream; and rendering the 3D data based on the inter-layer relative information to generate a viewpoint image viewed from an arbitrary viewpoint.

[0012] In an image processing device and method according to one aspect of the present technology, inter-layer relative information regarding a relative relationship between layers of 3D data composed of a layer image of a color component and a layer image of an alpha component is generated, the 3D data is encoded to generate a bitstream, and the inter-layer relative information is added to the bitstream.

[0013] In another aspect of the image processing device and method of the present technology, a bitstream is decoded to generate 3D data composed of a layer image of a color component and a layer image of an alpha component, inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream is analyzed, and the 3D data is rendered based on the inter-layer relative information to generate a viewpoint image viewed from any viewpoint.

[0014] 1 is a diagram illustrating 3D data using MPI. FIG. 1 is a diagram illustrating 3D data using MPI. FIG. 1 is a diagram illustrating NeX. FIG. 1 is a diagram illustrating NeX. FIG. 1 is a diagram illustrating NeX. FIG. 1 is a diagram illustrating an example of a method for extending a relative relationship between layers. FIG. 1 is a diagram illustrating an example of syntax. FIG. 1 is a diagram illustrating a method for packing a layer image. FIG. 1 is a diagram illustrating a method for packing a layer image. FIG. 1 is a diagram illustrating an example of a method for extending a ratio of the number of layers between color component layers and alpha component layers. FIG. 1 is a diagram illustrating an extension of a resolution ratio between color component layers and alpha component layers. FIG. 1 is a diagram illustrating an extension of a ratio of the number of layers and a resolution ratio between color component layers and alpha component layers. FIG. 1 is a diagram illustrating an example of a method for extending layer arrangement. FIG. 1 is a diagram illustrating an example of an additional layer. FIG. 1 is a diagram illustrating an example of a non-uniform arrangement function. FIG. 1 is a diagram illustrating an example of control information related to layer arrangement. FIG. 1 is a diagram illustrating an example of a method for extending a layer shape. FIG. 1 is a diagram illustrating an example of a layer shape. FIG. 1 is a diagram illustrating an example of control information related to a layer shape. FIG. 1 is a block diagram illustrating an example of a main configuration of an encoding device. It is a flowchart explaining an example of the flow of encoding processing. It is a block diagram showing an example of the main configuration of a decoding device. It is a flowchart explaining an example of the flow of decoding processing. It is a block diagram showing an example of the main configuration of a computer.

[0015] Below, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 1. Literature etc. supporting technical content and technical terminology 2. 3D data using MPI 3. Layer-related extensions 4. First embodiment (encoding device) 5. Second embodiment (decoding device) 6. Supplementary notes

[0016] <1. Literature, etc. supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the content described in the embodiments, but also the content described in the following non-patent documents, etc. that were publicly known at the time of filing, and the content of other documents referenced in the following non-patent documents.

[0017] Non-patent document 1: (described above) Non-patent document 2: (described above) Non-patent document 3: (described above) Non-patent document 4: (described above) Non-patent document 5: (described above)

[0018] In other words, the contents of the above-mentioned non-patent documents and the contents of other documents referenced in the above-mentioned non-patent documents are also used as the basis for determining the support requirements.

[0019] <2. 3D Data Using MPI> <NeRF> Conventionally, as a method for representing the three-dimensional shape of an object, there has been NeRF (Representing Scenes as Neural Radiance Fields for View Synthesis), which generates a radiance field corresponding to a space including the object and approximates the radiance field using a neural network, as described in Non-Patent Documents 1 to 3.

[0020] A radiation field represents a scene in a space containing an object by using the color of each direction at each position in that space and the opacity of each position in that space. Here, opacity is an indicator of the presence of some kind of object, and can also be thought of as density. In other words, the shape of a three-dimensional space can be represented by obtaining a radiation field in which the density of coordinates where an object exists is high. NeRF approximates such a radiation field using a neural network. In other words, NeRF is a technology that generates a radiation field corresponding to a 3D object, approximates that radiation field using a neural network, and uses that neural network to perform rendering.

[0021] In NeRF, an object in three-dimensional space is imaged from different positions, the camera pose of the obtained images is determined, and the image and camera pose are used to train a neural network, generating a neural network (MLP (Multi Layer Perceptron)) that approximates a radiation field representing the three-dimensional shape of the object. During inference, pose information for a desired new viewpoint is input to this neural network, and a rendered image for that new viewpoint is obtained. In this specification, such pose information for a certain viewpoint (arbitrary viewpoint) is also referred to as viewpoint information. A rendered image for a certain viewpoint (arbitrary viewpoint) is also referred to as viewpoint image.

[0022] <NeX> There is also a method called MPI (Multi Plane Image) that divides a three-dimensional space into multiple layer images (2D images) and expresses it. In Non-Patent Document 4, a method called NeX was proposed that converts MLP into this MPI and handles it in order to speed up the rendering process of NeRF. MPI is a group of images in which multiple 2D images are arranged in layers, and 3D data can be handled as a 2D image.

[0023] In MPI, multiple layers (2D planes) are formed in a three-dimensional space, such as layers 11 to 14 in Figure 1, and objects are represented on each layer. Each layer is 2D data representing the distribution of specific information on a 2D plane, and can therefore also be referred to as an image. In this specification, these layers are also referred to as layer images. MPI is composed of multiple RGBA images. An RGBA image is a combination of a color component (RGB) image and an alpha component (Alfa) image. The alpha component indicates the presence of an object. The alpha component may represent the presence or absence of an object using the values ​​"0" and "1," or it may represent the probability that an object exists (or does not exist). In other words, the alpha component can be considered transparency. This RGBA image is generated by learning multi-viewpoint images using, for example, a deep neural network (DNN). In other words, MPI is 3D data composed of color component layer images and alpha component layer images.

[0024] A viewpoint image (a rendering image from a specified viewpoint) shows the composite result (superimposition result) of objects on each layer as seen from that viewpoint. For example, as shown in FIG. 1 , assume that object 21 exists on layer 11, object 22 exists on layer 12, object 23 exists on layer 13, and object 24 exists on layer 14. In the rendering image, these objects are superimposed in the positional relationship when viewed from the viewpoint. Rendering image 41 in FIG. 2 shows an example of a rendering image from viewpoint 31 in FIG. 1 , and rendering image 42 shows a rendering image from viewpoint 32 in FIG. 1 . As such, the way in which objects 21 to 24 overlap changes depending on the viewpoint, thereby expressing the three-dimensional effect of three-dimensional space (the three-dimensional shape of the objects).

[0025] In NeX, NeRF's MLP is converted into this type of MPI (also called NeX MPI) for handling. In other words, three-dimensional space is represented by multiple layers. An example of the NeX layer structure is shown in Figure 3. In the case of NeX MPI, multiple main layers 51 are formed, such as main layer 51-1, main layer 51-2, ..., and multiple sublayers 52 are formed between the main layers 51, such as sublayer 52-1 to sublayer 52-11, sublayer 52-21, .... In other words, in the group of layers representing three-dimensional space, some layers at a predetermined interval are formed as main layers 51, and the remaining layers in between are formed as sublayers 52.

[0026] The number of layers in this layer group, the number of main layers, and the number of sub-layers between the main layers are all arbitrary. For example, a three-dimensional space (one scene) at a certain time may be represented by 192 layers, with one in 12 of those layers (i.e., a total of 16) being the main layer and the remaining 176 being sub-layers. In this case, one of the 12 consecutive layers in the layer group is the main layer and the remaining 11 are sub-layers. Also, for example, there may be 32 main layers. In this case, the number of layers in the layer group representing one scene is 384. Of course, this is not limited to these examples.

[0027] In the case of NeX, as shown in FIG. 4 , the data consists of an alpha image 71, an RGB_base image 72, an RGB_view image 73, and a base image 74. The alpha image 71 is distribution information of the alpha component indicating the presence of an object in a single layer among the layers representing a three-dimensional space. In other words, the alpha image 71 indicates where an object is present in the layer. The alpha image 71 may also indicate the probability that the object is present. The alpha image 71 may also indicate where an object is not present in the layer, or the probability that the object is not present. In other words, the alpha image 71 may indicate the transparency of each position in the layer. This alpha image 71 is generated for each sublayer. That is, this alpha image 71 is generated for each of all layers (main layer 51 and sublayer 52 in FIG. 3 ) in the layer group representing a three-dimensional space.

[0028] The RGB_base image 72 and the RGB_view image 73 are information indicating the distribution of RGB components that indicate color information of an object in multiple consecutive layers in a layer group. The RGB_base image and the RGB_view image are also collectively referred to as RGB images.

[0029] The RGB_base image 72 is color distribution information that does not depend on the angle of the viewpoint. The RGB_base image 72 is composed of a single image. In other words, the RGB_base image 72 is color distribution information with a strong color component when viewed from the front of the image.

[0030] The RGB_view image 73 is color distribution information that depends on the angle of the viewpoint. The RGB_view image 73 is composed of multiple images. The multiple images are color distribution information with strong color components when viewed from different viewpoints. For example, the RGB_view image 73 is composed of four images: an image k0 with strong color components when viewed from above, an image with strong color components when viewed from below, an image with strong color components when viewed from the left, and an image with strong color components when viewed from the right.

[0031] The RGB_base image 72 and the RGB_view image 73 are generated for each main layer, i.e., for each of a plurality of consecutive layers in the layer group. For example, in the case of Fig. 3, one RGB_base image 72 and one RGB_view image 73 are generated for each of the main layer 51-1 and the sublayers 52-1 to 52-11. In other words, the RGB_base image 72 represents color distribution information for 12 consecutive layers, and one image is generated for each of the 12 layers. Similarly, the RGB_view images 73 represent color distribution information for 12 consecutive layers, and four images are generated for each of the 12 layers.

[0032] The base images 74 are composed of the same number of images as the RGB_view images 73. For example, if four RGB_view images 73 are formed as described above, four base images 74 are also formed. Each base image 74 corresponds to a different RGB_view image 73, and indicates the blend ratio of the corresponding RGB_view image 73 for each viewpoint direction. In other words, the pixel position of the base image 74 is determined according to the viewpoint direction, and the RGB_view image 73 corresponding to that base image 74 is blended at the blend ratio indicated by the pixel value. For example, when the viewpoint is frontal, as shown in FIG. 5 , each RGB_view image 73 (k1 to k4) is multiplied by the blend ratio indicated by the central pixel value of the corresponding base image 74 (H1 to H4), and the multiplication results are added to the RGB_base image 72 (k0). Furthermore, when the viewpoint is at the top left, as shown in FIG. 6, each RGB_view image 73 (k1 to k4) is multiplied by the blend ratio indicated as the pixel value at the top left of the corresponding base image 74 (H1 to H4), and the multiplication results are added to the RGB_base image 72 (k0).

[0033] NeX MPI can improve expressiveness by changing the RGB parts depending on the viewpoint in this way. By treating the alpha image as a sublayer and the RGB image as a main layer, it is possible to suppress the increase in data volume while improving expressiveness and achieve more efficient learning.

[0034] <MPI Extensions and Constraints on Layer Characteristics> Non-Patent Document 5 also discloses a method for extending the MPI specifications by using SEI (Supplemental Enhancement Information) for a bitstream in which MPI is coded.

[0035] However, even when the extension method described in Non-Patent Document 5 was applied, there were restrictions on the characteristics of MPI layers. For example, there were restrictions on the arrangement, number, resolution, shape, etc. of layers. These restrictions could degrade the quality of images rendered from 3D data. For example, there was a risk of degradation in the image quality and resolution of the rendered image.

[0036] <3. Layer-related Extensions> <Method 1> For example, for 3D data configured with a color component layer image and an alpha component layer image as described above, the relative relationship between layers is fixed. In other words, the decoder needs to understand the relative relationship between the layers, which makes it difficult to optimize the relative relationship depending on, for example, the content. This can result in a reduction in the quality of the rendered image.

[0037] Therefore, in order to prevent a reduction in the quality of the rendered image due to such constraints, relative information regarding the color layer and the alpha layer is transmitted (Method 1), as shown in the top row of the table in Fig. 7. In this specification, a color component layer is also referred to as a color layer, and an alpha component layer is also referred to as an alpha layer.

[0038] For example, a first image processing device may include an inter-layer relative information generation unit that generates inter-layer relative information regarding the relative relationship between layers of 3D data configured with color component layer images and alpha component layer images, and an encoding unit that encodes the 3D data to generate a bitstream and adds the inter-layer relative information to the bitstream. For example, a first image processing method may include the first image processing device generating inter-layer relative information regarding the relative relationship between layers of 3D data configured with color component layer images and alpha component layer images, and the first image processing device encoding the 3D data to generate a bitstream and adding the inter-layer relative information to the bitstream. For example, a first program may cause the first image processing device to execute the following steps: generating inter-layer relative information regarding the relative relationship between layers of 3D data configured with color component layer images and alpha component layer images, encoding the 3D data to generate a bitstream, and adding the inter-layer relative information to the bitstream.

[0039] For example, the second image processing device may include a decoding unit that decodes a bitstream and generates 3D data composed of color component layer images and alpha component layer images, an inter-layer relative information analysis unit that analyzes inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream, and a viewpoint image generation unit that renders the 3D data based on the inter-layer relative information and generates a viewpoint image seen from an arbitrary viewpoint.For example, the second image processing device may include: the second image processing device decoding the bitstream and generating 3D data composed of color component layer images and alpha component layer images; the second image processing device analyzing the inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream; and the second image processing device rendering the 3D data based on the inter-layer relative information and generating a viewpoint image seen from an arbitrary viewpoint. For example, the second program causes the second image processing device to decode the bitstream and generate 3D data consisting of a layer image of a color component and a layer image of an alpha component, analyze inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream, render the 3D data based on the inter-layer relative information, and generate a viewpoint image viewed from any viewpoint.

[0040] This inter-layer relative information may be added to the bitstream in any manner. For example, this inter-layer relative information may be stored in an SEI added to the bitstream of 3D data. Of course, other methods may also be applied.

[0041] By transmitting the inter-layer relative information from the encoding side to the decoding side in this way, even if the relative relationship between layers is optimized according to the content, etc., the decoder can easily grasp the relative relationship based on the inter-layer relative information. Therefore, it is possible to suppress a decrease in the quality of the rendered image.

[0042] <Method 1-1> For example, in the method described in Non-Patent Document 5, the MPI specification is extended by adding an SEI with syntax as shown in Figure 8 to the bitstream. As shown in this syntax, this SEI does not transmit information that sets the relationship between the number of color layers and the number of alpha layers. In other words, the relationship between the number of color layers and the number of alpha layers is fixed (for example, one to one).

[0043] In the method described in Non-Patent Document 5, the layers constituting the MPI are packed and encoded using a predetermined method. At this time, as shown in A of Figure 9, a color layer group 111 in which the layers of the color components constituting the MPI are arranged, and an alpha layer group 112 in which the layers of the alpha component are arranged are formed. In the method described in Non-Patent Document 5, the ratio of the number of color layers to the number of alpha layers is 1:1, so the size and shape of the color layer group 111 and the alpha layer group 112 are the same.

[0044] Then, such color layer group 111 and alpha layer group 112 are arranged and packed. For example, as shown in B of FIG. 9 , the color layer group 111 and alpha layer group 112 may be arranged (stacked) in the vertical direction and packed. As shown in A of FIG. 10 , the color layer group 111 and alpha layer group 112 may be arranged (stacked) in the horizontal direction and packed. As shown in B of FIG. 10 , the color layer group 111 and alpha layer group 112 may be arranged (stacked) in the time (T) direction and packed. In other words, the color layer group 111 and alpha layer group 112 may be stacked as consecutive frames. The color layer group 111 and alpha layer group 112 are stacked according to a method selected from these predetermined options prepared in advance.

[0045] However, because the ratio of the number of color layers to the number of alpha layers is fixed in this way, there is a restriction that it is difficult to optimize the ratio depending on, for example, the content, etc. This may result in a reduction in the quality of the rendered image.

[0046] For example, an insufficient number of alpha layers can result in insufficient sampling density of the alpha component (shape information) depending on the viewpoint position, potentially resulting in a subjective quality degradation known as the stacked card effect, in which the outlines of objects become coarse in the rendered image. Increasing the number of layers can mitigate this stacked card effect, but it also increases the processing time and data volume for encoding, decoding, rendering, etc. Because the ratio of the number of color layers to the number of alpha layers is fixed, increasing the number of alpha layers also increases the number of color layers. Therefore, even if the number of color layers is sufficient, there is a risk of unnecessarily increasing the processing time and data volume. Furthermore, systems typically have limitations on the processing time and data volume for encoding, decoding, rendering, etc., making it difficult to increase the number of layers without limit. Therefore, when the processing time and data volume reach the system's upper limits, it is difficult to mitigate the degradation of the subjective quality of the rendered image due to an insufficient number of alpha layers.

[0047] Therefore, when Method 1 is applied, information regarding the relative number of alpha layers to color layers may be transmitted (Method 1-1), as shown in the second row from the top of the table in Fig. 7 . For example, in the first image processing device, the first image processing method, and the first program, the inter-layer relative information added to the bitstream may include information regarding the relative number of alpha component layers to color component layers. Similarly, in the second image processing device, the second image processing method, and the second program, the inter-layer relative information added to the bitstream may include information regarding the relative number of alpha component layers to color component layers.

[0048] The information about the relative number of alpha layers to color layers may be added to the bitstream in any manner, for example, it may be stored in an SEI that is added to the bitstream of the 3D data, although other methods may also be applied.

[0049] For example, a parameter "mpii_opacity_layer_num_minus1" as shown in box 113 in A of Fig. 11 may be stored in an SEI added to a bitstream of 3D data. The value of the parameter "mpii_opacity_layer_num_minus1" indicates the relative number of alpha layers to the number of color layers. In other words, the value of the parameter "mpii_opacity_layer_num_minus1" indicates the ratio between the number of color layers and the number of alpha layers. In other words, in 3D data, the color layers and the alpha layers are configured in the ratio indicated by this parameter "mpii_opacity_layer_num_minus1".

[0050] For example, if the value of this parameter "mpii_opacity_layer_num_minus1" is "5," the ratio of the number of color layers to the number of alpha layers in the 3D data is 1:4. In other words, during packing, four times as many alpha layers are stacked as color layers. For example, as shown in B to D of FIG. 11, four alpha layers 112 (A of FIG. 9) that are the same size and shape as the color layer group 111 are stacked on top of the color layer group 111 (A of FIG. 9). In FIG. 11, these four alpha layers 112 are also referred to as alpha layer group 112-1, alpha layer group 112-2, alpha layer group 112-3, and alpha layer group 112-4.

[0051] In this case, for example, as shown in B of Fig. 11 , the color layer group 111, the alpha layer group 112-1, the alpha layer group 112-2, the alpha layer group 112-3, and the alpha layer group 112-4 may be packed by being arranged vertically (stacked). Alternatively, as shown in C of Fig. 11 , the color layer group 111, the alpha layer group 112-1, the alpha layer group 112-2, the alpha layer group 112-3, and the alpha layer group 112-4 may be packed by being arranged horizontally (stacked). Alternatively, as shown in D of Fig. 11 , the color layer group 111, the alpha layer group 112-1, the alpha layer group 112-2, the alpha layer group 112-3, and the alpha layer group 112-4 may be packed by being arranged in the time (T) direction (stacked). That is, the color layer group 111, the alpha layer group 112-1, the alpha layer group 112-2, the alpha layer group 112-3, and the alpha layer group 112-4 may be stacked as consecutive frames. The color layer group 111, the alpha layer group 112-1, the alpha layer group 112-2, the alpha layer group 112-3, and the alpha layer group 112-4 may be stacked according to a method designated from these predetermined options prepared in advance.

[0052] By transmitting information about the relative number of alpha layers to color layers from the encoding side to the decoding side in this way, the decoder can easily grasp the ratio of the number of color layers to the number of alpha layers based on that information. In other words, the ratio of the number of color layers to the number of alpha layers can be made variable. For example, the ratio of the number of color layers to the number of alpha layers can be optimized depending on the content, etc.

[0053] In this way, by varying the ratio of the number of color layers to the number of alpha layers, it is possible to increase the number of alpha layers without increasing the number of color layers. For example, when the number of color layers is sufficient, it is possible to suppress the occurrence of a stacked card effect by suppressing a shortage of alpha layers without unnecessarily increasing processing time or data volume. Furthermore, even when the processing time or data volume has reached the system upper limit, it is possible to suppress the occurrence of a stacked card effect by suppressing a shortage of alpha layers.

[0054] That is, by transmitting information about the relative number of alpha layers to color layers from the encoding side to the decoding side, it is possible to suppress a reduction in the subjective quality of a rendered image due to an insufficient number of alpha layers.

[0055] <Method 1-2> For example, in the method described in Non-Patent Document 5, information that sets the relationship between the resolutions of the color layer and the alpha layer is not transmitted in the SEI, as shown in the syntax of Fig. 8. In other words, the relationship between the resolutions of the color layer and the alpha layer (which can also be said to be the relationship between the number of pixels or layer size) is fixed.

[0056] In the method described in Non-Patent Document 5, the resolutions (number of pixels, thickness, and size) of the color layer and the alpha layer are the same as each other, as shown in Figures 9 and 10. In other words, the ratio of the resolutions of the color layer and the alpha layer is fixed (1:1).

[0057] However, because the resolution ratio between the color layer and the alpha layer is fixed in this way, it is difficult to optimize the resolution ratio depending on the content, etc. This can result in a reduction in the quality of the rendered image.

[0058] Generally, a high resolution is more important for color layers than for alpha layers. Insufficient resolution of the color layer can result in blurred textures, resulting in poor reproducibility during rendering (i.e., reduced subjective quality of the rendered image). Increasing the resolution can reduce this texture blurring, but it also increases the processing time and data volume required for encoding, decoding, rendering, and other processes. Because the resolution ratio between the color layer and the alpha layer is fixed, increasing the resolution of the color layer also increases the resolution of the alpha layer. Therefore, even if the resolution of the alpha layer is sufficient, there is a risk of unnecessarily increasing the processing time and data volume. Furthermore, systems generally have limitations on the processing time and data volume required for encoding, decoding, rendering, and other processes, making it difficult to increase the resolution without limit. Therefore, when the processing time and data volume reach the system's upper limit, it is difficult to reduce the reduction in the subjective quality of the rendered image due to insufficient resolution of the color layer.

[0059] Therefore, when Method 1 is applied, information regarding the relative resolution of the alpha layer with respect to the color layer may be transmitted (Method 1-2), as shown in the third row from the top of the table in Fig. 7 . For example, in the first image processing device, the first image processing method, and the first program, the inter-layer relative information added to the bitstream may include information regarding the relative resolution of the alpha component layer with respect to the color component layer. Similarly, in the second image processing device, the second image processing method, and the second program, the inter-layer relative information added to the bitstream may include information regarding the relative resolution of the alpha component layer with respect to the color component layer.

[0060] The information about the relative resolution of the alpha layer with respect to the color layer may be added to the bitstream in any manner. For example, the information about the relative resolution of the alpha layer with respect to the color layer may be stored in an SEI added to the bitstream of the 3D data. Of course, other methods may also be applied.

[0061] For example, the parameter "mpii_opacity_layer_div_scale_log2" as shown in box 114 in A of FIG. 12 may be stored in the SEI added to the bitstream of 3D data. The value of the parameter "mpii_opacity_layer_div_scale_log2" indicates the relative resolution of the alpha layer with respect to the color layer. In other words, the value of the parameter "mpii_opacity_layer_div_scale_log2" indicates the ratio between the resolutions of the color layer and the alpha layer. In other words, the size ratio between the color layer and the alpha layer is the ratio indicated by this parameter "mpii_opacity_layer_div_scale_log2".

[0062] For example, if the resolution ratio between the color layers and the alpha layers is 4:1, the size of the alpha layer group will also be one-fourth of the color layer group, as shown in FIG. 11B or FIG. 11C. In this case, the layer groups may be formed so that they can be stacked in a rectangular shape. For example, if the color layer group and the alpha layer group are stacked vertically as shown in FIG. 11B, the alpha layer group 115 may be formed so that its horizontal length is the same as that of the color layer group 111. Also, if the color layer group and the alpha layer group are stacked horizontally as shown in FIG. 11C, the alpha layer group 116 may be formed so that its vertical length is the same as that of the color layer group 111.

[0063] The shape of the area to be coded must be rectangular. Therefore, by forming the layer group so that it can be stacked into a rectangle according to the value of the parameter "mpii_opacity_layer_div_scale_log2" (i.e., the resolution ratio), the layer group can be packed efficiently. In other words, the rectangular area to be coded can be filled with the layer group without leaving any empty space. This prevents unnecessary increases in the amount of data. In other words, if the existence of empty space is allowed within the rectangular area to be coded after packing, the color layer group and the alpha layer group do not need to be stacked into a rectangle.

[0064] By transmitting information about the relative resolution of the alpha layer relative to the color layer from the encoding side to the decoding side in this way, the decoder can easily grasp the resolution ratio between the color layer and the alpha layer based on that information. In other words, the resolution ratio between the color layer and the alpha layer can be made variable. For example, the resolution ratio between the color layer and the alpha layer can be optimized depending on the content, etc.

[0065] In this way, by varying the ratio of the resolution of the color layer to the alpha layer, it is possible to increase the resolution of the color layer without increasing the resolution of the alpha layer. For example, if the resolution of the alpha layer is sufficient, it is possible to suppress insufficient resolution of the color layer and prevent texture blurring without unnecessarily increasing processing time or data volume. Furthermore, even if the processing time or data volume reaches the system's upper limit, it is possible to suppress insufficient resolution of the color layer and prevent texture blurring.

[0066] That is, by transmitting information about the relative resolution of the alpha layer with respect to the color layer from the encoding side to the decoding side, it is possible to suppress a reduction in the subjective quality of the rendered image due to insufficient resolution of the color layer.

[0067] <Method 1-3> Furthermore, both the layer number ratio and the resolution ratio between color layers and alpha layers may be variable. That is, when Method 1 is applied, information regarding the relative number of alpha layers to color layers and information regarding the relative resolution may be transmitted (Method 1-3), as shown in the bottom row of the table in FIG. 7 . For example, in the first image processing device, the first image processing method, and the first program, the inter-layer relative information added to the bitstream may include information regarding the relative number of alpha component layers to color component layers and information regarding the relative resolution of the alpha component layers to the color component layers. Similarly, in the second image processing device, the second image processing method, and the second program, the inter-layer relative information added to the bitstream may include information regarding the relative number of alpha component layers to color component layers and information regarding the relative resolution of the alpha component layers to the color component layers.

[0068] In this case, too, the information regarding the relative number of alpha layers to color layers and the information regarding the relative resolution of the alpha layers to color layers may be added to the bitstream in any manner. For example, this information may be stored in an SEI added to the bitstream of the 3D data. Of course, other methods may also be applied.

[0069] For example, as shown in box 117 in A of Fig. 13, the parameters "mpii_opacity_layer_num_minus1" and "mpii_opacity_layer_div_scale_log2" may be stored in the SEI added to the bitstream of 3D data. These parameters are as described above in <Method 1-1> and <Method 1-2>.

[0070] By doing so, it is possible to obtain both the effects obtained by applying the above-mentioned method 1-1 and the effects obtained by applying the above-mentioned method 1-2.

[0071] Furthermore, in this case, since both the ratio of the number of layers and the resolution ratio between the color layers and the alpha layers are variable, a layer group consisting of alpha layers of different sizes from the color layers, such as the alpha layer group 118 shown in FIG. 13B, can be formed to the same size as the color layer group 111. By doing so, for example, not only can layers be stacked vertically as in the example shown in FIG. 13B or horizontally as in the example shown in FIG. 14A, but also in the time direction as in the example shown in FIG. 14B. In other words, the color layer group 111 and the alpha layer group 118 can be stacked as consecutive frames. In this way, the variety of stacks can be increased, allowing for more efficient packing.

[0072] <Method 2> The method described in Non-Patent Document 5 has the restriction that the layer arrangement pattern is fixed for 3D data configured with color component layer images and alpha component layer images as described above. For example, once the positions of the foremost and farthest layers are specified, the remaining layers are arranged at equal intervals between them.

[0073] However, in general, depth and parallax are inversely related. Therefore, in order to maintain the quality of an object regardless of its position in the depth direction, it is necessary to increase the density of layers toward the foreground. Therefore, when layers are arranged at equal intervals as described above, there is a risk that the quality of objects in the foreground may be reduced in scenes that include distant views, such as outdoors.

[0074] In addition, objects generally do not exist uniformly in the depth direction. Furthermore, even for a single object, the importance of its quality is not necessarily uniform in the depth direction. Therefore, arranging layers at equal intervals in the depth direction may result in a reduction in the quality of the rendered image.

[0075] For example, if you want to represent an infinitely distant background, such as the sky, you need to set the position of the deepest layer to infinity. In this case, even if other objects exist only within a short distance from the viewpoint, the layers are arranged at equal intervals over a wide range up to infinity, resulting in an insufficient number of layers allocated to the object positions, potentially making it difficult to represent the object's shape and texture with sufficient accuracy. Furthermore, for a single object, the quality of the near side, which is clearly visible from the viewpoint, has a greater impact on image quality than the quality of the far side. However, arranging layers at equal intervals in the depth direction can make it difficult to represent the shape and texture of the near side of the object with sufficient accuracy. This reduction in 3D data quality (reduced reproducibility) can potentially reduce the quality of the rendered image.

[0076] Increasing the number of layers can mitigate this degradation in 3D data quality, but it also increases the processing time and data volume for encoding, decoding, rendering, etc. In other words, simply increasing the total number of layers increases the layer density even at unnecessary depths. This means that an increase in unnecessary layers is required, which may unnecessarily increase processing time and data volume. Furthermore, systems generally have limitations on the processing time and data volume for encoding, decoding, rendering, etc., making it difficult to increase the number of layers without limit. Therefore, when the processing time and data volume reach the system's upper limit, it is difficult to mitigate the degradation in the subjective quality of the rendered image by arranging layers at equal intervals.

[0077] In the method described in Non-Patent Document 5, the position of each layer can be specified. This allows the layers to be arranged at non-equidistant intervals. However, with this method, the positions of all layers must be specified, which may unnecessarily increase the amount of data.

[0078] Therefore, in order to prevent a reduction in the quality of the rendered image due to such constraints, information regarding the arrangement of layers is transmitted as shown in the top row of the table in FIG. 15 (Method 2).

[0079] For example, a first image processing device may include a layer arrangement information generation unit that generates layer arrangement information regarding the arrangement of layers of 3D data composed of color component layer images and alpha component layer images, and an encoding unit that encodes the 3D data to generate a bitstream and adds the layer arrangement information to the bitstream. For example, a first image processing method may include the first image processing device generating layer arrangement information regarding the arrangement of layers of 3D data composed of color component layer images and alpha component layer images, and the first image processing device encoding the 3D data to generate a bitstream and adding the layer arrangement information to the bitstream. For example, a first program may cause the first image processing device to execute the steps of generating layer arrangement information regarding the arrangement of layers of 3D data composed of color component layer images and alpha component layer images, encoding the 3D data to generate a bitstream, and adding the layer arrangement information to the bitstream.

[0080] For example, the second image processing device may include a decoding unit that decodes a bitstream and generates 3D data composed of color component layer images and alpha component layer images, a layer arrangement information analysis unit that analyzes layer arrangement information regarding the arrangement of layers of the 3D data extracted from the bitstream, and a viewpoint image generation unit that renders the 3D data based on the layer arrangement information to generate a viewpoint image seen from an arbitrary viewpoint.For example, the second image processing method may include the second image processing device decoding the bitstream and generating 3D data composed of color component layer images and alpha component layer images, the second image processing device analyzing the layer arrangement information regarding the arrangement of layers of the 3D data extracted from the bitstream, and the second image processing device rendering the 3D data based on the layer arrangement information to generate a viewpoint image seen from an arbitrary viewpoint. For example, the second program causes the second image processing device to decode the bitstream, generate 3D data consisting of a layer image of a color component and a layer image of an alpha component, analyze layer arrangement information regarding the arrangement of layers of the 3D data extracted from the bitstream, render the 3D data based on the layer arrangement information, and generate a viewpoint image seen from an arbitrary viewpoint.

[0081] This layer placement information may be added to the bitstream in any manner. For example, this layer placement information may be stored in an SEI added to the bitstream of 3D data. Of course, other methods may also be applied.

[0082] By transmitting the layer arrangement information from the encoding side to the decoding side in this manner, the decoder can easily arrange each layer at an arbitrary position based on the layer arrangement information. Therefore, for example, it is possible to arrange each layer at non-equidistant intervals while suppressing an unnecessary increase in the amount of data. This makes it possible to optimize the layer arrangement according to the content, etc. Therefore, it is possible to suppress a reduction in the quality of the rendering image while suppressing an unnecessary increase in the amount of data.

[0083] <Method 2-1> For example, the layer arrangement information may include information about the arrangement of some layers of the 3D data. The some layers may be, for example, additional layers added to a basic layer group in which each layer is arranged according to a predetermined arrangement pattern. In other words, when Method 2 is applied, information about the arrangement of the additional layers may be transmitted (Method 2-1), as shown in the second row from the top of the table in FIG. 15 .

[0084] For example, in the first image processing device, the first image processing method, and the first program, the layer arrangement information added to the bitstream may include information regarding the arrangement of some layers of the 3D data. Similarly, in the second image processing device, the second image processing method, and the second program, the layer arrangement information added to the bitstream may include information regarding the arrangement of some layers of the 3D data.

[0085] For example, as shown in FIG. 16A, a group of layers of 3D data are arranged at a distance Z from the viewpoint. Near From Z Far The basic layer group is composed of layers 211 to 215 of the basic layer group, which are arranged in a range of 1 to 3, and additional layers 221 and 222. Additional layers 221 and 222 are layers other than the basic layer group (i.e., layers that do not belong to the basic layer group). For example, if each layer of the basic layer group (layers 211 to 215) is arranged according to a predetermined arrangement pattern, additional layers 221 and 222 are arranged in positions that do not follow the predetermined arrangement pattern.

[0086] By providing additional layers in addition to the base layer group in this way, it is possible to easily control the density of layers at each position in the depth direction as desired. For example, the additional layers can be used to easily increase the density of layers at positions in the depth direction where objects are present. It is also possible to easily place the base layer group in the range in the depth direction where objects are present, and place additional layers at infinity. This allows the number of layers to be increased at desired positions without increasing the number of layers at unnecessary positions, thereby preventing unnecessary increases in data volume and reducing degradation in the quality of the rendered image.

[0087] Note that, although the example of A in FIG. 16 shows a case where there are two additional layers (additional layer 221 and additional layer 222), any number of additional layers may be used. The layout pattern (predetermined layout pattern) of the base layer may be predetermined by a standard or the like, or may be specified by a user or the like. This predetermined layout pattern may be any pattern. For example, the layers may be arranged at equal intervals or at non-equidistant intervals. For example, a layout pattern based on a function may be used. For example, a layout pattern in which the distance between layers increases as the distance from the viewpoint increases (for example, an inverse proportional pattern or a logarithmic pattern) may be used.

[0088] As described above, information regarding the placement of the additional layers may be added to the bitstream as layer placement information.

[0089] For example, in the first image processing device, the first image processing method, and the first program, information about the arrangement of some layers added to the bitstream may include information indicating the positions of the some layers in the depth direction. Similarly, in the second image processing device, the second image processing method, and the second program, information about the arrangement of some layers added to the bitstream may include information indicating the positions of the some layers in the depth direction.

[0090] For example, in the first image processing device, the first image processing method, and the first program, information regarding the arrangement of some layers added to a bitstream may include information indicating the number of those layers. Similarly, in the second image processing device, the second image processing method, and the second program, information regarding the arrangement of some layers added to a bitstream may include information indicating the number of those layers.

[0091] For example, each parameter may be stored in an SEI added to a bitstream of 3D data using a syntax such as that shown in box 231 in B of FIG. 16. The value of the parameter "mpii_num_layers_explicit_depth" indicates the number of some layers (e.g., additional layers). The parameter "depth_rep_info_element" indicates the depth position of some layers (e.g., additional layers). In other words, in the layer placement information of this example, the position of each additional layer is specified.

[0092] In this way, by transmitting information about the arrangement of some layers (e.g., additional layers) from the encoding side to the decoding side, the decoder can easily set the some layers (additional layers) to positions other than those conforming to the predetermined arrangement pattern of the base layer group. Therefore, the density of layers at each position in the depth direction can be arbitrarily controlled, and it is possible to suppress a reduction in the quality of the rendered image while suppressing an unnecessary increase in the amount of data.

[0093] <Method 2-2> For example, each layer of a group of layers constituting 3D data may be arranged at non-equidistant intervals. In other words, the arrangement pattern of the group of layers constituting 3D data may be a non-equidistant arrangement pattern in which each layer is arranged at non-equidistant intervals. In other words, a non-equidistant arrangement pattern may be specified as the arrangement pattern of the group of layers constituting 3D data, and information specifying the non-equidistant arrangement pattern may be transmitted as layer arrangement information.

[0094] This non-uniform spacing pattern may be specified by any method. For example, this non-uniform spacing pattern may be specified using a predetermined function. Hereinafter, a function that specifies a non-uniform spacing pattern is also referred to as a non-uniform spacing function. In other words, a non-uniform spacing function is a function that arranges layers at non-uniform intervals. In other words, when Method 2 is applied, information regarding the arrangement of layers based on the non-uniform spacing function may be transmitted (as layer arrangement information) (Method 2-2), as shown in the third row from the top of the table in FIG. 15 .

[0095] For example, in the first image processing device, the first image processing method, and the first program, the layer placement information added to the bitstream may include information about placement of layers of the 3D data based on a non-uniform placement function. Similarly, in the second image processing device, the second image processing method, and the second program, the layer placement information added to the bitstream may include information about placement of layers of the 3D data based on a non-uniform placement function.

[0096] This non-uniform spacing function may be any function. For example, it may be a function in which layers are denser (i.e., the layer density is higher) toward the front. In other words, it may be a function in which layers are sparser (i.e., the layer density is lower) toward the back. For example, suppose that indices are assigned to each layer in ascending order from the front to the back. It may be a function (hereinafter also referred to as an inverse proportional spacing function) that arranges each layer in an arrangement pattern (hereinafter also referred to as an inverse proportional pattern) in which the difference in index between the innermost layer and the layer in question is inversely proportional to the spacing between the layer and its adjacent layer. It may also be a function (hereinafter also referred to as a logarithmic spacing function) that arranges each layer in an arrangement pattern (hereinafter also referred to as a logarithmic pattern) in which the spacing between layers exponentially increases depending on the index value.

[0097] For example, the relationship between the layer index and the layer depth (position in the depth direction) is shown in the graph of FIG. 17. In the graph of FIG. 17, a line 241 shows the relationship between the index (i) and the depth (Z) of each layer when an equal spacing function is applied, which specifies an equal spacing pattern that arranges each layer at equal intervals. Note that the index (i) of each layer is normalized between "0" and "1." A curve 242 shows the relationship between the index (i) and the depth (Z) of each layer when an inverse proportional placement function is applied. A curve 243 shows the relationship between the index (i) and the depth (Z) of each layer when a logarithmic placement function is applied.

[0098] For example, the depth (Z(i)) of each layer (i) may be expressed as in the following equation (1).

[0099] Z(i) = g(f(Z Near ) + i・(f(Z Far ) - f(Z Near ))) ... (1)

[0100] In the case of a uniform spacing function, f(x) = x, g(x) = x. In the case of an inverse proportional spacing function, f(x) = 1 / x, g(x) = 1 / x. In the case of a logarithmic spacing function, f(x) = log(x), g(x) = exp(x). By specifying the placement pattern using such functions, the position of each layer can be specified without increasing the amount of data. Furthermore, each layer can be easily placed using a predetermined placement pattern. Furthermore, by applying an inverse proportional spacing function or a logarithmic spacing function, it is possible to easily make layers denser (i.e., increase the layer density) toward the foreground. In other words, it is possible to easily make layers sparser (i.e., decrease the layer density) toward the background. Therefore, it is possible to suppress the reduction in quality of foreground objects due to the inverse relationship between depth and parallax.

[0101] <Method 2-2-1> The layer arrangement pattern as described above may be predetermined, or may be set by a user, etc. For example, when Method 2-2 is applied, information indicating whether the arrangement of layers is based on a non-uniform spacing function may be transmitted (as layer arrangement information) as shown in the fourth row from the top of the table in Fig. 15 (Method 2-2-1).

[0102] For example, in the first image processing device, the first image processing method, and the first program, the "information relating to placement based on a non-uniform placement function" added to the bitstream may include information indicating whether the layer is placed based on the non-uniform placement function. Similarly, in the second image processing device, the second image processing method, and the second program, the "information relating to placement based on a non-uniform placement function" added to the bitstream may include information indicating whether the layer is placed based on the non-uniform placement function.

[0103] This "information indicating whether layers are arranged based on their non-uniform spacing functions" may be information specifying a mode for arranging layers.

[0104] For example, in the first image processing device, the first image processing method, and the first program, the "information relating to placement based on a non-uniformly spaced function" added to the bitstream may include information specifying the mode of the layer placement method. Similarly, in the second image processing device, the second image processing method, and the second program, the "information relating to placement based on a non-uniformly spaced function" added to the bitstream may include information specifying the mode of the layer placement method.

[0105] For example, the parameter "mpii_layer_depth_type" as shown in a square 251 in A of FIG. 18 may be transmitted. This parameter "mpii_layer_depth_type" is information specifying a mode of layer arrangement. For example, the parameter "mpii_layer_depth_type" may specify a mode to be applied depending on its value, as shown in table 252 in A of FIG. 18. For example, when the value of the parameter "mpii_layer_depth_type" is "0," this indicates that a mode in which layers are arranged based on an equal spacing function is specified. When the value of the parameter "mpii_layer_depth_type" is "1," this indicates that a mode in which the positions of all layers are specified is specified. When the value of the parameter "mpii_layer_depth_type" is "2," this indicates that a mode in which layers are arranged based on an inverse proportional arrangement function (Inverse) is specified. When the value of the parameter "mpii_layer_depth_type" is "3," this indicates that a mode in which layers are arranged based on a logarithmic arrangement function (Logarithm) is specified. In other words, the parameter "mpii_layer_depth_type" can be said to be information indicating whether or not layers are arranged based on the non-uniform spacing function.

[0106] By transmitting this parameter "mpii_layer_depth_type" from the encoding side to the decoding side as "information about placement based on a non-uniform placement function," the decoder can place layers in the mode specified by the parameter "mpii_layer_depth_type." This allows for a wider variety of placement patterns to be realized, enabling layer placement to be optimized in a wider variety of cases. In other words, in a wider variety of cases, it is possible to suppress unnecessary increases in data volume while suppressing degradation in the quality of rendered images.

[0107] <Method 2-2-2> Furthermore, when Method 2-2 is applied, as shown in the fifth row from the top of the table in FIG. 15 , information indicating whether or not the interval is not specified, information indicating whether or not the pattern arrangement is nonlinear, and information indicating whether or not the pattern arrangement is inversely proportional may be transmitted (as layer arrangement information) (Method 2-2-2).

[0108] For example, in the first image processing device, the first image processing method, and the first program, the "information regarding placement based on a non-uniformly spaced function" added to the bitstream may include information indicating whether the layer is placed using an inversely proportional pattern. Similarly, in the second image processing device, the second image processing method, and the second program, the "information regarding placement based on a non-uniformly spaced function" added to the bitstream may include information indicating whether the layer is placed using an inversely proportional pattern.

[0109] In the first image processing device, the first image processing method, and the first program, the "information regarding placement based on a non-uniform placement function" added to the bitstream may include information indicating whether the layer is placed using a non-linear pattern. Similarly, in the second image processing device, the second image processing method, and the second program, the "information regarding placement based on a non-uniform placement function" added to the bitstream may include information indicating whether the layer is placed using a non-linear pattern.

[0110] In the first image processing device, the first image processing method, and the first program, the "information relating to placement based on non-uniformly spaced functions" added to the bitstream may include information indicating whether the spacing between layers is not specified. Similarly, in the second image processing device, the second image processing method, and the second program, the "information relating to placement based on non-uniformly spaced functions" added to the bitstream may include information indicating whether the spacing between layers is not specified.

[0111] For example, the parameter "mpii_layer_depth_equal_distance_flag", the parameter "mpii_layer_depth_nonlinear_flag", and the parameter "mpii_layer_depth_inverse_flag" may be transmitted as "information regarding arrangement based on a non-uniform arrangement function".

[0112] The parameter "mpii_layer_depth_equal_distance_flag" is flag information indicating whether the layer spacing is not specified. If this flag information is true (for example, "1"), it indicates that the layer spacing is not specified. If this flag information is false (for example, "0"), it indicates that the layer spacing is specified. For example, if this flag information is false, it may indicate that a mode specifying the positions of all layers is specified.

[0113] The parameter "mpii_layer_depth_nonlinear_flag" is flag information indicating whether layers are arranged using a nonlinear pattern. A nonlinear pattern is an arrangement pattern in which each layer is arranged at nonlinear intervals. If this flag information is true (e.g., "1"), it indicates that layers are arranged using a nonlinear pattern. If this flag information is false (e.g., "0"), it indicates that layers are not arranged using a nonlinear pattern. For example, if this flag information is false, it may indicate that a mode in which layers are arranged based on an equal spacing function is specified.

[0114] The parameter "mpii_layer_depth_inverse_flag" is flag information indicating whether layers are arranged using an inverse proportional pattern. If this flag information is true (e.g., "1"), it indicates that layers are arranged using an inverse proportional pattern. If this flag information is false (e.g., "0"), it indicates that layers are not arranged using an inverse proportional pattern. In other words, if this flag information is true, it may indicate that a mode in which layers are arranged based on an inverse proportional arrangement function (Inverse) is specified. Also, if this flag information is false, it may indicate that a mode in which layers are arranged based on a logarithmic arrangement function (Logarithm) is specified.

[0115] These parameters may be stored in an SEI added to the bitstream of 3D data using a syntax such as that shown in a box 253 in B of Fig. 18. That is, if the parameter "mpii_layer_depth_equal_distance_flag" is true, the parameter "mpii_layer_depth_nonlinear_flag" may be set, and if the value of the parameter "mpii_layer_depth_inverse_flag" is true, the parameter "mpii_layer_depth_equal_distance_flag" may be set. If the parameter "mpii_layer_depth_equal_distance_flag" is false, a mode specifying the positions of all layers may be specified; if the parameter "mpii_layer_depth_nonlinear_flag" is false, a mode specifying the placement of layers based on an equal spacing function may be specified; if the parameter "mpii_layer_depth_inverse_flag" is true, a mode specifying the placement of layers based on an inverse proportional placement function (Inverse) may be specified; and if the parameter "mpii_layer_depth_inverse_flag" is false, a mode specifying the placement of layers based on a logarithmic placement function (Logarithm) may be specified.

[0116] By transmitting these parameters from the encoding side to the decoding side as "information about placement based on a non-uniform placement function," the decoder can place layers in the mode indicated by these parameters. This allows for a wider variety of placement patterns to be realized, enabling layer placement to be optimized in a wider variety of cases. In other words, it is possible to prevent unnecessary increases in data volume and to prevent degradation in the quality of rendered images in a wider variety of cases.

[0117] <Method 2-2-3> Note that the parameter "mpii_layer_depth_explicit_distance_flag" may be applied instead of the parameter "mpii_layer_depth_equal_distance_flag." The parameter "mpii_layer_depth_explicit_distance_flag" is flag information indicating whether the layer spacing is to be specified. When this flag information is true (e.g., "1"), it indicates that the layer spacing is specified. When this flag information is false (e.g., "0"), it indicates that the layer spacing is not specified. For example, when this flag information is true, it may indicate that a mode specifying the positions of all layers is specified. In other words, when Method 2-2 is applied, as shown in the bottom row of the table in FIG. 15 , information indicating whether the spacing is to be specified, information indicating whether a nonlinear pattern arrangement is used, and information indicating whether an inverse proportional pattern arrangement is used may be transmitted (as layer arrangement information) (Method 2-2-3).

[0118] For example, in the first image processing device, the first image processing method, and the first program, the "information relating to placement based on non-uniformly spaced functions" added to the bitstream may include information indicating whether the spacing between layers is specified. Similarly, in the second image processing device, the second image processing method, and the second program, the "information relating to placement based on non-uniformly spaced functions" added to the bitstream may include information indicating whether the spacing between layers is specified.

[0119] This parameter may be stored in an SEI added to the bitstream of the 3D data in a syntax such as that shown in a box 254 C in Fig. 18. That is, when the parameter "mpii_layer_depth_explicit_distance_flag" is false, the parameter "mpii_layer_depth_nonlinear_flag" may be set, and when the value of the parameter "mpii_layer_depth_inverse_flag" is true, the parameter "mpii_layer_depth_explicit_distance_flag" may be set. Then, when the parameter "mpii_layer_depth_explicit_distance_flag" is true, a mode specifying the positions of all layers may be specified; when the parameter "mpii_layer_depth_nonlinear_flag" is false, a mode specifying the placement of layers based on a uniform spacing function may be specified; when the parameter "mpii_layer_depth_inverse_flag" is true, a mode specifying the placement of layers based on an inverse proportional placement function (Inverse) may be specified; and when the parameter "mpii_layer_depth_inverse_flag" is false, a mode specifying the placement of layers based on a logarithmic placement function (Logarithm) may be specified.

[0120] By transmitting these parameters from the encoding side to the decoding side as "information about placement based on a non-uniform placement function," the decoder can place layers in the mode indicated by these parameters. This allows for a wider variety of placement patterns to be realized, enabling layer placement to be optimized in a wider variety of cases. In other words, it is possible to prevent unnecessary increases in data volume and to prevent degradation in the quality of rendered images in a wider variety of cases.

[0121] <Method 3> The method described in Non-Patent Document 5 has the constraint that the shape of the layers is fixed for 3D data composed of color component layer images and alpha component layer images, as described above. In this method, MPI is applied to the 3D data, and the shape of the layers is planar. Therefore, in scenes with a wide angle of view, the resolution decreases near the center of the distant image, which could reduce the quality of the rendered image including that part.

[0122] Therefore, in order to prevent a reduction in the quality of the rendered image due to such constraints, information regarding the selection of the layer shape is transmitted (Method 3), as shown in the top row of the table in Fig. 19. In other words, it may be possible to select the shape of the layer.

[0123] For example, a first image processing device may include a layer shape selection information generation unit that generates layer shape selection information regarding a selection of a layer shape of 3D data composed of color component layer images and alpha component layer images, and an encoding unit that encodes the 3D data to generate a bitstream and adds the layer shape selection information to the bitstream. For example, a first image processing method may include the first image processing device generating layer shape selection information regarding a selection of a layer shape of 3D data composed of color component layer images and alpha component layer images, and the first image processing device encoding the 3D data to generate a bitstream and adding the layer shape selection information to the bitstream. For example, a first program may cause the first image processing device to execute the steps of generating layer shape selection information regarding a selection of a layer shape of 3D data composed of color component layer images and alpha component layer images, encoding the 3D data to generate a bitstream, and adding the layer shape selection information to the bitstream.

[0124] For example, the second image processing device may include a decoding unit that decodes a bitstream and generates 3D data composed of color component layer images and alpha component layer images, a layer shape selection information analysis unit that analyzes layer shape selection information extracted from the bitstream regarding selection of a layer shape of the 3D data, and a viewpoint image generation unit that renders the 3D data based on the layer shape selection information to generate a viewpoint image seen from an arbitrary viewpoint.For example, the second image processing method may include the second image processing device decoding the bitstream and generating 3D data composed of color component layer images and alpha component layer images, the second image processing device analyzing layer shape selection information extracted from the bitstream regarding selection of a layer shape of the 3D data, and the second image processing device rendering the 3D data based on the layer shape selection information to generate a viewpoint image seen from an arbitrary viewpoint. For example, the second program causes the second image processing device to decode the bitstream and generate 3D data consisting of a layer image of a color component and a layer image of an alpha component, analyze layer shape selection information extracted from the bitstream regarding the selection of the shape of the layer of the 3D data, render the 3D data based on the layer shape selection information, and generate a viewpoint image viewed from any viewpoint.

[0125] As shown in A of FIG. 20 , the MPI is composed of planar layers. If the resolution of each layer is constant, there is a risk of resolution reduction near the distant center in a scene with a wide angle of view. In the example of A of FIG. 20 , when viewing the center of the scene from viewpoint 311, the rectangle 312 is larger in the layer at the back than in the layer at the front. This rectangle 312 is a schematic representation of the pixels of each layer. Therefore, if the coverage area of ​​each layer is constant, the size of the rectangle 312 of each layer will be the same. However, because the layers are planar, the rectangle 312 becomes larger the farther the layer is from viewpoint 311. In other words, the coverage area of ​​one pixel becomes wider. In other words, the number of pixels included in the angle of view of the viewpoint image of viewpoint 311 becomes smaller the farther the layer is from viewpoint 311. In other words, the resolution decreases the farther the layer is from viewpoint 311. This could result in a reduction in the quality of the rendering image including that portion.

[0126] Therefore, as shown in B of Fig. 20 , the layer may have a curved surface shape. In the example of B of Fig. 20 , the layer is curved with a predetermined curvature, and the difference in size of the rectangle 312 between the front layer and the back layer is smaller than in the example of A of Fig. 20 . In other words, the difference in the range covered by one pixel between the front layer and the back layer is smaller than in the example of A of Fig. 20 . Reduction in the resolution of layers far from the viewpoint 311, which are included in the angle of view of the viewpoint image of the viewpoint 311, is suppressed. Therefore, it is possible to suppress reduction in the quality of the rendering image including that portion.

[0127] Examples of 3D data composed of curved layers include MCI (Multi Cylinder Image), which is composed of cylindrical layers, and MSI (Multi Sphere Image), which is composed of spherical layers. MCI layers can be said to be curved in one direction, while MSI layers can be said to be curved in two directions.

[0128] As mentioned above, MCI and MSI can suppress the reduction in resolution near the center of distant objects more than MPI. However, MPI makes it easier to set up layers and suppresses the increase in processing load.

[0129] As described above, by transmitting information regarding layer shape selection and enabling selection of layer shapes, it is possible to select and apply the method that is best suited to the content from, for example, MPI, MCI, and MSI. This makes it possible to suppress unnecessary increases in the processing load of encoding, decoding, playback, etc., while also suppressing degradation in the quality of rendered images.

[0130] <Method 3-1> For example, when Method 3 is applied, as shown in the second row from the top of the table in Figure 19, curvature designation information in a single direction and information indicating whether the shape is cylindrical may be transmitted (Method 3-1).

[0131] For example, in the first image processing device, the first image processing method, and the first program, the layer shape selection information added to the bitstream may include curvature designation information that designates the curvature of the layer in a single direction and information indicating whether the layer is cylindrical. Similarly, in the second image processing device, the second image processing method, and the second program, the layer shape selection information added to the bitstream may include curvature designation information that designates the curvature of the layer in a single direction and information indicating whether the layer is cylindrical.

[0132] For example, as shown in a box 321 in A of Fig. 21 , the parameters "mpii_layer_curvature" and "mpii_cylindrical_layer_flag" may be added to the bitstream as layer shape selection information. The parameter "mpii_layer_curvature" is curvature specification information that specifies the curvature of the layer in a single direction. The parameter "mpii_cylindrical_layer_flag" is flag information that indicates whether the layer is cylindrical. If this flag information is true (e.g., "1"), it indicates that the layer is cylindrical (i.e., MCI). If this flag information is false (e.g., "0"), it indicates that the layer is not cylindrical (i.e., not MCI).

[0133] These parameters may be stored in an SEI added to the bitstream of 3D data according to the syntax shown in box 321A in FIG. 21 . That is, the parameter “mpii_layer_curvature” may be set. A value of “0” for this parameter indicates that the layer is planar (i.e., MPI). Conversely, a value of this parameter greater than “0” indicates that the layer is curved. In this case, the parameter “mpii_cylindrical_layer_flag” may be set. A value of this flag indicates that the layer is cylindrical (i.e., MCI). In this case, the curvature indicated by the parameter “mpii_layer_curvature” may be applied to the curvature direction of the layer. This curvature direction may be any single direction. For example, it may be horizontal, vertical, or another direction. The curvature direction may be predetermined or may be specified by a user, etc. Furthermore, a value of “false” for the parameter “mpii_cylindrical_layer_flag” indicates that the layer is not cylindrical (i.e., not MCI). For example, if this flag information is false, it may indicate that the layer is spherical (i.e., MSI). In this case, the curvature indicated by the parameter "mpii_layer_curvature" may be applied to the curvature direction of the layer. The curvature direction may be any direction as long as it is multiple directions (e.g., two directions perpendicular to each other). For example, it may be horizontal and vertical, or may be another direction. The curvature direction may be predetermined or may be specified by a user or the like.

[0134] <Method 3-2> For example, when Method 3 is applied, as shown in the third row from the top of the table in FIG. 19 , information indicating whether the surface is non-planar, curvature designation information in a single direction, and information indicating whether the surface is cylindrical may be transmitted (Method 3-2).

[0135] For example, in the first image processing device, the first image processing method, and the first program, the layer shape selection information added to the bitstream may include information indicating whether the layer is non-planar, curvature designation information specifying the curvature of the layer in a single direction, and information indicating whether the layer is cylindrical. Similarly, in the second image processing device, the second image processing method, and the second program, the layer shape selection information added to the bitstream may include information indicating whether the layer is non-planar, curvature designation information specifying the curvature of the layer in a single direction, and information indicating whether the layer is cylindrical.

[0136] For example, as shown in a box 322 in B of Fig. 21 , the parameters "mpii_non_plane_flag", "mpii_layer_curvature", and "mpii_cylindrical_layer_flag" may be added to the bitstream as layer shape selection information. The parameter "mpii_non_plane_flag" is flag information indicating whether or not the layer is non-planar. If this flag information is true (e.g., "1"), it indicates that the layer is non-planar. If this flag information is false (e.g., "0"), it indicates that the layer is not non-planar (i.e., is an MPI). The parameters "mpii_layer_curvature" and "mpii_cylindrical_layer_flag" are the information as described above.

[0137] These parameters may be stored in an SEI added to the bitstream of 3D data according to the syntax shown in box 322 in FIG. 21B. That is, the parameter "mpii_non_plane_flag" may be set. When this flag information is false, it indicates that the layer is planar (i.e., MPI). On the other hand, when this flag information is true, it indicates that the layer is non-planar. In this case, the parameters "mpii_layer_curvature" and "mpii_cylindrical_layer_flag" may be set. When the parameter "mpii_cylindrical_layer_flag" is true, it indicates that the layer is cylindrical. Therefore, the curvature indicated by the parameter "mpii_layer_curvature" may be applied to a single curvature direction of the layer. This curvature direction may be any single direction. For example, it may be horizontal, vertical, or another direction. Furthermore, this curvature direction may be predetermined or may be specified by a user or the like. Furthermore, if the parameter "mpii_cylindrical_layer_flag" is false, it indicates that the layer is spherical. Therefore, the curvature indicated by the parameter "mpii_layer_curvature" may be applied to multiple curvature directions of the layer. The curvature directions may be any direction as long as they are multiple directions (e.g., two directions perpendicular to each other). For example, the curvature directions may be horizontal and vertical, or may be other directions. The curvature directions may be predetermined or may be specified by a user or the like.

[0138] <Method 3-3> For example, when Method 3 is applied, curvature designation information in a plurality of directions may be transmitted as shown in the fourth row from the top of the table in FIG. 19 (Method 3-3).

[0139] For example, in the first image processing device, the first image processing method, and the first program, the layer shape selection information added to the bitstream may include curvature designation information that designates the curvature of the layer in multiple directions. Similarly, in the second image processing device, the second image processing method, and the second program, the layer shape selection information added to the bitstream may include curvature designation information that designates the curvature of the layer in multiple directions.

[0140] For example, as shown in a square 331 in A of FIG. 22 , the parameters "mpii_layer_curvature_h" and "mpii_layer_curvature_v" may be added to the bitstream as layer shape selection information. The parameter "mpii_layer_curvature_h" is curvature specification information that specifies the curvature of the layer in the horizontal direction. The parameter "mpii_layer_curvature_v" is curvature specification information that specifies the curvature of the layer in the vertical direction. In other words, the parameters "mpii_layer_curvature_h" and "mpii_layer_curvature_v" may be set, and the curvatures indicated by these parameters may be applied to the horizontal and vertical directions of the layer. When both of these parameters have a value of "0," this indicates that the layer is planar (i.e., MPI). When one of these parameters has a value of "0" and the other has a value other than "0," this indicates that the layer has a cylindrical shape (i.e., MCI). A non-zero value for both of these parameters indicates that the layer is spherical (i.e., MSI).

[0141] <Method 3-4> For example, when Method 3 is applied, as shown in the fifth row from the top of the table in Figure 19, information indicating whether the surface is non-planar or not and curvature specification information in multiple directions may be transmitted (Method 3-4).

[0142] For example, in the first image processing device, the first image processing method, and the first program, the layer shape selection information added to the bitstream may include information indicating whether the layer is non-planar and curvature designation information that designates the curvature of the layer in multiple directions. Similarly, in the second image processing device, the second image processing method, and the second program, the layer shape selection information added to the bitstream may include information indicating whether the layer is non-planar and curvature designation information that designates the curvature of the layer in multiple directions.

[0143] For example, as shown in a box 332 in Fig. 22B, the parameters "mpii_non_plane_flag", "mpii_layer_curvature_h", and "mpii_layer_curvature_v" may be added to the bitstream as layer shape selection information. These parameters are the information described above.

[0144] These parameters may be stored in an SEI added to the 3D data bitstream according to the syntax shown in box 332 in FIG. 22B. That is, the parameter "mpii_non_plane_flag" may be set. When this flag information is false, it indicates that the layer is planar (i.e., MPI). On the other hand, when this flag information is true, it indicates that the layer is non-planar. In this case, the parameters "mpii_layer_curvature_h" and "mpii_layer_curvature_v" may be set, and the curvatures indicated by these parameters may be applied to the horizontal and vertical directions of the layer. When one of these parameters has a value of "0" and the other has a non-zero value, it indicates that the layer is cylindrical (i.e., MCI). When both of these parameters have a value of non-zero, it indicates that the layer is spherical (i.e., MSI).

[0145] <Method 3-5> For example, when Method 3 is applied, as shown in the bottom row of the table in Figure 19, information indicating whether the shape is non-planar or not, information indicating whether the shape is cylindrical or not, and curvature designation information in one or more directions may be transmitted (Method 3-5).

[0146] For example, in the first image processing device, the first image processing method, and the first program, the layer shape selection information added to the bitstream may include information indicating whether the layer is non-planar, information indicating whether the layer is cylindrical, and curvature designation information that specifies the curvature of the layer in multiple directions. Similarly, in the second image processing device, the second image processing method, and the second program, the layer shape selection information added to the bitstream may include information indicating whether the layer is non-planar, information indicating whether the layer is cylindrical, and curvature designation information that specifies the curvature of the layer in multiple directions.

[0147] For example, as shown in a box 333 C in Fig. 22, the parameters "mpii_non_plane_flag", "mpii_cylindrical_layer_flag", and "mpii_layer_curvature", or the parameters "mpii_layer_curvature_h" and "mpii_layer_curvature_v" may be added to the bitstream as layer shape selection information. These parameters are the information described above.

[0148] These parameters may be stored in an SEI added to the bitstream of 3D data according to the syntax shown in box 333 C of FIG. 23 . That is, the parameter “mpii_non_plane_flag” may be set. When this flag information is false, it indicates that the layer is planar (i.e., MPI). On the other hand, when this flag information is true, it indicates that the layer is non-planar. In this case, the parameter “mpii_cylindrical_layer_flag” may be set. When this parameter “mpii_cylindrical_layer_flag” is true, it indicates that the layer is cylindrical. In this case, the parameter “mpii_layer_curvature” may be set, and the curvature indicated by the parameter “mpii_layer_curvature” may be applied to a single curvature direction of the layer. This curvature direction may be any single direction. For example, it may be horizontal, vertical, or another direction. Furthermore, this curvature direction may be predetermined or may be specified by a user or the like. Additionally, if the parameter "mpii_cylindrical_layer_flag" is false, it indicates that the layer is spherical. In that case, the parameters "mpii_layer_curvature_h" and "mpii_layer_curvature_v" may be set, and the curvature indicated by those parameters may be applied to the horizontal and vertical directions of the layer.

[0149] By applying any of the above methods 3-1 to 3-5, it is possible to select and apply the method that is best suited to the content from, for example, MPI, MCI, or MSI, thereby suppressing unnecessary increases in the processing load of encoding, decoding, playback, etc., while also suppressing degradation in the quality of rendered images.

[0150] <Application of Each Method> Each of the above-described methods can be applied in combination with other methods, as long as no contradictions arise. For example, two or more of the methods shown in the table of FIG. 7 may be applied in appropriate combination. Also, two or more of the methods shown in the table of FIG. 15 may be applied in appropriate combination. Also, two or more of the methods shown in the table of FIG. 19 may be applied in appropriate combination. Also, two or more of Method 1 shown in the table of FIG. 7, Method 2 shown in the table of FIG. 15, or Method 3 shown in the table of FIG. 19 may be applied in appropriate combination. Also, a lower-level method shown in a table of each figure may be applied in appropriate combination with a method shown in a table other than that figure. Of course, these methods may be applied in combination with other methods not shown.

[0151] In this specification, a description of a higher-level method may include a description of a lower-level method. For example, a description such as "applying method 1" may also include the application of one or more of methods 1-1, 1-2, and 1-3. Similarly, a description such as "applying method 2" may also include the application of one or more of methods 2-1, 2-2, 2-2-1, 2-2-2, and 2-2-3. Similarly, a description such as "applying method 2-2" may also include the application of one or more of methods 2-2-1, 2-2-2, and 2-2-3. Similarly, a description such as "applying method 3" may also include the application of one or more of methods 3-1, 3-2, 3-3, 3-4, and 3-5.

[0152] Therefore, for example, Method 1 shown in the table of FIG. 7 and Method 2 shown in the table of FIG. 15 may be applied in combination.

[0153] For example, the first image processing device may include an inter-layer relative information generation unit that generates inter-layer relative information regarding a relative relationship between layers of 3D data configured from layer images of color components and layer images of an alpha component, a layer arrangement information generation unit that generates layer arrangement information regarding an arrangement of layers of the 3D data, and an encoding unit that encodes the 3D data to generate a bitstream and adds the inter-layer relative information and the layer arrangement information to the bitstream.For example, the first image processing method may include the first image processing device generating inter-layer relative information regarding a relative relationship between layers of 3D data configured from layer images of color components and layer images of an alpha component, the first image processing device generating layer arrangement information regarding the arrangement of layers of the 3D data, and the first image processing device encoding the 3D data to generate a bitstream and adding the inter-layer relative information and the layer arrangement information to the bitstream. For example, the first program may cause the first image processing device to generate inter-layer relative information regarding the relative relationship between layers of 3D data composed of color component layer images and alpha component layer images, generate layer arrangement information regarding the arrangement of the layers of the 3D data, encode the 3D data to generate a bitstream, and add the inter-layer relative information and layer arrangement information to the bitstream.

[0154] For example, the second image processing device may include a decoding unit that decodes a bitstream and generates 3D data composed of color component layer images and alpha component layer images; an inter-layer relative information analysis unit that analyzes inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream; a layer arrangement information analysis unit that analyzes layer arrangement information regarding the arrangement of layers of the 3D data extracted from the bitstream; and a viewpoint image generation unit that renders the 3D data based on the inter-layer relative information and layer arrangement information and generates a viewpoint image seen from any viewpoint. For example, the second image processing method may include the second image processing device decoding the bitstream to generate 3D data composed of color component layer images and alpha component layer images, the second image processing device analyzing inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream, the second image processing device analyzing layer arrangement information regarding the arrangement of layers of the 3D data extracted from the bitstream, and the second image processing device rendering the 3D data based on the inter-layer relative information and layer arrangement information to generate a viewpoint image seen from an arbitrary viewpoint. For example, the second program may cause the second image processing device to decode the bitstream, generate 3D data composed of color component layer images and alpha component layer images, analyze inter-layer relative information extracted from the bitstream regarding the relative relationship between layers of the 3D data, analyze layer arrangement information extracted from the bitstream regarding the arrangement of layers of the 3D data, render the 3D data based on the inter-layer relative information and layer arrangement information, and generate a viewpoint image seen from any viewpoint.

[0155] Also, Method 1 shown in the table of FIG. 7 and Method 3 shown in the table of FIG. 19 may be applied in combination.

[0156] For example, a first image processing device may include an inter-layer relative information generation unit that generates inter-layer relative information regarding a relative relationship between layers of 3D data configured from layer images of color components and layer images of an alpha component, a layer shape selection information generation unit that generates layer shape selection information regarding a selection of a shape of a layer of the 3D data, and an encoding unit that encodes the 3D data to generate a bitstream and adds the inter-layer relative information and the layer shape selection information to the bitstream.For example, a first image processing method may include the first image processing device generating inter-layer relative information regarding a relative relationship between layers of 3D data configured from layer images of color components and layer images of an alpha component, the first image processing device generating layer shape selection information regarding a selection of a shape of a layer of the 3D data, and the first image processing device encoding the 3D data to generate a bitstream and adding the inter-layer relative information and the layer shape selection information to the bitstream. For example, the first program may cause the first image processing device to generate inter-layer relative information regarding the relative relationship between layers of 3D data composed of color component layer images and alpha component layer images, generate layer shape selection information regarding selection of the shape of the layers of the 3D data, encode the 3D data to generate a bitstream, and add the inter-layer relative information and layer shape selection information to the bitstream.

[0157] For example, the second image processing device may include a decoding unit that decodes a bitstream and generates 3D data composed of color component layer images and alpha component layer images; an inter-layer relative information analysis unit that analyzes inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream; a layer shape selection information analysis unit that analyzes layer shape selection information regarding the selection of the shape of layers of the 3D data extracted from the bitstream; and a viewpoint image generation unit that renders the 3D data based on the inter-layer relative information and layer shape selection information and generates a viewpoint image seen from any viewpoint. For example, the second image processing method may include the second image processing device decoding the bitstream to generate 3D data composed of color component layer images and alpha component layer images, the second image processing device analyzing inter-layer relative information extracted from the bitstream regarding the relative relationship between layers of the 3D data, the second image processing device analyzing layer shape selection information extracted from the bitstream regarding the selection of the shape of the layers of the 3D data, and the second image processing device rendering the 3D data based on the inter-layer relative information and layer shape selection information to generate a viewpoint image seen from any viewpoint. For example, the second program may cause the second image processing device to decode the bitstream and generate 3D data composed of color component layer images and alpha component layer images, analyze inter-layer relative information extracted from the bitstream regarding the relative relationship between layers of the 3D data, analyze layer shape selection information extracted from the bitstream regarding the selection of the shape of layers of the 3D data, render the 3D data based on the inter-layer relative information and layer shape selection information, and generate a viewpoint image seen from any viewpoint.

[0158] Also, method 1 shown in the table of FIG. 7, method 2 shown in the table of FIG. 15, and method 3 shown in the table of FIG. 19 may be applied in combination.

[0159] For example, the first image processing device may include an inter-layer relative information generation unit that generates inter-layer relative information regarding the relative relationship between layers of 3D data composed of color component layer images and alpha component layer images, a layer arrangement information generation unit that generates layer arrangement information regarding the arrangement of layers of the 3D data, a layer shape selection information generation unit that generates layer shape selection information regarding selection of the shape of layers of the 3D data, and an encoding unit that encodes the 3D data to generate a bitstream and adds the inter-layer relative information, layer arrangement information, and layer shape selection information to the bitstream. For example, the first image processing method may include: a first image processing device generating inter-layer relative information regarding a relative relationship between layers of 3D data composed of color component layer images and alpha component layer images; a first image processing device generating layer arrangement information regarding an arrangement of layers of the 3D data; a first image processing device generating layer shape selection information regarding a selection of a shape of a layer of the 3D data; and a first image processing device encoding the 3D data to generate a bitstream, and adding the inter-layer relative information, layer arrangement information, and layer shape selection information to the bitstream. For example, the first program may cause the first image processing device to generate inter-layer relative information regarding the relative relationship between layers of 3D data composed of color component layer images and alpha component layer images, generate layer arrangement information regarding the arrangement of layers of the 3D data, generate layer shape selection information regarding the selection of the shape of layers of the 3D data, encode the 3D data to generate a bitstream, and add the inter-layer relative information, layer arrangement information, and layer shape selection information to the bitstream.

[0160] For example, the second image processing device may include a decoding unit that decodes a bit stream and generates 3D data composed of color component layer images and alpha component layer images, an inter-layer relative information analysis unit that analyzes inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bit stream, a layer arrangement information analysis unit that analyzes layer arrangement information extracted from the bit stream regarding the arrangement of layers of the 3D data, a layer shape selection information analysis unit that analyzes layer shape selection information extracted from the bit stream regarding the selection of the shape of layers of the 3D data, and a viewpoint image generation unit that renders the 3D data based on the inter-layer relative information, layer arrangement information, and layer shape selection information, and generates a viewpoint image seen from an arbitrary viewpoint. For example, the second image processing method may include the second image processing device decoding the bitstream to generate 3D data composed of color component layer images and alpha component layer images; the second image processing device analyzing inter-layer relative information extracted from the bitstream that relates to the relative relationship between layers of the 3D data; the second image processing device analyzing layer arrangement information extracted from the bitstream that relates to the arrangement of layers of the 3D data; the second image processing device analyzing layer shape selection information extracted from the bitstream that relates to the selection of the shape of layers of the 3D data; and the second image processing device rendering the 3D data based on the inter-layer relative information, layer arrangement information, and layer shape selection information to generate a viewpoint image seen from an arbitrary viewpoint.For example, the second program may cause the second image processing device to decode the bitstream, generate 3D data composed of color component layer images and alpha component layer images, analyze inter-layer relative information extracted from the bitstream regarding the relative relationship between layers of the 3D data, analyze layer arrangement information extracted from the bitstream regarding the arrangement of layers of the 3D data, analyze layer shape selection information extracted from the bitstream regarding the selection of the shape of layers of the 3D data, render the 3D data based on the inter-layer relative information, layer arrangement information, and layer shape selection information, and generate a viewpoint image seen from any viewpoint.

[0161] <4. First Embodiment> <Encoding Device> The present technology can be applied to any device. For example, the present technology can be applied to an encoding device that encodes 3D data composed of color component layer images and alpha component layer images and generates a bitstream. FIG. 23 is a block diagram showing an example of the configuration of an encoding device, which is one aspect of an image processing device to which the present technology is applied. The encoding device 500 (first image processing device) shown in FIG. 23 is a device that generates 3D data composed of color component layer images and alpha component layer images using multi-viewpoint images, encodes the 3D data, and generates a bitstream. Therefore, the encoding device 500 can also be called a 3D data generation device that generates 3D data, or a bitstream generation device that generates a bitstream.

[0162] Fig. 23 shows the main processing units, data flows, etc., but does not necessarily include all of them. In other words, in encoding device 500, there may be processing units that are not shown as blocks in Fig. 23, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 23.

[0163] As shown in FIG. 23 , the encoding device 500 (first image processing device) includes a layer image generation unit 511 , a layer image packing unit 512 , an SEI generation unit 513 , an encoding unit 514 , a storage unit 515 , and a transmission unit 516 .

[0164] The layer image generation unit 511 executes processing related to the generation of layer images that constitute 3D data. For example, the layer image generation unit 511 may acquire multi-viewpoint images supplied from outside the encoding device 500. Furthermore, the layer image generation unit 511 may acquire setting information related to layer generation that is input by, for example, a user or an application.

[0165] This setting information may include any information related to layer generation. For example, this setting information may include settings related to the resolution and size of layer images. This setting information may also include settings related to the number of layers to be generated.

[0166] Furthermore, when the above-described method 1 is applied, the setting information may include settings related to the relative relationship between layers, such as the relative number of alpha layers to color layers, or the relative resolution of the alpha layers to color layers.

[0167] Furthermore, when the above-described method 2 is applied, the setting information may include settings related to the arrangement of layers, such as settings related to the arrangement of additional layers and settings related to the arrangement of layers based on a non-uniform spacing function.

[0168] Furthermore, when the above-described method 3 is applied, the setting information may include settings related to the selection of the layer shape. For example, the setting information may include settings related to the designation of the curvature, whether the shape is cylindrical, whether the shape is non-planar, etc.

[0169] Based on this setting information, the layer image generation unit 511 may generate layer images that constitute 3D data using multi-viewpoint images. The layer image generation unit 511 may also generate metadata for the layer images (which may also be called metadata for the 3D data).

[0170] This metadata may contain any content. For example, this metadata may contain information indicating the resolution and size of the layer images that make up the 3D data generated by the layer image generation unit 511. This metadata may also contain information indicating the number of layers that make up the 3D data generated by the layer image generation unit 511.

[0171] Furthermore, when the above-described method 1 is applied, the metadata may include information regarding the relative relationship between layers of the 3D data generated in the layer image generation unit 511. For example, the metadata may include information indicating the relative number of alpha layers to color layers in the 3D data generated in the layer image generation unit 511, information indicating the relative resolution of the alpha layers to color layers, and the like.

[0172] Furthermore, when the above-described method 2 is applied, the metadata may include information regarding the arrangement of layers in the 3D data generated by the layer image generation unit 511. For example, the metadata may include information regarding the arrangement of additional layers in the 3D data generated by the layer image generation unit 511, information regarding the arrangement of layers based on a non-uniform arrangement function, and the like.

[0173] Furthermore, when the above-described method 3 is applied, information regarding the selection of the layer shape may be included in this metadata. For example, information indicating the curvature of the layer in the 3D data generated by the layer image generation unit 511, information indicating whether the layer has a cylindrical shape, information indicating whether the layer is non-planar, etc. may be included in this metadata.

[0174] Any method may be used to generate these layer images (3D data) and their metadata. For example, the layer image generation unit 511 may have a learning model (neural network) that receives setting information and multi-viewpoint images as input and outputs 3D data and its metadata composed of color component layer images and alpha component layer images, and may input the supplied setting information and multi-viewpoint images to the learning model, which may then output the layer images and their metadata. The layer image generation unit 511 may supply the group of color component and alpha component layer images (i.e., 3D data composed of color component and alpha component layer images) and their metadata generated in this manner to the layer image packing unit 512.

[0175] The layer image packing unit 512 executes processing related to packing of layer images. For example, the layer image packing unit 512 may acquire a group of layer images of color components and alpha components (i.e., 3D data configured from layer images of color components and alpha components) and their metadata supplied from the layer image generation unit 511. The layer image packing unit 512 may also acquire control information related to packing of layer images input by, for example, a user or an application.

[0176] This control information may include any information related to packing of layer images. For example, this control information may include information specifying a packing method (e.g., whether to stack in the spatial direction (vertical direction or horizontal direction) or the time direction).

[0177] The layer image packing unit 512 may pack the layer images based on the supplied control information, and generate a packed image which is the packed layer image group. The layer image packing unit 512 may supply the generated packed image to the encoding unit 514. Furthermore, the layer image packing unit 512 may supply the supplied metadata to the SEI generation unit 513.

[0178] The SEI generation unit 513 executes processing related to generation of an SEI. For example, the SEI generation unit 513 may acquire metadata supplied from the layer image packing unit 512. The SEI generation unit 513 may generate an SEI based on the metadata. In other words, the SEI generation unit 513 may generate an SEI that includes the contents of the metadata. Any information may be included in this SEI.

[0179] The SEI generation unit 513 may include, for example, an inter-layer relative information generation unit 521 , a layer placement information generation unit 522 , and a layer shape selection information generation unit 523 .

[0180] The inter-layer relative information generation unit 521 executes processing related to the generation of inter-layer relative information. For example, the inter-layer relative information generation unit 521 may generate inter-layer relative information for the 3D data based on information regarding the relative relationship between layers of the 3D data generated by the layer image generation unit 511, which information is included in the metadata supplied to the SEI generation unit 513. For example, the inter-layer relative information generation unit 521 may generate, as inter-layer relative information, information regarding the relative number of alpha layers with respect to color layers for the 3D data, information regarding the relative resolution of the alpha layer with respect to the color layers, or the like.

[0181] The layer placement information generation unit 522 executes processing related to the generation of layer placement information. For example, the layer placement information generation unit 522 may generate layer placement information for the 3D data based on information regarding the placement of layers of the 3D data generated by the layer image generation unit 511, which information is included in the metadata supplied to the SEI generation unit 513. For example, the layer placement information generation unit 522 may generate, as layer placement information, information regarding the placement of additional layers for the 3D data, information regarding the placement of layers based on a non-uniform placement function, or the like.

[0182] The layer shape selection information generator 523 executes processing related to the generation of layer shape selection information. For example, the layer shape selection information generator 523 may generate layer shape selection information for the 3D data generated by the layer image generator 511 based on information regarding selection of the shape of the layer of the 3D data, which is included in the metadata supplied to the SEI generator 513. For example, the layer shape selection information generator 523 may generate, as layer shape selection information for the 3D data, curvature specification information for a single direction or multiple directions of the layer, information indicating whether the layer is cylindrical, information indicating whether the layer is non-planar, and the like.

[0183] In this way, the SEI generation unit 513 may generate an SEI that includes inter-layer relative information, layer arrangement information, and layer shape selection information. Of course, the SEI may also include information other than these pieces of information. For example, the SEI may include information indicating the resolution and size of the layer images that constitute the 3D data generated by the layer image generation unit 511, information indicating the number of layers that constitute the 3D data, and the like.

[0184] The SEI generation unit 513 may supply the generated SEI to the encoding unit 514 .

[0185] In addition, if the above-mentioned method 1 is not applied in the SEI generation unit 513, the inter-layer relative information generation unit 521 may be omitted. Furthermore, if the above-mentioned method 2 is not applied, the layer placement information generation unit 522 may be omitted. Furthermore, if the above-mentioned method 3 is not applied, the layer shape selection information generation unit 523 may be omitted.

[0186] The encoding unit 514 performs processing related to encoding of 3D data. For example, the encoding unit 514 may acquire packed images supplied from the layer image packing unit 512. The encoding unit 514 may acquire SEI supplied from the SEI generation unit 513. The encoding unit 514 may encode the acquired packed images and generate a bitstream that is the encoded data. Any encoding method may be used. For example, the encoding unit 514 may encode the packed images using a predetermined 2D codec. The encoding unit 514 may also add the acquired SEI to the bitstream. The encoding unit 514 may supply the bitstream generated in this manner and to which the SEI is added to the storage unit 515. The encoding unit 514 may also supply the bitstream to the transmission unit 516.

[0187] The storage unit 515 includes a storage medium such as a hard disk or semiconductor memory, and performs processing related to the storage and reading of information using the storage medium. For example, the storage unit 515 may acquire a bitstream supplied from the encoding unit 514 and store it in the storage medium. The storage unit 515 may read the bitstream stored in the storage medium at a predetermined timing or based on a request from an external source (e.g., the transmission unit 516) and supply it to the transmission unit 516. The storage unit 515 may also include removable media and a drive therefor and store the acquired bitstream in the removable media attached to the drive. In this case, the removable media can be removed from the drive. In other words, the bitstream may be output to the outside of the encoding device 500 while stored on the removable media.

[0188] The transmitting unit 516 has a communication device that communicates with other devices and uses the communication device to perform processing related to the transmission of information, etc. For example, the transmitting unit 516 may acquire a bitstream supplied from the encoding unit 514. The transmitting unit 516 may also acquire a bitstream read from the storage unit 515. The transmitting unit 516 may transmit the acquired bitstream to other devices using the above-mentioned communication device.

[0189] In such an encoding device 500, the present technology described above in <3. Layer-Related Extension> may be applied. For example, Method 1 may be applied to the encoding device 500. In this case, the inter-layer relative information generation unit 521 may generate inter-layer relative information regarding the relative relationship between layers of 3D data configured from layer images of color components and layer images of alpha components. Furthermore, the encoding unit 514 may encode the 3D data to generate a bitstream and add the inter-layer relative information to the bitstream.

[0190] In this way, even if the relative relationship between layers is optimized depending on the content, etc., the decoder can easily grasp the relative relationship based on the inter-layer relative information. Therefore, the encoding device 500 can suppress a decrease in the quality of the rendered image.

[0191] Alternatively, the encoding device 500 may apply Method 2. In this case, the layer arrangement information generation unit 522 may generate layer arrangement information regarding the arrangement of layers of 3D data configured from layer images of color components and layer images of alpha components. Alternatively, the encoding unit 514 may encode the 3D data to generate a bitstream and add the layer arrangement information to the bitstream.

[0192] By doing so, the decoder can easily arrange each layer at an arbitrary position based on the layer arrangement information. Therefore, the encoding device 500 can arrange each layer at non-equidistant intervals while suppressing, for example, an unnecessary increase in data amount. This allows the encoding device 500 to optimize the layer arrangement according to, for example, content, etc. Therefore, the encoding device 500 can suppress a decrease in the quality of the rendered image while suppressing an unnecessary increase in data amount.

[0193] Alternatively, Method 3 may be applied to the encoding device 500. In this case, the layer shape selection information generation unit 523 may generate layer shape selection information regarding selection of the shape of a layer of 3D data configured from a layer image of a color component and a layer image of an alpha component. Alternatively, the encoding unit 514 may encode the 3D data to generate a bitstream and add the layer shape selection information to the bitstream.

[0194] In this way, the encoding device 500 can select the shape of the layer. For example, the encoding device 500 can select and apply the method that is most suitable for the content from among MPI, MCI, and MSI. This allows the encoding device 500 to suppress a reduction in the quality of the rendered image while suppressing an unnecessary increase in the processing load of encoding, decoding, playback, and the like.

[0195] <Flow of Encoding Process> An example of the flow of the encoding process executed by the encoding device 500 will be described with reference to the flowchart of FIG.

[0196] When the encoding process starts, in step S501, the layer image generation unit 511 generates layer images and their metadata using multi-view images in accordance with the input settings related to layer generation.

[0197] In step S502, the layer image packing unit 512 packs the layer images generated in step S501 in accordance with the input control information, and generates a packed image.

[0198] In step S503, the inter-layer relative information generation unit 521 generates inter-layer relative information as SEI using the metadata generated in step S501.

[0199] In step S504, the layer placement information generation unit 522 generates layer placement information as SEI using the metadata generated in step S501.

[0200] In step S505, the layer shape selection information generation unit 523 generates layer shape selection information as SEI using the metadata generated in step S501.

[0201] In step S506, the encoding unit 514 encodes the packed image generated in step S502 to generate a bitstream thereof. Furthermore, the encoding unit 514 adds the SEI generated in each process from step S503 to step S505 to the bitstream.

[0202] In step S507, the storage unit 515 stores the bitstream generated in step S506 and to which the SEI is added.

[0203] In step S508, the storage unit 515 reads out the bitstream stored in step S507.

[0204] In step S509, the transmission unit 516 transmits the bitstream generated in step S506 and to which the SEI has been added, or the bitstream read from the storage unit 515 in step S508.

[0205] When the process of step S509 is completed, the encoding process ends.

[0206] By performing each process in this manner, the encoding device 500 can suppress a decrease in the quality of the rendered image.

[0207] If the above-described method 1 is not applied, the process of step S503 may be skipped (omitted). If the above-described method 2 is not applied, the process of step S504 may be skipped (omitted). If the above-described method 3 is not applied, the process of step S505 may be skipped (omitted).

[0208] 5. Second Embodiment Decoding Device The present technology can be applied to a decoding device that decodes a bitstream obtained by encoding 3D data composed of color component layer images and alpha component layer images. Fig. 25 is a block diagram showing an example of the configuration of a decoding device, which is one aspect of an image processing device to which the present technology is applied. The decoding device 600 (second image processing device) shown in Fig. 25 is a device that decodes, for example, a bitstream generated by the encoding device 500 (Fig. 23), generates (restores) 3D data composed of color component layer images and alpha component layer images, and renders the 3D data at a desired viewpoint to generate and output a viewpoint image.

[0209] Fig. 25 shows the main processing units, data flows, etc., but does not necessarily include all of them. In other words, in the decoding device 600, there may be processing units that are not shown as blocks in Fig. 25, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 25.

[0210] As shown in FIG. 25, the decoding device 600 (second image processing device) has a receiving unit 611, a storage unit 612, a decoding unit 613, an SEI analysis unit 614, a layer image unpacking unit 615, a rendering unit 616, and an output unit 617.

[0211] The receiving unit 611 has a communication function for communicating with other devices and uses the communication function to perform processing related to the reception of information, etc. For example, the receiving unit 611 may obtain a bitstream transmitted from (the transmitting unit 516 of) the encoding device 500. The receiving unit 611 may supply the received bitstream to the storage unit 612. The receiving unit 611 may also supply the received bitstream to the decoding unit 613.

[0212] The storage unit 612 includes a storage medium such as a hard disk or semiconductor memory, and performs processing related to the storage and reading of information using the storage medium. For example, the storage unit 612 may acquire a bitstream supplied from the receiving unit 611 and store it in the storage medium. The storage unit 612 may read the bitstream stored in the storage medium at a predetermined timing or based on a request from an external device (e.g., the decoding unit 613, etc.) and supply it to the decoding unit 613. The storage unit 612 may also include removable media and a drive therefor, and may read the bitstream stored in the removable media attached to the drive at a predetermined timing or based on a request from an external device (e.g., the decoding unit 613, etc.) and supply it to the decoding unit 613. In other words, the bitstream may be supplied from the encoding device 500, etc. to the decoding device 600 while stored on removable media.

[0213] The decoding unit 613 executes processing related to decoding of a bitstream. For example, the decoding unit 613 may acquire a bitstream supplied from the receiving unit 611 or the storage unit 612. The decoding unit 613 may extract SEI added to the acquired bitstream and supply it to the SEI analysis unit 614. The decoding unit 613 may also decode the acquired bitstream and generate (restore) packed images of layers of 3D data configured from layer images of color components and layer images of alpha components. Note that any decoding method may be used. For example, the decoding unit 613 may decode the bitstream using a predetermined 2D codec. The decoding unit 613 may supply the packed images generated (restored) in this manner to the layer image unpacking unit 615.

[0214] The SEI analysis unit 614 executes processing related to analysis of the SEI. For example, the SEI analysis unit 614 may acquire the SEI supplied from the decoding unit 613. Any information may be included in this SEI. For example, this SEI may include inter-layer relative information, layer placement information, layer shape selection information, etc. Of course, this SEI may also include information other than these pieces of information. For example, this SEI may include information indicating the resolution and size of layer images constituting the 3D data, information indicating the number of layers constituting the 3D data, etc.

[0215] The SEI analysis unit 614 may analyze the SEI and generate metadata for the 3D data according to the analysis results. That is, the SEI analysis unit 614 may generate metadata that reflects the content of the acquired SEI.

[0216] The SEI analysis unit 614 may include, for example, an inter-layer relative information analysis unit 621 , a layer placement information analysis unit 622 , and a layer shape selection information analysis unit 623 .

[0217] The inter-layer relative information analysis unit 621 performs processing related to the analysis of the inter-layer relative information included in the SEI. For example, the inter-layer relative information analysis unit 621 may analyze the inter-layer relative information included in the SEI supplied from the decoding unit 613 and generate information regarding the relative relationship between layers of the 3D data as metadata of the 3D data according to the analysis results. For example, the inter-layer relative information analysis unit 621 may generate, as metadata of the 3D data, information indicating the relative number of alpha layers with respect to color layers in the 3D data, information indicating the relative resolution of the alpha layer with respect to the color layers, etc.

[0218] The layer placement information analysis unit 622 performs processing related to the analysis of layer placement information included in the SEI. For example, the layer placement information analysis unit 622 may analyze the layer placement information included in the SEI supplied from the decoding unit 613 and generate information about the placement of layers of the 3D data as metadata of the 3D data according to the analysis results. For example, the layer placement information analysis unit 622 may generate, as metadata of the 3D data, information about the placement of additional layers in the 3D data, information about the placement of layers based on a non-uniform placement function, and the like.

[0219] The layer shape selection information analysis unit 623 performs processing related to the analysis of the layer shape selection information included in the SEI. For example, the layer shape selection information analysis unit 623 may analyze the layer shape selection information included in the SEI supplied from the decoding unit 613 and generate information regarding the selection of the shape of the layer of the 3D data as metadata of the 3D data according to the analysis results. For example, the layer shape selection information analysis unit 623 may generate, as metadata of the 3D data, information indicating the curvature of the layer in the 3D data, information indicating whether the layer is cylindrical, information indicating whether the layer is non-planar, etc.

[0220] The SEI analysis unit 614 may supply the metadata generated as described above to the layer image unpacking unit 615 .

[0221] The layer image unpacking unit 615 performs processing related to unpacking of packed images. For example, the layer image unpacking unit 615 may acquire packed images supplied from the decoding unit 613. The layer image unpacking unit 615 may also acquire metadata supplied from the SEI analysis unit 614. The layer image unpacking unit 615 may unpack the acquired packed images in accordance with control information included in the acquired metadata, and generate unpacked layer image groups (color component layer image groups and alpha component layer image groups). The layer image unpacking unit 615 may arrange the layer image groups based on information included in the metadata, and generate 3D data configured of color component layer images and alpha component layer images. The layer image unpacking unit 615 may supply the generated 3D data (layer image groups) and the metadata to the rendering unit 616.

[0222] The rendering unit 616 performs processing related to the rendering of 3D data. For example, the rendering unit 616 may acquire 3D data (layer images) and their metadata supplied from the layer image unpacking unit 615. The rendering unit 616 may also acquire viewpoint position information input by, for example, a user or an application. This viewpoint position information is information indicating the position and orientation of the viewpoint when rendering the 3D data. The rendering unit 616 may render the 3D data based on the viewpoint position information and the metadata. In other words, the rendering unit 616 may grasp the content of the 3D data based on the metadata, perform rendering at the viewpoint indicated by the viewpoint position information, and generate the viewpoint image (rendered image). In other words, the rendering unit 616 can also be considered a viewpoint image generation unit that generates a viewpoint image. The rendering unit 616 may supply the generated viewpoint image to the output unit 617.

[0223] The output unit 617 executes processing related to the output of viewpoint images. For example, the output unit 617 may have a display device such as a monitor and use the display device to display the viewpoint images supplied from the rendering unit 616. The output unit 617 may have a communication device that communicates with other devices and use the communication device to transmit the viewpoint images supplied from the rendering unit 616 to an external device (e.g., another device) outside the decoding device 600. The output unit 617 may have removable media and a drive for the removable media and store the viewpoint images supplied from the rendering unit 616 in the removable media attached to the drive via the drive. The removable media is configured to be detachable from the drive.

[0224] In such a decoding device 600, the present technology described above in <3. Layer-Related Extension> may be applied. For example, Method 1 may be applied to the decoding device 600. In this case, the decoding unit 613 may decode the bitstream and generate 3D data configured of a layer image of a color component and a layer image of an alpha component. Then, the inter-layer relative information analysis unit 621 may analyze inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream. Then, the rendering unit 616 may render the 3D data based on the inter-layer relative information and generate a viewpoint image seen from an arbitrary viewpoint.

[0225] In this way, even if the encoder optimizes the relative relationship between layers according to the content, etc., the decoding device 600 can easily grasp the relative relationship based on the inter-layer relative information. Therefore, the decoding device 600 can suppress a decrease in the quality of the rendered image.

[0226] Alternatively, Method 2 may be applied to the decoding device 600. In this case, the decoding unit 613 may decode the bitstream and generate 3D data configured from layer images of color components and layer images of alpha components. Then, the inter-layer relative information analysis unit 621 may analyze layer arrangement information regarding the arrangement of layers of the 3D data extracted from the bitstream. Then, the rendering unit 616 may render the 3D data based on the layer arrangement information and generate a viewpoint image seen from an arbitrary viewpoint.

[0227] By doing so, the decoding device 600 can easily arrange each layer at an arbitrary position based on the layer arrangement information. Therefore, the decoding device 600 can arrange each layer at non-equidistant intervals while suppressing, for example, an unnecessary increase in the amount of data. This allows the decoding device 600 to optimize the layer arrangement according to, for example, content, etc. Therefore, the decoding device 600 can suppress a decrease in the quality of the rendering image while suppressing an unnecessary increase in the amount of data.

[0228] Alternatively, Method 3 may be applied to the decoding device 600. In this case, the decoding unit 613 may decode the bitstream and generate 3D data configured from a layer image of a color component and a layer image of an alpha component. Then, the inter-layer relative information analysis unit 621 may analyze layer shape selection information regarding selection of the shape of the layer of the 3D data extracted from the bitstream. Then, the rendering unit 616 may render the 3D data based on the layer shape selection information and generate a viewpoint image seen from an arbitrary viewpoint.

[0229] This allows the encoder to select the shape of the layer. For example, the encoder can select and apply the method that is most suitable for the content from among MPI, MCI, and MSI. The decoding device 600 can then easily grasp the selection result. This allows the decoding device 600 to suppress a reduction in the quality of the rendered image while suppressing an unnecessary increase in the processing load of encoding, decoding, playback, and the like.

[0230] <Flow of Decoding Process> An example of the flow of the decoding process executed by the decoding device 600 will be described with reference to the flowchart of FIG.

[0231] When the decoding process starts, the receiving unit 611 receives a bitstream in step S601.

[0232] In step S602, the storage unit 612 stores the bitstream.

[0233] In step S603, the storage unit 612 reads out the bitstream.

[0234] In step S604, the decoding unit 613 decodes the bitstream and generates a packed image.

[0235] In step S605, the inter-layer relative information analysis unit 621 of the SEI analysis unit 614 analyzes the inter-layer relative information included in the SEI extracted from the bitstream, and generates information regarding the relative relationship between layers of the 3D data as metadata of the 3D data according to the analysis results.

[0236] In step S606, the layer placement information analysis unit 622 of the SEI analysis unit 614 analyzes the layer placement information included in the SEI extracted from the bitstream, and generates information about the layer placement of the 3D data as metadata of the 3D data according to the analysis results.

[0237] In step S607, the layer shape selection information analysis unit 623 of the SEI analysis unit 614 analyzes the layer shape selection information included in the SEI extracted from the bit stream, and generates information regarding the selection of the shape of the layer of the 3D data as metadata of the 3D data according to the analysis results.

[0238] In step S608, the layer image unpacking unit 615 unpacks the packed image generated in step S604 to generate a group of color component layer images and a group of alpha component layer images. The layer image unpacking unit 615 arranges the group of layer images based on the information included in the metadata generated in each of the processes from step S605 to step S607, and generates 3D data composed of the color component layer images and the alpha component layer images.

[0239] In step S609, the rendering unit 616 renders the 3D data at the viewpoint indicated by the viewpoint position information based on the metadata generated in the processes of steps S605 to S607, and generates a viewpoint image.

[0240] In step S610, the output unit 617 outputs the viewpoint image.

[0241] When the process of step S610 is completed, the decoding process ends.

[0242] By performing each process in this manner, the decoding device 600 can suppress degradation in the quality of the rendered image.

[0243] If the above-described method 1 is not applied, the process of step S605 may be skipped (omitted). If the above-described method 2 is not applied, the process of step S606 may be skipped (omitted). If the above-described method 3 is not applied, the process of step S607 may be skipped (omitted).

[0244] <6. Supplementary Notes> <3D Data> In the above, MPI, MCI, MSI, and NeX using these have been described as examples of 3D data, but the present technology is not limited to these examples and can be applied to any 3D data that is composed of a layer image of a color component and a layer image of an alpha component.

[0245] <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs constituting the software are installed on a computer (also called an information processing device or image processing device). Here, the term "computer" includes computers built into dedicated hardware and general-purpose personal computers, etc., that can execute various functions by installing various programs.

[0246] FIG. 27 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0247] In a computer 900 shown in FIG. 27, a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, and a RAM (Random Access Memory) 903 are interconnected via a bus 904.

[0248] An input / output interface 910 is also connected to the bus 904. To the input / output interface 910, an input unit 911, an output unit 912, a storage unit 913, a communication unit 914, and a drive 915 are connected.

[0249] The input unit 911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, and an input terminal. The output unit 912 includes, for example, a display, a speaker, and an output terminal. The storage unit 913 includes, for example, a hard disk, a RAM disk, and a non-volatile memory. The communication unit 914 includes, for example, a network interface. The drive 915 drives removable media 921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0250] In a computer configured as described above, the CPU 901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 913 into the RAM 903 via the input / output interface 910 and the bus 904 and executing the program. The RAM 903 also stores data necessary for the CPU 901 to execute various processes as appropriate.

[0251] The program executed by the computer can be applied by recording it on, for example, a removable medium 921 such as a package medium. In this case, the program can be installed in the storage unit 913 via the input / output interface 910 by inserting the removable medium 921 into the drive 915.

[0252] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 914 and installed in the storage unit 913.

[0253] Alternatively, this program can be installed in advance in the ROM 902 or the storage unit 913 .

[0254] <Applicable Targets of the Present Technology> The present technology can be applied to any encoding / decoding method.

[0255] Furthermore, the present technology can be applied to any configuration, for example, various electronic devices.

[0256] Furthermore, for example, the present technology can also be implemented as part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set in which other functions are added to a unit (e.g., a video set).

[0257] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, an AV (Audio Visual) device, a portable information processing terminal, or an IoT (Internet of Things) device.

[0258] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0259] <Fields and uses to which this technology can be applied> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, for example, transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, nature monitoring, etc. In addition, the uses thereof are also arbitrary.

[0260] For example, the present technology can be applied to systems and devices used to provide viewing content, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for transportation, such as monitoring traffic conditions and controlling automatic driving. Furthermore, for example, the present technology can also be applied to systems and devices used for security. Furthermore, for example, the present technology can also be applied to systems and devices used for automatic control of machines, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, the present technology can also be applied to systems and devices used to monitor natural conditions, such as volcanoes, forests, and oceans, and wildlife. Furthermore, for example, the present technology can also be applied to systems and devices used for sports.

[0261] <Others> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be included not only in a bitstream, but also in a bitstream that includes differential information of the identification information relative to certain reference information. Therefore, in this specification, "flag" and "identification information" encompass not only the information itself, but also differential information relative to the reference information.

[0262] Furthermore, various types of information (e.g., metadata) related to captured images may be transmitted or recorded in any form as long as they are associated with the captured images. Here, the term "associate" means, for example, that one piece of data can be used (linked) when processing the other piece of data. That is, pieces of associated data may be combined into one piece of data or may be separate pieces of data. For example, information associated with encoded data (image) may be transmitted over a transmission path separate from that of the encoded data (image). Furthermore, for example, information associated with encoded data (image) may be recorded on a recording medium separate from that of the encoded data (image) (or on a different recording area of ​​the same recording medium). Note that this "association" may refer to only a portion of the data, rather than the entire data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.

[0263] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.

[0264] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0265] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0266] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and is able to obtain the necessary information.

[0267] Also, for example, each step of a single flowchart may be executed by a single device, or may be shared and executed by multiple devices. Furthermore, when a single step includes multiple processes, the multiple processes may be executed by a single device, or may be shared and executed by multiple devices. In other words, multiple processes included in a single step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.

[0268] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.

[0269] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.

[0270] Note that the present technology can also be configured as follows. (1) An image processing device comprising: an inter-layer relative information generation unit that generates inter-layer relative information regarding a relative relationship between layers of 3D data configured from color component layer images and alpha component layer images; and an encoding unit that encodes the 3D data to generate a bitstream and adds the inter-layer relative information to the bitstream. (2) The image processing device according to (1), in which the inter-layer relative information includes information regarding a relative number of the alpha component layers with respect to the color component layers. (3) The image processing device according to (1) or (2), in which the inter-layer relative information includes information regarding a relative resolution of the alpha component layers with respect to the color component layers. (4) An image processing method including: generating inter-layer relative information regarding a relative relationship between layers of 3D data configured from color component layer images and alpha component layer images; encoding the 3D data to generate a bitstream and adding the inter-layer relative information to the bitstream. (5) A program for causing a computer to execute a process including: generating inter-layer relative information regarding a relative relationship between layers of 3D data composed of a layer image of a color component and a layer image of an alpha component; encoding the 3D data to generate a bitstream; and adding the inter-layer relative information to the bitstream.

[0271] (11) An image processing device comprising: a layer placement information generation unit that generates layer placement information regarding the placement of layers of 3D data configured from layer images of color components and layer images of an alpha component; and an encoding unit that encodes the 3D data to generate a bitstream and adds the layer placement information to the bitstream. (12) The image processing device according to (11), wherein the layer placement information includes information regarding the placement of some layers of the 3D data. (13) The image processing device according to (12), wherein the information regarding the placement of some layers includes information indicating positions of some layers in a depth direction. (14) The image processing device according to (13), wherein the information regarding the placement of some layers includes information indicating the number of some layers. (15) The image processing device according to any of (11) to (14), wherein the layer placement information includes information regarding placement of layers of the 3D data based on a non-uniform placement function. (16) The image processing device according to (15), wherein the information regarding placement based on a non-uniform placement function includes information indicating whether the layers are placed based on the non-uniform placement function. (17) The image processing device according to (16), wherein the information regarding the arrangement based on the non-uniform arrangement function includes information specifying a mode of an arrangement method for the layers. (18) The image processing device according to (15), wherein the information regarding the arrangement based on the non-uniform arrangement function includes information indicating whether the layers are arranged using an inverse proportional pattern. (19) The image processing device according to (18), wherein the information regarding the arrangement based on the non-uniform arrangement function includes information indicating whether the layers are arranged using a non-linear pattern. (20) The image processing device according to (19), wherein the information regarding the arrangement based on the non-uniform arrangement function includes information indicating whether the spacing between the layers is not specified. (21) An image processing method comprising: generating layer arrangement information regarding an arrangement of layers of 3D data composed of color component layer images and alpha component layer images; encoding the 3D data to generate a bitstream; and adding the layer arrangement information to the bitstream.(22) A program for causing a computer to execute a process including: generating layer arrangement information regarding the arrangement of layers of 3D data composed of layer images of color components and layer images of alpha components; encoding the 3D data to generate a bitstream; and adding the layer arrangement information to the bitstream.

[0272] (31) An image processing device comprising: a layer shape selection information generation unit that generates layer shape selection information regarding selection of a shape of a layer of 3D data constituted by a layer image of a color component and a layer image of an alpha component; and an encoding unit that encodes the 3D data to generate a bit stream and adds the layer shape selection information to the bit stream. (32) The image processing device according to (31), wherein the layer shape selection information includes curvature designation information that designates a curvature of the layer in a single direction and information indicating whether the layer has a cylindrical shape. (33) The image processing device according to (31), wherein the layer shape selection information includes information indicating whether the layer is non-planar, curvature designation information that designates a curvature of the layer in a single direction, and information indicating whether the layer has a cylindrical shape. (34) The image processing device according to (31), wherein the layer shape selection information includes curvature designation information that designates curvature of the layer in multiple directions. (35) The image processing device according to (31), wherein the layer shape selection information includes information indicating whether the layer is non-planar and curvature designation information specifying curvatures of the layer in multiple directions. (36) The image processing device according to (31), wherein the layer shape selection information includes information indicating whether the layer is non-planar, information indicating whether the layer is cylindrical, and curvature designation information specifying curvatures of the layer in multiple directions. (37) An image processing method comprising: generating layer shape selection information regarding selection of a layer shape of 3D data configured from color component layer images and alpha component layer images; encoding the 3D data to generate a bitstream, and adding the layer shape selection information to the bitstream. (38) A program for causing a computer to execute processes including: generating layer shape selection information regarding selection of a layer shape of 3D data configured from color component layer images and alpha component layer images; encoding the 3D data to generate a bitstream, and adding the layer shape selection information to the bitstream.

[0273] (41) An image processing device comprising: a decoding unit that decodes a bitstream and generates 3D data configured of color component layer images and alpha component layer images; an inter-layer relative information analysis unit that analyzes inter-layer relative information regarding a relative relationship between layers of the 3D data extracted from the bitstream; and a viewpoint image generation unit that renders the 3D data based on the inter-layer relative information and generates a viewpoint image seen from an arbitrary viewpoint. (42) The image processing device according to (41), wherein the inter-layer relative information includes information regarding a relative number of the alpha component layers with respect to the color component layers. (43) The image processing device according to (41) or (42), wherein the inter-layer relative information includes information regarding a relative resolution of the alpha component layer with respect to the color component layers. (44) An image processing method comprising: decoding a bitstream to generate 3D data composed of color component layer images and alpha component layer images, analyzing inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream, rendering the 3D data based on the inter-layer relative information to generate a viewpoint image seen from an arbitrary viewpoint. (45) A program for causing a computer to execute processes including: decoding a bitstream to generate 3D data composed of color component layer images and alpha component layer images, analyzing inter-layer relative information regarding the relative relationship between layers of the 3D data extracted from the bitstream, and rendering the 3D data based on the inter-layer relative information to generate a viewpoint image seen from an arbitrary viewpoint.

[0274] (51) An image processing device comprising: a decoding unit that decodes a bitstream and generates 3D data configured of color component layer images and alpha component layer images; a layer arrangement information analysis unit that analyzes layer arrangement information regarding the arrangement of layers of the 3D data extracted from the bitstream; and a viewpoint image generation unit that renders the 3D data based on the layer arrangement information and generates a viewpoint image seen from an arbitrary viewpoint. (52) The image processing device according to (51), wherein the layer arrangement information includes information regarding the arrangement of some layers of the 3D data. (53) The image processing device according to (52), wherein the information regarding the arrangement of some layers includes information indicating positions of some layers in the depth direction. (54) The image processing device according to (53), wherein the information regarding the arrangement of some layers includes information indicating the number of some layers. (55) The image processing device according to any of (51) to (54), wherein the layer arrangement information includes information regarding the arrangement of layers of the 3D data based on a non-uniform arrangement function. (56) The image processing device according to (55), wherein the information regarding placement based on the non-uniform placement function includes information indicating whether the layers are placed based on the non-uniform placement function. (57) The image processing device according to (56), wherein the information regarding placement based on the non-uniform placement function includes information specifying a mode of a method for placing the layers. (58) The image processing device according to (55), wherein the information regarding placement based on the non-uniform placement function includes information indicating whether the layers are placed using an inverse proportional pattern. (59) The image processing device according to (58), wherein the information regarding placement based on the non-uniform placement function includes information indicating whether the layers are placed using a non-linear pattern. (60) The image processing device according to (59), wherein the information regarding placement based on the non-uniform placement function includes information indicating whether the spacing between the layers is not specified.(61) An image processing method comprising: decoding a bitstream to generate 3D data composed of color component layer images and alpha component layer images, analyzing layer arrangement information relating to the arrangement of layers of the 3D data extracted from the bitstream, rendering the 3D data based on the layer arrangement information to generate a viewpoint image seen from an arbitrary viewpoint. (62) A program for causing a computer to execute processes comprising: decoding a bitstream to generate 3D data composed of color component layer images and alpha component layer images, analyzing layer arrangement information relating to the arrangement of layers of the 3D data extracted from the bitstream, and rendering the 3D data based on the layer arrangement information to generate a viewpoint image seen from an arbitrary viewpoint.

[0275] (71) An image processing device comprising: a decoding unit that decodes a bitstream and generates 3D data composed of a layer image for a color component and a layer image for an alpha component; a layer shape selection information analysis unit that analyzes layer shape selection information regarding selection of a shape of a layer of the 3D data extracted from the bitstream; and a viewpoint image generation unit that renders the 3D data based on the layer shape selection information and generates a viewpoint image seen from an arbitrary viewpoint. (72) The image processing device according to (71), wherein the layer shape selection information includes curvature designation information that designates a curvature of the layer in a single direction and information indicating whether the layer has a cylindrical shape. (73) The image processing device according to (71), wherein the layer shape selection information includes information indicating whether the layer is non-planar, curvature designation information that designates a curvature of the layer in a single direction, and information indicating whether the layer has a cylindrical shape. (74) The image processing device according to (71), wherein the layer shape selection information includes curvature designation information that designates curvature of the layer in multiple directions. (75) The image processing device according to (71), wherein the layer shape selection information includes information indicating whether the layer is non-planar or not, and curvature designation information specifying curvature of the layer in multiple directions. (76) The image processing device according to (71), wherein the layer shape selection information includes information indicating whether the layer is non-planar or not, information indicating whether the layer is cylindrical or not, and curvature designation information specifying curvature of the layer in multiple directions. (77) An image processing method comprising: decoding a bitstream to generate 3D data configured of a color component layer image and an alpha component layer image; analyzing layer shape selection information related to selection of a layer shape of the 3D data extracted from the bitstream; and rendering the 3D data based on the layer shape selection information to generate a viewpoint image seen from an arbitrary viewpoint.(78) A program for causing a computer to execute a process including: decoding a bitstream to generate 3D data composed of a layer image of a color component and a layer image of an alpha component; analyzing layer shape selection information regarding selection of a shape of a layer of the 3D data extracted from the bitstream; and rendering the 3D data based on the layer shape selection information to generate a viewpoint image seen from an arbitrary viewpoint.

[0276] 500 Encoding device, 511 Layer image generation unit, 512 Layer image packing unit, 513 SEI generation unit, 514 Encoding unit, 515 Storage unit, 516 Transmission unit, 521 Inter-layer relative information generation unit, 522 Layer placement information generation unit, 523 Layer shape selection information generation unit, 600 Decoding device, 611 Receiving unit, 612 Storage unit, 613 Decoding unit, 614 SEI analysis unit, 615 Layer image unpacking unit, 616 Rendering unit, 617 Output unit, 621 Inter-layer relative information analysis unit, 622 Layer placement information analysis unit, 623 Layer shape selection information analysis unit, 900 Computer

Claims

1. An image processing device comprising: an inter-layer relative information generating unit that generates inter-layer relative information regarding a relative relationship between layers of 3D data constituted by a layer image of a color component and a layer image of an alpha component; and an encoding unit that encodes the 3D data to generate a bit stream and adds the inter-layer relative information to the bit stream.

2. The image processing device according to claim 1, wherein the inter-layer relative information includes information regarding the relative number of layers of the alpha component to the layers of the color components.

3. The image processing device according to claim 1, wherein the inter-layer relative information includes information regarding a relative resolution of the alpha component layer with respect to the color component layer.

4. The image processing device according to claim 1, further comprising a layer arrangement information generating unit configured to generate layer arrangement information regarding an arrangement of layers of the 3D data, wherein the encoding unit is configured to add the layer arrangement information to the bit stream.

5. The image processing device according to claim 4, wherein the layer arrangement information includes information regarding the arrangement of some layers of the 3D data.

6. The image processing device according to claim 5, wherein the information relating to the arrangement of the part of layers includes information indicating the positions of the part of layers in the depth direction.

7. The image processing device according to claim 6, wherein the information relating to the arrangement of the part of layers includes information indicating the number of the part of layers.

8. The image processing device according to claim 4, wherein the layer arrangement information includes information regarding an arrangement of layers of the 3D data based on a non-uniform arrangement function.

9. The image processing device according to claim 8, wherein the information relating to the arrangement based on the non-uniform arrangement function includes information indicating whether the layers are arranged based on the non-uniform arrangement function.

10. The image processing device according to claim 9, wherein the information relating to the arrangement based on the non-uniform arrangement function includes information specifying a mode of an arrangement method for the layers.

11. The image processing device of claim 8, wherein the information regarding the arrangement based on the non-uniform arrangement function includes information indicating whether the spacing between the layers is not specified, information indicating whether the layers are arranged according to a non-linear pattern, and information indicating whether the layers are arranged according to an inverse proportional pattern.

12. The image processing device according to claim 1, further comprising a layer shape selection information generation unit configured to generate layer shape selection information regarding selection of a shape of a layer of the 3D data, wherein the encoding unit is configured to add the layer shape selection information to the bit stream.

13. The image processing device according to claim 12, wherein the layer shape selection information includes curvature designation information that designates a curvature in a single direction of the layer, and information indicating whether the layer has a cylindrical shape.

14. The image processing device according to claim 12, wherein the layer shape selection information includes information indicating whether the layer is non-planar, curvature designation information specifying a curvature in a single direction of the layer, and information indicating whether the layer is cylindrical.

15. The image processing device according to claim 12, wherein the layer shape selection information includes curvature designation information that designates curvatures of the layer in multiple directions.

16. The image processing device according to claim 12, wherein the layer shape selection information includes information indicating whether the layer is non-planar or not, and curvature designation information that designates curvatures of the layer in multiple directions.

17. The image processing device according to claim 12, wherein the layer shape selection information includes information indicating whether the layer is non-planar, information indicating whether the layer is cylindrical, and curvature designation information that specifies the curvature of the layer in multiple directions.

18. An image processing method comprising: generating inter-layer relative information regarding a relative relationship between layers of 3D data constituted by a layer image of a color component and a layer image of an alpha component; encoding the 3D data to generate a bitstream; and adding the inter-layer relative information to the bitstream.

19. An image processing device comprising: a decoding unit that decodes a bit stream and generates 3D data composed of a color component layer image and an alpha component layer image; an inter-layer relative information analysis unit that analyzes inter-layer relative information regarding a relative relationship between layers of the 3D data extracted from the bit stream; and a viewpoint image generation unit that renders the 3D data based on the inter-layer relative information and generates a viewpoint image seen from an arbitrary viewpoint.

20. An image processing method comprising: decoding a bitstream to generate 3D data composed of a color component layer image and an alpha component layer image; analyzing inter-layer relative information regarding a relative relationship between layers of the 3D data extracted from the bitstream; and rendering the 3D data based on the inter-layer relative information to generate a viewpoint image seen from an arbitrary viewpoint.