Information processing device, information processing method, and program
The information processing apparatus allows selective display of annotation information in 3D objects based on viewpoint position, addressing the inconsistency in existing technologies by generating a 3D image file with metadata that controls annotation display.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-12
- Publication Date
- 2026-05-22
Smart Images

Figure 2026085046000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and a program.
Background Art
[0002] In recent years, technologies for generating three-dimensional data such as free-viewpoint videos and point cloud data from images of multiple imaging devices and data measured by LiDAR (Light Detection And Ranging) sensors have been known. Three-dimensional data such as volumetric media data is usually compression-encoded to reduce the data size. The MPEG (Moving Pictures Experts Group) is standardizing formats for volumetric media such as three-dimensional images and three-dimensional videos. Examples of encoding methods for three-dimensional data include G-PCC (Geometry-based Point Cloud Compression) for compressing point cloud data or V3C (Visual Volumetric Video-based Coding) for compressing volume media. Three-dimensional data compression-encoded using G-PCC or V3C can be stored in a file in a derived format of the Base Media File Format (ISOBMFF) of ISO / IEC 14496-12.
[0003] In recent years, annotations related to objects within image content have been generated by analyzing the image content. Annotations are, for example, annotation information that indicates the result of object recognition as a human- or computer-readable string, or as parameters for identifying and classifying the object. Annotation generation can sometimes be done by humans looking at images and making judgments, but it is often performed by AI-based image recognition processing. In this process, the objects recognized by the AI processing are shown as sub-regions within the image, and annotations are attached to or associated with these sub-regions. In such image recognition processing, it has become possible to identify 3D objects even in 3D volumetric data, and it is also possible to generate information indicating the 3D sub-regions that represent the identified objects, and information indicating the annotations to be attached to those sub-regions.
[0004] Patent Document 1 discloses a technique for adding annotation information to any 3D region in 3D data. Patent Document 2 also discloses a technique for adding annotation information (ancillary information) to any 3D region. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Japanese Patent Publication No. 2006-211531 [Patent Document 2] Japanese Patent Publication No. 2013-232730 [Overview of the project] [Problems that the invention aims to solve]
[0006] For example, when displaying a 3D object enclosed by a 3D region to which multiple annotations have been applied on a 2D display device, it is necessary to determine whether or not to display the annotations, or to selectively display the annotations, depending on the distance between the viewpoint position and the object position in 3D space, and how the 3D object is displayed within the current viewport (display area). However, Patent Document 1 does not describe a method for selectively switching which annotations to display when multiple annotations have been applied to an arbitrary 3D region. In other words, it is not possible to save selectable annotations in Patent Document 1.
[0007] Furthermore, Patent Document 2 does not describe a method for selectively switching and assigning annotation information to any 3D region in 3D space, when adding annotation information that indicates the interior of a 3D region. For example, for a 3D object represented in a 3D region, it is not possible to define separately as selectable information the annotation information when displayed from a wider space and the annotation information when the 3D region is displayed from inside that 3D region. In other words, even if the annotation information indicates the interior when viewed from a wider space, it will still be associated with the 3D media data as annotation information.
[0008] The present invention aims to provide a 3D image file in which annotation information to be displayed can be selected according to the viewpoint. [Means for solving the problem]
[0009] To achieve the object of the present invention, for example, an information processing apparatus according to one embodiment has the following configuration: an acquisition means for acquiring encoded data of a three-dimensional image, metadata corresponding to the encoded data of a three-dimensional image, which includes region information relating to a partial region in the three-dimensional image, annotation information associated with the partial region, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the three-dimensional space of the three-dimensional image, associated with the region information or the annotation information, and a generation means for generating a three-dimensional image file storing the encoded data of the three-dimensional image and the metadata. [Effects of the Invention]
[0010] This provides a 3D image file that allows users to select annotation information to display according to their viewpoint. [Brief explanation of the drawing]
[0011] [Figure 1] A block diagram showing an example of the hardware configuration of the information processing device according to Embodiment 1. [Figure 2] A diagram illustrating the 3D file according to Embodiment 1. [Figure 3] A diagram illustrating the three-dimensional region according to Embodiment 1. [Figure 4] A flowchart showing an example of the file creation process according to Embodiment 1. [Figure 5] A flowchart illustrating an example of the detailed process for setting viewpoint conditions. [Figure 6] A diagram showing an example of the file structure according to Embodiment 1. [Figure 7] A diagram showing an example of the item format for annotation information related to Embodiment 1. [Figure 8] A diagram showing an example of the format of a 3D domain set according to Embodiment 1. [Figure 9] A diagram showing an example of the format of viewpoint condition information according to Embodiment 1. [Figure 10] A diagram showing an example of the format of the dimensional coordinate range according to Embodiment 1. [Figure 11] A diagram for explaining the display of annotation information according to Embodiment 2. [Figure 12] A diagram for explaining the viewport according to Embodiment 2. [Figure 13] A diagram for explaining the viewport according to Embodiment 2. [Figure 14] A flowchart showing an example of file creation processing according to Embodiment 2. [Figure 15] A diagram showing an example of the structure of a file according to Embodiment 2. [Figure 16] A diagram showing an example of the description format of an item ID or item type. [Figure 17] A diagram showing an example of the description format of the storage location of an item. [Figure 18] A diagram showing an example of the description format of attribute information. [Figure 19] A diagram showing an example of the description of a three-dimensional region according to Embodiment 2. [Figure 20] A diagram showing an example of the description format of the position and size of a rectangular parallelepiped partial region. [Figure 21] A diagram showing an example of the format for describing the coordinates of a coordinate system of a reference point. [Figure 22] A diagram showing an example of the format for describing the rotation of a rectangular parallelepiped in quaternion. [Figure 23] A diagram showing an example of the description format of a virtual viewport. [Figure 24] A diagram showing an example of the description format of camera information of a virtual viewport. [Figure 25] A diagram showing an example of the description format of the range of the viewpoint position of a virtual viewport. [Figure 26] A diagram showing an example of the description format of the size of a virtual viewport. [Figure 27] A diagram for explaining the range of the viewpoint position of a virtual viewport. [Figure 28] A flowchart showing an example of file playback processing according to Embodiment 3. [Figure 29]A diagram illustrating an example of display according to viewpoint position according to Embodiment 3. [Figure 30] A diagram illustrating another example of display according to viewpoint position according to Embodiment 3. [Figure 31] A diagram illustrating another example of display according to viewpoint position according to Embodiment 3. [Figure 32] A diagram illustrating another example of display according to viewpoint position according to Embodiment 3. [Figure 33] A diagram showing an example of a description format for the position and size of a rectangular prism sub-region. [Figure 34] A diagram showing an example of the file structure according to Embodiment 4. [Figure 35] A figure showing an example of 3D image data according to Embodiment 4. [Figure 36] A schematic diagram showing an example of the metadata structure of a file according to Embodiment 4. [Figure 37] A flowchart showing an example of the file creation process according to Embodiment 4. [Figure 38] A flowchart showing an example of the file playback process according to Embodiment 5. [Modes for carrying out the invention]
[0012] The embodiments will be described in detail below with reference to the attached drawings. Note that the following embodiments do not limit the invention as defined in the claims. While the embodiments describe multiple features, not all of these features are essential to the invention, and the features may be combined in any way. Furthermore, in the attached drawings, identical or similar configurations are given the same reference numerals, and redundant descriptions are omitted.
[0013] [Embodiment 1] The information processing device according to Embodiment 1 acquires encoded data of a 3D image and metadata corresponding to the encoded data of the 3D image, and generates a 3D image file storing that data. Here, the metadata includes region information relating to a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight in the 3D image, which is associated with the region information or annotation information. Here, the format of the 3D image data is not particularly limited, but the following explanation assumes that it is data including information on the shape of an object in 3D space and texture information.
[0014] Figure 1 is a block diagram showing an example of the hardware configuration of the information processing device 100 according to this embodiment. Each functional unit of the information processing device 100 is connected via a system bus 109 to enable the transmission and reception of information. Although each functional unit of the information processing device 100 is implemented as hardware including a processor, the configuration is not limited to this as long as similar processing can be performed. For example, some or all of the functions described below may be implemented as software by implementing a program that performs the same functions as each functional unit. In this case, it is not necessary to separate the processing performed by each functional unit into hardware and software units according to the illustrated functional configuration.
[0015] The CPU 101 controls the operation of each functional unit of the information processing device 100. The ROM 102 is a non-volatile memory device. The RAM 103 is a volatile memory device capable of temporary data storage. The CPU 101 reads system programs, control programs for each functional unit, or application programs recorded in the ROM 102, loads them into the RAM 103, and executes them. The ROM 102 also stores parameter information and display data necessary for processing each function. The RAM 103 is also used as an input / output buffer to temporarily store data input or output in the processing of each functional unit. For example, the RAM 103 is used as a data buffer in the image file storage process described later, or as an output destination for temporary storage of image data or metadata to be stored in an image file.
[0016] The imaging unit 104 is, for example, an image sensor such as a CMOS sensor or a CCD. The imaging unit 104 converts the optical image formed on the imaging surface of the image sensor into an optical system (not shown) into an optical system (not shown). The imaging unit 104 also includes circuits for noise reduction and gain processing of the output signal of the image sensor, and an A / D conversion circuit for converting an analog signal into a digital signal, and outputs a digital image signal (image data).
[0017] The image processing unit 105 performs various image processing operations on the image data. The image processing operations according to this embodiment include, for example, development-related operations such as gamma conversion, color space conversion, white balance, or exposure correction. The image processing unit 105 may also be capable of performing image data analysis operations or synthesis operations that combine two or more image data. Furthermore, the image processing unit 105 includes an encoding / decoding unit 111, a metadata processing unit 112, a generation unit 113, and a recognition processing unit 114. In this embodiment, to facilitate understanding of the invention, each image processing operation is assumed to be performed by a single hardware unit, the image processing unit 105; however, some or all of these image processing operations may be configured to be performed by separate hardware units.
[0018] The encoding / decoding unit 111 is a video or still image codec that conforms to H.265 (HEVC), H.264 (AVC), H.266 (VVC), AV1, JPEG, G-PCC, or V3C, etc. The encoding / decoding unit 111 performs encoding and decoding of 3D point cloud images, captured still images, or video data handled by the information processing device 100.
[0019] The metadata processing unit 112 acquires the data encoded by the encoding / decoding unit 111 (encoded data). Next, the metadata processing unit 112 generates an image file in a file format compliant with ISOBMFF. Specifically, the metadata processing unit 112 performs analysis processing of the encoded data stored in the image file, which includes a 3D point cloud image, a still image, or an image sequence, and acquires parameter information related to the encoded data. Then, the metadata processing unit 112 performs metadata creation processing to be stored in the file along with the encoded data. Note that the metadata processing unit 112 can generate metadata not only for ISOBMFF-compliant files, but also for other video file formats or JPEG files, for example. Note that the encoded data acquired here may be data previously stored in the ROM 102 or non-volatile memory 110, or data acquired via the communication unit 108 and stored in the buffer of RAM 103.
[0020] Furthermore, the metadata processing unit 112 processes the metadata stored in the file. In particular, the metadata processing unit 112 performs processing to acquire and analyze metadata such as parameters necessary for decoding image data, or parameters necessary for displaying or playing it back.
[0021] The generation unit 113 generates region information indicating a sub-region in which an object detected by the recognition processing unit 114 can be identified. Here, since a 3D image is used as the image to be processed, a 3D sub-region in 3D space is treated as a sub-region. When generating information data for sub-regions, not only the detection results by the recognition processing unit 114 may be used, but input from the user operating the information processing device 100 via the operation input unit 107 may also be used. For example, when setting the region of a sub-region, the region recognized by the recognition processing unit 114 may be used, or the region specified by user input may be used. In the following explanation, when simply referred to as "sub-region," it refers to a 3D sub-region in a 3D image.
[0022] The metadata processing unit 112 acquires (generates) metadata that includes region information relating to a subregion in a 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight in the 3D image, which is associated with the region information or annotation information. As will be described in detail later, the metadata processing unit 112 generates information indicating the coordinates for identifying the 3D subregion and the shape of the 3D subregion from the 3D subregion information generated by the generation unit 113 as region information. Furthermore, the metadata processing unit 112 performs analysis processing of this metadata when the 3D image data is being played back.
[0023] The recognition processing unit 114 processes the image data acquired as storage target and performs object detection and recognition processing (for example, using a machine learning model). Note that the object recognition processing described as being performed by the recognition processing unit 114 may be executed by another device, such as an image recognition server, and the results of this processing may be obtained. The recognition processing unit 114 acquires various information, such as the position or range of the detected object within the image. Since a three-dimensional image is used as the image to be processed, information about a sub-region in the three-dimensional coordinate space is acquired as a sub-region representing an object. Note that the image recognition processing according to this embodiment includes the detection and recognition of multiple objects, as well as the classification of objects.
[0024] The display unit 106 is a display device, such as a liquid crystal display (LCD), that is integrated with the information processing device 100 or detachably attached to the information processing device 100. The display unit 106 is used for displaying the GUI for operating the information processing device 100, displaying a live view during shooting, or displaying the screen for playback of generated image files.
[0025] The operation input unit 107 is one of the various user interfaces provided by the information processing device 100, such as operation buttons, switches, a mouse, or a keyboard. Alternatively, the display unit 106 and the operation input unit 107 may be implemented as an integrated unit, such as a touch panel and a touch panel sensor. When the operation input unit 107 detects that an operation input has been made to the user interface, it outputs a control signal to the CPU 101 indicating that an operation input has been made.
[0026] The communication unit 108 is a communication interface for the information processing device 100 to external devices. The communication unit 108 may be, for example, a network interface that connects to a network and transmits and receives transmission frames. In this case, the communication unit 108 may be, for example, a PHY and MAC (transmission media control processing) capable of wired LAN connection via Ethernet®. Furthermore, if the communication unit 108 is capable of connecting to a wireless LAN, it may include a controller that performs wireless LAN control such as IEEE802.11a / b / g / n / ac / ax, an RF circuit, or an antenna.
[0027] The non-volatile memory 110 is a non-volatile recording device with a large storage capacity, such as an SD card, CompactFlash®, or flash memory. The non-volatile memory 110 according to this embodiment stores generated image files or image files acquired via the communication unit 108.
[0028] The information processing device 100 according to this embodiment generates a 3D image file that stores 3D image data. In the following, such 3D image data may be simply referred to as "3D data".
[0029] As described above, the information processing device 100 according to this embodiment acquires metadata including region information relating to a sub-region in a 3D image, annotation information associated with the sub-region, and conditional information according to the viewpoint position or line of sight direction indicating whether or not to display the annotation information. Here, the sub-region is the region corresponding to an object in the 3D image data. Furthermore, the annotation information is assumed to be text information displayed in association with the object, but it is not particularly limited to this as long as it is information associated with the sub-region.
[0030] In this case, conditional information is set, which includes the range of viewpoint positions or the direction of the line of sight for displaying certain annotation information. Such conditional information is associated with the annotation information or with a sub-region representing an object. The various types of information set for such 3D data will be explained below with reference to Figures 2 to 10.
[0031] Figure 2 is a diagram illustrating the three-dimensional image data according to this embodiment. In Figure 2, an example is shown of an object 201 (signboard), an object 202 (tree), and a reference point 203 that are recognized within the three-dimensional data 200. In the example in Figure 2, annotation information displayed when viewed from viewpoint 204 and annotation information displayed when viewed from viewpoint 205 are separately assigned (associated) to object 201.
[0032] Here, viewpoints 204 and 205 indicate the location (range) of the viewpoint that is the condition for displaying the corresponding annotation information. Here, viewpoints 204 and 205 are illustrated as points, but this is just an example; the condition information may also be specified as a region or using coordinate conditions. Such conditions that specify the viewpoint location as a condition for displaying annotation information (regarding the display of annotation information) will be referred to as "viewpoint conditions (information)" below. Note that viewpoints 204 and 205 may be set within or outside the region of the 3D data.
[0033] Figure 3 shows a 3D region 300, which is a subregion representing object 201 in 3D data 200, and 3D regions 301 and 302 (corresponding to viewpoints 204 and 205, respectively), which represent the regions corresponding to viewpoint conditions when a region is set as a viewpoint condition. In the example of Figure 3, each region is set as a rectangular parallelepiped region, but it is not limited to this shape as long as it can be set as a region; for example, each region may be set as a spherical region or a planar region.
[0034] In the example shown in Figure 3, the user specifies the regions of 3D region 301 and 3D region 302 as viewpoint condition information for setting different annotation information associated with the 3D region 300 of object 201. One method for specifying the region is to use the syntax "3D RegionSet" defined in MPEG, but this is not particularly limited as long as the region can be specified in a similar manner. Here, the user sets the annotation information for when the viewpoint position is included in 3D region 301 and the annotation information for when the viewpoint position is included in 3D region 302. The CPU 101 in this embodiment stores this information about the subregions in the 3D image in the 3D image file, including it in the metadata.
[0035] Furthermore, 3D region 303 is a broad 3D region that encompasses all of 3D regions 300 to 302. Details on how to set conditions so that different annotation information is displayed depending on whether the viewpoint position is inside or outside the 3D region 303 will be described later. In addition, default annotation information may be set to be displayed when the viewpoint position is not included in any of 3D regions 301, 3D region 302, or 3D region 303.
[0036] The CPU 101 according to this embodiment stores annotation information associated with such sub-regions in the 3D image file, further including it in the metadata. Furthermore, the CPU 101 stores the above-mentioned viewpoint condition information in the 3D image file, including it in the metadata, in association with the region information or annotation information. The following describes such processing by the CPU 101. As mentioned above, there are viewpoint conditions that depend on the viewpoint position and conditions that depend on the line of sight, but in the following description, we will use the viewpoint conditions that depend on the viewpoint position. An example of using the line of sight will be described in Embodiment 2.
[0037] Figure 4 is a flowchart illustrating an example of a process that associates default annotation information with annotation information corresponding to the viewpoint position for objects recognized in 3D data. The process shown in Figure 4 is executed by the CPU 101, for example, in response to an input from the user to initiate an action.
[0038] Figure 4 is a flowchart showing an example of the 3D image file creation process of the information processing device 100. The process corresponding to this flowchart is realized by the CPU 101 reading the program stored in ROM 102, expanding it into RAM 103, and executing it, thereby operating each block. The process shown in Figure 4 is executed by the CPU 101 in response to, for example, an operation to start the operation input by the user. The explanation of each box in the file, as shown in Figure 6, will be described later.
[0039] In S401, the CPU 101 controls the imaging unit 104 or the image processing unit 105 to acquire 3D image data to be stored in a file.
[0040] In S402, the recognition processing unit 114 performs object detection processing within the 3D image. Here, predetermined objects, such as a person or a specific object, are detected. The recognition processing by the recognition processing unit 114 may, for example, perform matching processing based on reference image data of a specific object that has been registered in advance to detect whether the object is present in the image. In addition to object detection processing, processing including the generation of object attribute information may also be performed. Furthermore, in addition to detection processing from the image, the process of identifying objects may also involve specifying the object's region by user input using a 3D image editing application or the like.
[0041] In S403, the generation unit 113 generates region information relating to the subregions in the 3D image that indicate objects detected by the recognition processing unit 114. For example, the region information may be the structure of the 3D RegionSet shown in Figure 8, or the 3D RegionSet shown in Figure 19. The generated region information is stored in the output buffer to be stored in the 'mdat' box shown in Figure 6.
[0042] In S404, CPU 101 determines whether or not to add (associate) annotation information to the sub-region generated in S403. For example, if attribute information for an object was generated in S402, it determines whether to add that attribute information as annotation information. Here, the conditions for determining whether or not to add annotation information to the sub-region (or corresponding object) being processed can be set arbitrarily. For example, it may be determined whether the detected object is pre-configured to have a predetermined (for example, according to the type of object) annotation information added. Alternatively, CPU 101 may obtain the user's selection operation regarding whether or not to add annotation information to the detected object, and determine whether or not to associate the annotation information according to that operation. If it is determined that annotation information should be added, the process proceeds to S405; otherwise, the process proceeds to S409.
[0043] In S405, the metadata processing unit 112 sets annotation information to be associated with the sub-region generated in S403. Here, the metadata processing unit 112 may, for example, set annotation information to be displayed for each viewpoint condition described later, set default annotation information to be displayed by default, or create entry data for the 'ipco' box described later as attribute information. Here, the metadata processing unit 112 may accept user input to set the content of the annotation information (e.g., text) and set the annotation information based on said user input.
[0044] In S406, the metadata processing unit 112 sets conditions (viewpoint conditions) that indicate the conditions for displaying the annotation information set in S405. Details of S405 in this embodiment will be described later with reference to Figure 5.
[0045] In S407, the metadata processing unit 112 determines whether to add other viewpoint conditions to the annotation information set in S405. If it does, the process returns to S406; otherwise, the process proceeds to S408. Here, for example, if user input indicating the addition of viewpoint conditions has been obtained, it may be determined that other viewpoint conditions should be added.
[0046] In S408, the metadata processing unit 112 determines whether to associate additional annotation information with the sub-region generated in S403. If additional annotation information is to be associated, the process returns to S405, and the processes from S405 to S407 are repeated. This loop process makes it possible to associate multiple pieces of annotation information with a single sub-region. It is also possible to create information for multiple annotation information display conditions (and virtual viewports used in Embodiment 2) for each piece of annotation information and store it in a file. If no additional annotation information is to be associated in S408, the process proceeds to S409. The process from S402 to S408 completes the process of associating annotation information with a single sub-region.
[0047] In S409, the metadata processing unit 112 determines whether to complete the object detection process for the 3D image. If it is completed, the process proceeds to S410; otherwise, the process returns to S402.
[0048] In S410, the encoding / decoding unit 111 performs encoding processing on the 3D image data and saves the encoded data to the output buffer. The metadata processing unit 112 then integrates the metadata created up to S406 and the metadata necessary for decoding the encoded data, creates data with a 'meta' box structure, and saves it to the output buffer.
[0049] In S411, the metadata processing unit 112 combines the information in the 'ftyp' box related to the 3D image file, the information in the 'meta' box which stores the final metadata, and the information in the 'mdat' box which stores items such as encoded data and viewpoint condition information. The CPU 101 then writes the image file generated by storing the combined metadata and image data from RAM 103 to non-volatile memory 110, saves it, and terminates the process shown in Figure 4.
[0050] In this embodiment, the 3D image data to be stored in the 3D image file was described as being obtained by controlling the imaging unit 104 and the image processing unit 105. However, the invention is not limited to this example, as long as the data can be processed in a similar manner. For example, the 3D image data may be an image pre-stored in the ROM 102 or non-volatile memory 110, or an image received via the communication unit 108.
[0051] Next, with reference to Figure 5, the process of setting viewpoint condition information will be explained. Figure 5 is a flowchart that shows in detail an example of the process performed in S406 according to this embodiment.
[0052] In S501, CPU101 determines whether to set viewpoint condition information by specifying a region or by specifying coordinates. For example, CPU101 may present the user with a display allowing them to choose between specifying a region or specifying coordinates, and make the determination in S501 based on the user input. If viewpoint condition information is set by specifying a region, the process proceeds to S502; if viewpoint condition information is set by specifying coordinates, the process proceeds to S506.
[0053] In S502, the metadata processing unit 112 sets the shape of the region to be used as a viewpoint condition (annotation information will be displayed when the viewpoint is in that region). Hereinafter, the region used as such a viewpoint condition may be simply referred to as the "condition region". Here, the shape of the region can be, for example, a point, line, plane, cuboid, sphere, or ellipsoid, but other shapes may also be used. For example, the metadata processing unit 112 may set the shape of the condition region to be used by default (for example, as a cuboid), and when it receives input from the user to change the shape of the condition region, it may reset the shape of the condition region based on that input.
[0054] In S503, the metadata processing unit 112 sets the position of the reference point of the condition region. The reference point is represented by three-dimensional coordinate information. The reference point here is not particularly limited as long as it is a coordinate that allows the position of the condition region to be set according to predetermined rules. For example, if the condition region is a rectangular prism, the reference point may be set as the coordinate of a predetermined vertex of the rectangular prism, or as the coordinate of the center of the rectangular prism, and if the condition region is a sphere or ellipsoid, the reference point may be set as the coordinate of their centers. The position of the reference point may be set, for example, as a predetermined position in the coordinate system within the three-dimensional image data, or it may be set based on user input.
[0055] In S504, the metadata processing unit 112 sets the size of the condition area. The size of the condition area may be, for example, an initial size provided according to the shape of the condition area, or it may be set based on user input. The size of the condition area can be expressed as offset information from a reference point if the condition area is a rectangular prism, as a radius if the condition area is a sphere, or as the radii of the X, Y, and Z axes if it is an ellipsoid.
[0056] In S505, the metadata processing unit 112 sets the amount of rotation of the condition region from the reference orientation. The amount of rotation of the condition region can be set separately for X-axis rotation, Y-axis rotation, and Z-axis rotation. In this case, the amount of rotation is represented by a quaternion, but it may also be represented using different parameters, such as Euler angles. After S505, the process proceeds to S509.
[0057] In S506-S508, the metadata processing unit 112 sets coordinates to be used as viewpoint conditions (annotation information will be displayed when the viewpoint is located at those coordinates). Hereinafter, such coordinates used as viewpoint conditions may simply be referred to as "condition coordinates". Here, ranges are set for the X, Y, and Z coordinates as condition coordinates, and annotation information will be displayed when the viewpoint's position satisfies all of these ranges. Note that the upper limit, lower limit, or both of the possible coordinates may be set as condition coordinates. Furthermore, it is not necessary to set all of the X, Y, and Z coordinates as condition coordinates; it is sufficient if at least one of them is set. Here, the metadata processing unit 112 sets the X coordinate in S506, the Y coordinate in S507, and the Z coordinate in S508 as condition coordinates. The metadata processing unit 112 can set condition coordinates based on user input, for example. After S508, the process proceeds to S509.
[0058] In S509, the metadata processing unit 112 sets the priority between viewpoint conditions. Here, the priority between viewpoint conditions is information used to determine whether or not to display annotation information corresponding to any of the viewpoint conditions when multiple viewpoint conditions are met. Here, a different numerical value (minimum is 0) is assigned as the priority for each viewpoint condition, and the annotation information corresponding to the viewpoint condition with the smallest priority value is displayed.
[0059] Here, we will explain assuming that only one annotation is displayed according to priority, but multiple annotations may be displayed. For example, the metadata processing unit 112 may select a predetermined number of viewpoint conditions in order of priority and display all annotations corresponding to the selected viewpoint conditions.
[0060] Furthermore, although we have described the sub-regions and conditional regions as being set based on user input, this is not necessarily limited to them, as long as these regions can be set as sub-regions in three-dimensional image data. For example, the region of an object recognized by a general object recognition process may be obtained as a sub-region.
[0061] Figure 6 shows an example of a file structure generated when annotation information, as explained in Figures 4 and 5, is set for object 201 using the 3D data and viewpoint condition information shown in Figures 2 and 3. File 600 represents the entire 3D image file. The FileTypeBox at the beginning of the file stores the brand name for the reader to identify the file specifications. MetaBox601('meta') is a box that summarizes information about the 3D data and stores multiple hierarchical boxes. HandlerBox602('hdlr') stored at the beginning of MetaBox601 stores a declaration of the handler type for analyzing the structure of MetaBox601. In this embodiment, HandlerBox602 stores the handler type 'volv', which identifies that it is metadata for 3D data. PrimaryItemBox603('pitm') specifies the identifier of the representative item in file 600. In this embodiment, PrimaryItemBox603 stores item ID=1 of the 3D data 200. ItemInfoBox604('iinf') stores information such as the item ID or item type for all items contained in file 600. Item ID=1 corresponds to the 3D data 200 shown in Figure 2, and the item type stored is 'gpe1', which indicates volume media.
[0062] Item ID=2 corresponds to the 3D region 300 shown in Figure 3, and the item type stored is 'vran', indicating 3D region annotation information (VolumetricRegionItem). The item format for 3D region annotation information is shown in Figure 7. 3D region annotation information can store flags used to specify the region's location and multiple 3D region sets (3D RegionSet). The format for a 3D region set is shown in Figure 8. A 3D region set can store multiple 3D regions with different shapes. The number of stored 3D regions is indicated by region_count. The shape of the 3D region is indicated by geometry_type. For example, geometry_type=0 represents a point, 1 represents a line, 2 represents a plane, and 3 represents a cuboid. The division stored in the 3D region set differs depending on the shape of the 3D region; for example, if the 3D region is a cuboid, the reference point, size, and rotation information of the cuboid are stored.
[0063] In the example in Figure 6, when geometry_type=5, the shape of the 3D region is specified by attribute information. Here, the region information optionally specifies a cuboid and the value of the attribute that defines the region (region_identifier_value). When a cuboid is specified, the 3D region is indicated by a bounding box.
[0064] Item ID=3 is the 3D region 301 shown in Figure 3, and the item type stored is 'vvra', which indicates the view point condition information (VolumetricViewPointRegionItem) according to this embodiment. An example of the definition format for view point condition information is shown in Figure 9. In view point condition information, there are cases where the region of the view point is specified and cases where the coordinate conditions of the view point are specified. range_type901 is an item where the lower 4 bits are valid values, and if all of these bits are 0, it indicates that the region of the view point is specified. When specifying the region of the view point, the 3D region set (3D RegionSet) explained in Figure 8 is stored thereafter. When specifying the coordinate conditions of the view point, at least one bit of the 0th to 2nd bits (x, y, z) of range_type901 must be 1. The values of these 0th to 2nd bits indicate whether the range of the X coordinate, Y coordinate, and Z coordinate are specified, respectively. For example, if the value of the 0th to 2nd bits is 5 (0b101), it indicates that the range of the X coordinate and the range of the Z coordinate will be stored thereafter. The third bit (f) indicates whether the condition used to set the coordinate range is an AND condition (f=1) or an OR condition (f=0). If the condition used to set the coordinate range is an AND condition, the viewpoint condition is determined to be true if all the conditions of the specified coordinate range for the viewpoint are true. If the condition used to set the coordinate range is an OR condition, the viewpoint condition is determined to be true if at least one of the conditions of the specified coordinate range for the viewpoint is true.
[0065] priority902 stores the priority set in S509 in Figure 5. When viewpoint condition information is specified using the coordinate conditions of the viewpoint position, the dimensional coordinate range (DimensionRange) (903~905) is stored subsequently as the range information of the dimensional coordinates specified in range_type901. An example of the definition format for the dimensional coordinate range is shown in Figure 10. Item 1001 (presicion_bytes_minus1) determines how many bits are used to represent the lower and upper limits of the coordinate range specification, which will be described later. The lower and upper limits of the coordinate range specification may be determined by selecting, for example, 8 bits, 16 bits, 24 bits, or 32 bits.
[0066] The limit_type1002 item has its lower two bits as a valid value, and at least one bit from the 0th to the 1st bit (L, U) must be 1. These 0th to 1st bits indicate whether a lower limit and upper limit are specified for the dimensional coordinates, respectively. For example, if the value of the 0th to 1st bits is 2 (0b10), it indicates that the upper limit will be stored later, and if the value of the 0th to 1st bits is 3 (0b11), it indicates that both the lower limit and the upper limit will be stored later.
[0067] Item ID=4 corresponds to the 3D region 302 shown in Figure 3, and its item type is 'vvra', which indicates viewpoint condition information, similar to item ID=3. Item ID=5 corresponds to the 3D region 303 shown in Figure 3, and its item type is 'vvra', which indicates viewpoint condition information, similar to items ID=3 and 4.
[0068] ItemLocationBox605('iloc') stores information indicating the storage location of each item in file 600, including 3D data. The information in ItemLocationBox605 allows the location within file 600 of 3D data, 3D region information, and viewpoint condition information, which will be stored in MediaDataBox610 later.
[0069] ItemReferenceBox606('iref') stores information describing the relationships between items contained in file 600. Relationships between items are established by specifying the item reference type, which allows for the identification of the type of item reference. Furthermore, the reference relationships between each item are described by writing the item ID specified in ItemInfoBox604 in from_item_ID and to_item_ID. In this embodiment, item ID=2 (3D area 300) is associated with item ID=1 (3D data 200). Also, item ID=3 (3D area 301) and item ID=4 (3D area 302) are associated with item ID=5 (3D area 303) and item ID=2 (3D area 300).
[0070] ItemPropertiesBox607('iprp') stores various attribute information (item properties) about the items contained in file 600. ItemPropertiesBox607 further includes ItemPropertyContainerBox608('ipco') which describes the attribute information, and ItemPropertyAssociation609('ipma') which shows the association between the attribute information and each item.
[0071] In this embodiment, the properties assigned to item ID=1 (3D data 200) are 'gpcC', which indicates the settings of the 3D data, and 'gpsr', which indicates the size of the 3D data. 'udes' stands for UserDescriptionProperty and is attribute information that can store arbitrary text information. In this embodiment, annotation information for the 3D region is set using 'udes'.
[0072] In this embodiment, 'Object' is stored as the default annotation information assigned to item ID=2 (3D region 300). Additionally, 'Tokyo' is stored as the annotation information assigned to item ID=3 (3D region 301). 'Kanagawa' is stored as the annotation information assigned to item ID=4 (3D region 302). 'Nameboard' is stored as the annotation information assigned to item ID=5 (3D region 303).
[0073] In MediaDataBox610('mdat'), each item data is stored at the location specified in ItemLocationBox605. Volumetric Media611 is the encoded data of 3D data 200. Region Annotation612 is the 3D region annotation information of 3D region 300.
[0074] Viewpoint Region Annotation613 is the viewpoint condition information for the 3D region 301, and is specified as the region of the viewpoint position (when range_type=0). Here, the viewpoint condition for the 3D region 301 is set as having the highest priority (priority=0). Here, the region shape is set as a cuboid (geometry_type=3), and its size and rotation information are set.
[0075] Viewpoint Region Annotation614 sets viewpoint condition information in the same format as Viewpoint Region Annotation613. Viewpoint Region Annotation615 is viewpoint condition information for the 3D region 303 and is specified as the coordinate condition for the viewpoint position (when range_type=15(0b1111)). Here, the viewpoint condition for the 3D region 303 sets lower and upper limits for the range of the X, Y, and Z coordinates, respectively.
[0076] With this configuration, it is possible to obtain encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes region information about a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating viewpoint conditions, and to generate a 3D image file that stores them. In particular, it is possible to generate a 3D image file in which the display of annotation information differs depending on the viewpoint position for any 3D region of the 3D data.
[0077] [Embodiment 2] In Embodiment 1, an example was described in which annotation information is displayed when the viewpoint position satisfies predetermined viewpoint conditions. In Embodiment 2, an example will be described in which annotation information is displayed when the line of sight direction satisfies predetermined conditions as a viewpoint condition. The information processing device 100 according to this embodiment has basically the same configuration as that of Embodiment 1 and can perform the same processing, so redundant explanations will be omitted.
[0078] The display of annotation information according to the difference in viewing direction in a 3D image by the information processing device 100 according to this embodiment will be explained with reference to Figure 11. Image 1101 in Figure 11 shows a 3D image of a traffic light installed at a road intersection in Japan. The 3D region 1102 represents a 3D sub-region surrounding the three-lamp traffic light and information sign included in the traffic light in image 1101. Annotation information 1103 is annotation information associated with the 3D region 1102. Annotation information 1103 indicates that three pieces of annotation information, "Scramble Crossing," "Shinjuku 3-chome West," and "Shinjuku 3 W.", are associated with the 3D region 1102.
[0079] Viewports 1111, 1113, 1115, and 1117 in Figure 2 are viewports that display a 3D image of a traffic light. These viewports consist of four patterns, including examples in which annotation information is displayed in association with the display of a sub-region of the 3D region 1102, and examples in which it is not displayed. In this embodiment, the viewports represent the projection range of the 3D image onto the display plane as seen from a specific viewpoint in 3D space.
[0080] Viewport 1111 is a viewport where the viewpoint is close to the traffic light in the 3D coordinate space of the traffic light image. In viewport 1111, the 3D region 1102 is displayed within the viewport, and two annotation information items 1112, "Shinjuku 3-chome West" and "Shinjuku 3 W.", are displayed as callouts. Viewport 1113 is a viewport from a viewpoint slightly further away from the traffic light than in the example of viewport 1111 in the same 3D space. In viewport 1113, the 3D region 1102 is displayed, and only the annotation information item 1114 for "Shinjuku 3-chome West" is displayed. Not only the viewpoint position but also the line of sight direction (relative direction from the traffic light) differs between viewport 1111 and viewport 1113. Viewport 1115 is a viewport from a viewpoint located further away from the traffic light than in the example of viewport 1113. In viewport 1115, the 3D region 1102 is displayed smaller than in the example of viewport 1111. In viewport 1115, annotation information 1116 for "Scramble Crossing" is displayed. In viewport 1117, neither the 3D area 1102 nor the annotation information is displayed within the viewport.
[0081] Thus, there are cases where whether an object represented by a three-dimensional region is actually contained within the viewport being displayed changes depending on the viewing position as well as the viewing direction. Taking into account such changes in display content depending on the viewing direction, the information processing device 100 according to this embodiment can display annotations if the object is actually contained within the image (viewport) being displayed, and not display annotation information (or display different annotation information) if it is not contained within the viewport being displayed, based on the viewing direction.
[0082] Although not shown in Figure 11, if the 3D region 1102 is not displayed in the viewport based on the line of sight, an arrow shape indicating the direction in which the 3D region 1102 exists in 3D space may be displayed in the viewport as annotation information.
[0083] As described above, the information processing device 100 according to this embodiment uses a condition that specifies the line of sight direction as a viewpoint condition, rather than a condition that specifies the viewpoint position (or a condition that specifies the viewpoint position in addition to a condition that specifies the viewpoint position). In this embodiment, the 3D image file stores 3D area information, annotation information, virtual viewport information (AnViewport), and viewpoint condition information data. In addition, metadata is stored as an association between the virtual viewport information and the annotation information.
[0084] The viewpoint conditions according to this embodiment can be treated in the same way as the viewpoint conditions according to Embodiment 1, except that they specify the line of sight direction rather than the viewpoint position. An example using the viewpoint conditions according to this embodiment will be described below with reference to Figure 12.
[0085] In Figure 12, a 3D region 1201 exists as a sub-region representing an object in 3D data, viewpoint 1202 is shown as the viewpoint position, and arrow line 1203 is shown as the direction of the line of sight from viewpoint 1202. Viewport 1204 represents the display screen when the 3D image is played back when the viewpoint is set to viewpoint 1202. Reference point 1205 is the reference point when setting the 3D region 1201, and in this case it is the center point of the rectangular prism 3D region 1201. Virtual viewport 1206 is a region that is virtually generated in the direction of the object (reference point 1205) from viewpoint 1202, and in this case it is set as the region assumed to be the projected image of the object onto viewport 1204. Dashed line 1207 is a line indicating the direction from viewpoint 1202 to reference point 1205. Here, assuming that the object in the 3D region 1201 will be displayed in viewport 1204 when the dashed line 1207 passes through viewport 1204, annotation information associated with the 3D region is displayed in viewport 1204.
[0086] In this embodiment, the viewpoint position and viewpoint direction can be changed by user operation. When playing back a 3D image, the content projected onto the display screen may also change depending on the viewpoint position in 3D space, or the direction or rotation of the line of sight, which changes the position and tilt of the viewport. Here, when the line of sight direction is changed, both viewport 1204 and virtual viewport 1206 change. Such an example will be explained with reference to Figure 13.
[0087] In Figure 13, similar to Figure 12, a 3D region 1301 exists as a sub-region representing an object in the 3D data, and a viewpoint 1302 is shown as the viewpoint position, with an arrow line 1303 indicating the line of sight from viewpoint 1302. Also in Figure 13, viewport 1304, reference point 1305, and virtual viewport 1306 are illustrated, similar to viewport 1204, reference point 1205, and virtual viewport 1206 in Figure 12. In Figure 12, the dashed line 1207 passes through viewport 1204, whereas in Figure 13, the dashed line extending from viewpoint 1302 to reference point 1305 does not pass through viewport 1304. Therefore, the CPU 101 can choose not to display annotation information in such cases.
[0088] In the examples shown in Figures 12 and 13, annotation information is displayed when a straight line from the viewpoint to an object (reference point of a partial region) passes through a viewport that is set as a predetermined range according to the viewing direction. Thus, the information processing device 100 according to this embodiment may set viewpoint conditions to determine whether or not to display annotation information according to the viewing direction. Here, the information processing device 100 can set viewpoint conditions so that annotation information is displayed when the relationship between the viewing direction and the direction from the viewpoint to the 3D region (object) satisfies a predetermined relationship. This predetermined relationship may be, for example, when the dashed line 1207 shown in Figure 3 passes through a viewport 1204 that is a predetermined range centered on the viewing direction, or when the angle between the arrow line 1203, which is the viewing direction, and the dashed line 1207 is within a predetermined angle (for example, 30°), and can be set arbitrarily.
[0089] Figure 14 is a flowchart showing an example of the 3D image file creation process of the information processing device 100 according to this embodiment. The process shown in Figure 14 is the same as the process shown in Figure 4 of Embodiment 1, except that S1401 to S1403 are performed instead of S406, so redundant explanations are omitted.
[0090] In S1401, the metadata processing unit 112 creates virtual viewport information corresponding to the annotation information set in S405. Here, the metadata processing unit 112 creates information about the reference point of a sub-region as data for item attribute information to be stored in the 'ipco' box 1510 in the example shown in Figure 15, which will be described later. The generation unit 113 also creates virtual viewport information as data for the AnViewport data structure in Figure 15.
[0091] In S1402, the generation unit 113 creates an item of viewpoint condition information that specifies the direction of the line of sight. Here, the generation unit 113 creates viewpoint condition information as data of the structure of VolumeticRegionConditionforAnnotation in Figure 13. The viewpoint condition information according to this embodiment includes the viewport information for selecting annotation information generated in S1401. The generation unit 113 also sets flag values indicating the conditions for displaying annotation information set in S405 to condition_flags(1302), shown in Figure 23, which will be described later. The generation unit 113 then saves the created annotation information display condition information data to the output buffer for storage in the 'mdat' box 1503.
[0092] In S1403, the metadata processing unit 112 creates data that associates the annotation information set in S405 with the attribute information of the viewpoint condition information item created in S1402, and proceeds to S407. Specifically, the metadata processing unit 112 creates data that associates the entry in the 'ipco' box 1510 containing the annotation information set in S405 with the item ID of the viewpoint condition information created in S1402. This attribute information association data is the entry data to be stored in the 'ipma' box 1511.
[0093] Figure 15 shows an example of a file structure generated when the processing shown in Figure 14 is performed using the 3D data and viewpoint condition information shown in Figures 12 and 13. Here, the 3D image data shown in Figure 12 is stored in a file, and it includes one 3D image, one sub-region, and three annotation information entries associated with that sub-region. The file structure shown in Figure 15 contains some content in common with the file structure shown in Figure 5 of Embodiment 1, so some of the redundant explanations are omitted.
[0094] File 1500 in Figure 15 shows the entire 3D image file. MetaBox 1502 ('meta') is a box that summarizes information about the 3D data and has basically the same functionality as MetaBox 601. MetaBox 1501 in Figure 15 includes HandlerBox 1504, PrimaryItemBox 1505, ItemInfoBox 1506, ItemLocationBox 1507, ItemReferenceBox 1508, and ItemPropertiesBox 1509. FileTypeBox 1501 ('ftyp') stores the brand name for the reader to identify the specification of the image file. In the example in Figure 15, the 'ftyp' box contains 'gpci' as the brand name and 'mif1' as the compatible brand name.
[0095] ItemInfoBox1506('iinf') defines the item ID or item type for each item in the image file. Description 1600 in Figure 16 shows the structure of the 'iinf' box. Description 1600 includes entry_count1601, which indicates the number of elements in the file, and the ItemInfoEntry data array 1602. In the example in Figure 15, the file contains 5 items, so entry_count1601 is 5. Also, description 1610 in Figure 16 shows the structure of ItemInfoEntry. Description 1610 includes item_ID1611, which indicates the item ID, item_type1612, which indicates the item type, and item_name1613, which indicates the item name parameter, as elements of the data array. Here, the item type for the G-PCC 3D point cloud image is 'gpe1', the item type for the 3D region information is 'vran', and the item type for the viewpoint condition information is 'vrca'.
[0096] ItemLocationBox1507('iloc') contains information indicating the storage location of data for each item, such as an image, within the file. An example of the structure of the 'iloc' box according to this embodiment is shown in Figure 17. 1701 to 1705 in Figure 17 are information that identifies the file location where the item data exists. item_count1706 indicates the number of items. Description 1700 indicates that the item data corresponding to item_ID1701 is represented by a byte offset from the beginning of the file, according to the information shown in description 1702. Description 1703 indicates that there is one item data, description 1704 indicates the offset position, and description 1704 indicates the data length. Since file 1500 stores five item data in 'mdat', the value of item_count1706 in 'iloc' is 5, and it contains five sets of parameters from 1701 to 1705.
[0097] ItemPropertiesBox1509('iprp') contains items from ItemPropertyContainerBox1510('ipco') and ItemPropertyAssociationBox1511('ipma'). The 'ipco' box contains a list of various attribute information (property) data. The 'ipma' box contains information that associates the attribute information with items. The 'ipma' box may be written in such a way that a single attribute information written in the 'ipco' box can be associated with multiple items.
[0098] The 'ipco' box contains, for example, information indicating the width and height of an image item in pixels, or data for a set of parameters necessary for decoding the encoded data of a 3D image. In this embodiment, the 'ipco' box can store the reference point of a subregion as attribute information of the 3D region. In the file format shown in Figure 15, the 'ipco' box stores an entry ('vrrp') of the coordinate data of the reference point of the subregion, and the 'ipma' box stores information about the association of the 3D region information with the item ID.
[0099] Figure 18 shows an example of the structure of VolumetricRegionRepresentationPointProperty('vrrp'), which is attribute information for the reference point of a subregion. rep_pos1801 is Vector3 data describing the Cartesian coordinate system coordinates of the reference point of the subregion. The Vector3 data structure is shown in Figure 21.
[0100] Furthermore, in the example in Figure 15, three annotation information items ('udes') are described as attribute information associated with each viewpoint condition information item. Here, the UserDescriptionProperty('udes') defined in ISO / IEC 23008-12 is used to describe the annotation information.
[0101] MediaDataBox1503('mdat') contains encoded data 1512 for the 3D image, 3D region information 1513, and viewpoint condition information 1514, 1515, and 1516. As described above, the information that identifies the file location where the item data in 'iloc' resides associates the image item with the file storage location of each item data in 'mdat'.
[0102] The encoded data 1512, Region Annotation, is the encoded data for a G-PCC 3D point cloud image item.
[0103] The 3D region information 1513 has the structure of a 3D RegionSet as shown in Figure 19 and stores data for 3D region information items. In the 3D region information shown in Figure 19, the shape of one or more subregions can be defined. region_count1901 indicates the number of subregions defined in the data, and in the example in Figure 19, the number of regions is 1. geometry_type1902 is a numerical value that represents the shape type of the subregion. In the example in Figure 19, the shape of the subregion is a rectangular prism, and the value of geometry_type1902 is "3", which corresponds to a rectangular prism.
[0104] Furthermore, description 1903 describes the detailed shape of the cuboid of the sub-region. cuboid1904 is CuboidRegion data indicating the position and size of the sub-region of the cuboid. Figure 20 shows an example of the structure of CuboidRegion data. In Figure 20, anchor2001 is Vector3 data describing the position coordinates of the cuboid. Also, size_x, size_y, and size_z shown in description 2002 are the lengths of the three sides of the cuboid of the sub-region along the x, y, and z axes of the Cartesian coordinate system. The CuboidRegion data structure shown in Figure 20 also includes a flag (anchor_included) indicating whether an anchor is included, a flag (scale_included) indicating whether a scale is included, and precision. In the example in Figure 20, since the anchor (the point specifying the region) is included, the anchor_include flag is set to 1, and the position coordinates of the cuboid constituting the region are specified. These position coordinates only need to be configured to indicate the position of the rectangular prism, and may, for example, one of the predetermined vertices (e.g., the upper left back in the left-hand model, or the lower right front in the right-hand model) or the center point may be used.
[0105] Furthermore, rotation1905 in Figure 19 is QuaternionRotation data used to represent the rotation of a three-dimensional subregion of a rectangular prism using quaternion representation. An example of the structure of QuaternionRotation data is shown in Figure 22. As shown in Figure 22, the x, y, and z components of the quaternion representation are described by real values. The other w component of the quaternion can be calculated using the values of the x, y, and z components by a predetermined calculation.
[0106] The descriptions shown in Figure 15, items 1514-1516, represent the data for viewpoint condition information items. Each of these data items is stored in the structure of VolumeticRegionConditionforAnnotation shown in Figure 23. The data structure shown in Figure 23 contains data for one of the virtual viewport information items (AnViewport) from 2303 to 2035, depending on the flag value of anviewport_flags2301. Here, the flag value of anviewport_flags2301 is one of 1, 2, or 3. The data structure shown in Figure 23 also includes the flag value of condition_flags2302. The bit flags indicated by condition_flags2302 control the display of annotation information associated with viewpoint condition information on the viewport during image file playback.
[0107] Next, Figures 24 to 26 show an example of the data structure of virtual viewport information (AnViewport). Figure 24 shows an example of the structure of AnViewport. AnViewport data includes both or either extCamInfo2401, which is the range data of external camera information, and intCamInfo2402, which is the information of internal camera information.
[0108] The range data of the external camera information indicates the range of the viewpoint position and the range of the line of sight direction when playing back a 3D image file. Here, the range data of the external camera information is stored in the AnExtCamRange structure shown in Figure 25. In Figure 25, the range of the viewpoint position is described by the coordinates shown in description 2501 and the ranges of the x, y, and z axes from that coordinate position in 2502. On the other hand, the line of sight direction is described by the real values of the x, y, and z components of the quaternion representation shown in description 2503 and the real value ranges of each component in description 2504 based on the direction determined in description 2503. The range of the other w component of the quaternion can be calculated by a predetermined calculation using the values of the x, y, and z components.
[0109] The range data of the internal camera information included in the AnViewport data structure is the range of the virtual viewport's size, and such information is stored in the AnIntCamRange structure shown in Figure 26. In Figure 26, the vertical / horizontal aspect ratio 2601, the minimum horizontal length 2602, and the maximum horizontal length 2603 of the virtual viewport are described. For example, if the aspect ratio is 0.75 and the horizontal length is from 512 to 1024, the vertical length will be from 384 to 768.
[0110] Furthermore, the `condition_flags2302` in the `VolumeticRegionConditionforAnnotation` data structure shown in Figure 23 contains a flag indicating whether or not to display annotation information when the 3D region is included in the virtual viewport. `condition_flags2302` also includes a flag indicating whether or not the 3D region is included in both the virtual viewport and the viewport itself, which is a condition for displaying the annotation condition. Additionally, if multiple annotations are associated with the same subregion, `condition_flags2302` may include a flag indicating a display condition where, if multiple reference points are associated with the 3D region, the virtual viewport contains reference points for all subregions, and the annotation information is displayed based on a determination using that viewpoint condition. Alternatively, if multiple reference points are associated with the 3D region, `condition_flags2302` may include a flag indicating a display condition where, if the virtual viewport from the viewpoint position contains the most reference points for the subregions, the annotation information is displayed, and the annotation condition is displayed based on a determination using that viewpoint condition.
[0111] Furthermore, Figure 27 provides supplementary information regarding the viewpoint position range in the virtual viewport information (AnViewport). In Figure 27, viewpoint 2701 represents the viewpoint in three-dimensional space, and reference point 2702 represents the reference point of a sub-region. The viewpoint position range is the range indicated by the relative positional relationship from the coordinates of reference point 2702. In Figure 27, this positional relationship is represented by the distance between two points, which is the length of the straight line connecting the viewpoint and the reference point in three-dimensional space, indicated by arrow 2703 (annotation information is displayed when the distance between these two points satisfies a predetermined condition (for example, being below a predetermined threshold)). Alternatively, this positional relationship may be represented by the distance between two points obtained by projecting each point of the viewpoint and the reference point onto the xy-plane (the plane where the z coordinate is zero), indicated by arrow line 2704. If we use the distance between these two points (for example, to display annotation information when the distance between these two points satisfies a predetermined condition), we can set the z-coordinate value of 2501 in the AnExtCamRange structure of Figure 25 to 0 and the z_range value of 2502 to 0.
[0112] With this configuration, it is possible to generate a 3D image file that stores metadata including the viewing direction as a viewpoint condition. Therefore, for any 3D region of the 3D data, it is possible to generate a 3D image file in which the display of annotation information differs depending on the viewing direction.
[0113] [Embodiment 3] Embodiment 3 describes the process of playing back a 3D image file generated by the information processing device 100 according to Embodiment 1 or 2. In this embodiment, the device for playing back the 3D image file is described as the information processing device 100, but an external device different from the information processing device 100 may be used as the playback device.
[0114] Here, the information processing device 100 generates an image with annotation information superimposed on it, corresponding to the viewpoint position (or viewing direction) set by the user during playback of a 3D image file, and outputs it to the display device. In the following description, we will assume that playback of a file containing the 3D data 200 described in Embodiment 1 is performed, but the same processing can be performed even when playing back a 3D image file generated by the information processing device 100 according to Embodiment 2.
[0115] Figure 28 is a flowchart showing an example of the playback process performed by the information processing device 100 according to this embodiment. The process shown in Figure 28 can be realized by the CPU 101 reading the corresponding processing program stored in the ROM 102, expanding it into the RAM 103, and executing it, thereby operating each block. The playback process shown in Figure 28 will be described as starting, for example, when the information processing device 100 is set to playback mode and an operation input related to a playback instruction for the image file is detected.
[0116] In S2801, the information processing device 100 acquires the viewpoint position for playback. Here, the information processing device 100 can acquire the viewpoint position based on user input. Alternatively, for example, the information processing device 100 may set the viewpoint position based on information indicating where the user is in a pre-identified space.
[0117] In S2802, the information processing device 100 generates an image from the viewpoint position using the encoded data (Volumetric Media 611) of the 3D data 200 contained in the file. Since the method for generating an image from 3D data can be performed using general 3D image playback processing, a detailed explanation is omitted here.
[0118] Loop1 is a loop process that scans all 3D region annotation information contained in the file, including S2803 to S2807. In file 600, item ID=2, defined as 'vran', corresponds to the 3D region annotation information.
[0119] In S2803, the information processing device 100 sets default annotation information for the current annotation information. The default annotation information here is the annotation information directly attached as a property to the 3D domain annotation information. In file 600, the default annotation information is 'Object' defined as 'udes' with property_index=3. Also, the lowest priority (for example, 255 out of 0 to 255) is set as the current priority. Loop2 is a nested loop process within Loop1, including S2804 to S2806, and is a loop that scans all viewpoint condition information associated with the 3D domain annotation information currently being processed. In this case, all viewpoint condition information in file 600 corresponds to item ID=3, item ID=4, and item ID=5 defined in 'vvra' associated with item ID=2 ('vran').
[0120] In S2804, the information processing device 100 determines whether the viewpoint position satisfies the viewpoint condition currently being processed. The viewpoint condition used is the one described in Embodiment 1. If the viewpoint condition is not satisfied, the process returns to the beginning of Loop 2, and S2804 starts again with the next viewpoint condition information as the target of processing. If the viewpoint condition is satisfied, the process proceeds to S2805.
[0121] In S2805, the information processing device 100 compares the current priority with the priority of the viewpoint condition currently being processed. If the current priority is greater than the priority of the viewpoint condition currently being processed, the process moves to S2806; otherwise, the process returns to the beginning of Loop2, and S2804 starts again with the next viewpoint condition information as the target of processing.
[0122] In S2806, the information processing device 100 sets the annotation information of the viewpoint condition information currently being processed as the current annotation information. The information processing device 100 also sets the priority of the viewpoint condition information currently being processed as the current priority. Once the scanning of all viewpoint condition information associated with the 3D region annotation information currently being processed in Loop2 is complete, the process proceeds to S2807.
[0123] In S2807, the information processing device 100 superimposes the current annotation information onto the viewpoint position image generated in S2802. Once Loop1 has completed scanning all 3D region annotation information contained in the file, the process proceeds to S2808. In S2808, the information processing device 100 outputs the viewpoint position image to the display device, and the process shown in Figure 28 is completed.
[0124] Next, referring to Figures 29 to 32, we will explain examples of viewpoint positions specified by the user and the resulting viewpoint position images.
[0125] Figure 29 shows an example where the user sets the viewpoint position 2902 inside the 3D region 301 on the viewpoint position manipulation screen 2901. The file playback screen 2903 outputs an image in which the viewpoint position image of viewpoint position 2902 is superimposed with the object 2904 (corresponding to item ID=2, 'vran') included in the viewpoint position image, and the annotation information 2905 of object 2904. Since the viewpoint position 2902 is set inside the 3D region 301, the viewpoint conditions item ID=3 ('Tokyo') and item ID=5 ('Nameboard') are true (satisfied). Of these, item ID=3 has a higher priority, so the annotation information ('Tokyo') of item ID=3 is superimposed on the image and output.
[0126] Figure 30 shows an example where the user sets the viewpoint position 3002 inside the 3D region 302 in the viewpoint position manipulation screen 3001. The file playback screen 3003 outputs an image in which the viewpoint position image of viewpoint position 3002 is superimposed with the object 3004 (corresponding to item ID=2, 'vran') included in the viewpoint position image, and the annotation information 3005 of object 3004. Since the viewpoint position 3002 is set inside the 3D region 302, the viewpoint conditions Item ID=4 ('Kanagawa') and Item ID=5 ('Nameboard') are true. Of these, Item ID=4 has a higher priority, so the annotation information ('Kanagawa') of Item ID=4 is superimposed on the image and output.
[0127] Figure 31 shows an example where the user sets the viewpoint position 3102 inside the 3D region 303 on the viewpoint position manipulation screen 3101. It is assumed that viewpoint position 3102 is not included in either the 3D region 301 or the 3D region 302. The file playback screen 3103 outputs an image where the viewpoint position image of viewpoint position 3102 is superimposed with object 3104 (corresponding to item ID=2, 'vran') included in the viewpoint position image, and annotation information 3105 for object 3104. Here, since viewpoint position 3102 is set inside the 3D region 303 and is not included in either the 3D region 301 or the 3D region 302, only the viewpoint condition item ID=5 ('Nameboard') is true. Therefore, the annotation information ('Nameboard') for item ID=5 is superimposed on the image and output.
[0128] Figure 32 shows an example where the user sets the viewpoint position 3202 outside the 3D region 303 in the viewpoint position manipulation screen 3201. The file playback screen 3203 outputs an image in which the viewpoint position image of viewpoint position 3202 is superimposed with the object 3204 (corresponding to item ID=2, 'vran') included in the viewpoint position image, and the annotation information 3205 of object 3204. Since the viewpoint position 3202 is set outside the 3D region 303, there are no items that satisfy the viewpoint condition. Therefore, the default annotation information ('Object') for item ID=2 is superimposed on the image and output.
[0129] This configuration makes it possible to play back a 3D image file that stores encoded 3D image data and metadata corresponding to the encoded 3D image data, including region information, annotation information, and condition information. In particular, it becomes possible to set appropriate annotation information to be displayed according to the viewpoint position and output the annotation information superimposed on the viewpoint position image.
[0130] [Embodiment 4] The information processing devices according to Embodiments 1 and 2 acquire metadata including region information, annotation information, and condition information, and generate a 3D image file containing 3D image data and metadata. On the other hand, the information processing device according to Embodiment 4 acquires encoded data of a 3D image and metadata corresponding to the encoded data of the 3D image, including first region information relating to a first sub-region in the 3D image, second region information relating to a second sub-region included in the first sub-region, first annotation information associated with the first sub-region, and second annotation information associated with the second sub-region. The information processing device then generates a 3D image file storing the acquired encoded data and metadata of the 3D image.
[0131] The information processing device 100 according to this embodiment has the same configuration as shown in Figure 1 and can perform the same processing, so redundant explanations will be omitted. In addition, the configuration of the 3D image file generated by the information processing device 100 according to this embodiment is basically the same as that shown in Figure 6 of Embodiment 1, but the differences between them will be explained below.
[0132] The item format for the 3D region annotation information according to this embodiment is the same as that shown in Figure 7. The positional information within the 3D image according to this embodiment is described by a defined Vector3, as shown in Figure 33. The data structure of Vector3 includes x, y, and z coordinate information, and the bit size of the parameters for each x, y, and z coordinate information is determined by precision information.
[0133] The CuboidRegion data, which shows the position and size of a subregion when its shape is a rectangular prism, is the same as that shown in Figure 20. Similarly, the QuaternionRotation data, which shows the rotation of a 3D subregion of a rectangular prism using quaternion representation, is the same as that shown in Figure 22.
[0134] The information processing device 100 according to this embodiment describes, as metadata, information indicating a first annotation information associated with a first sub-region in a 3D image, and a second annotation information associated with a second sub-region contained within the first sub-region. To associate a 3D region item with a 3D region item that indicates a 3D region inside it (i.e., contained within such a 3D region), information about the containment relationship is stored in the 'iref' box. The reference type indicating such a sub-region containment relationship (one sub-region being contained within another sub-region) is indicated by 'svrg', although this will be described in detail later. Note that 3D region items contained within other regions may be identified using a different 4CC (e.g., 'vvra') instead of the item type 'vran'. On the other hand, when associating a 3D region contained within another 3D region with another 3D image item, if a different 4CC is used for identification, it is necessary to separately specify a region item whose item type is vran. In other words, by defining a 3D region item as 'vran' regardless of whether it is contained within another region or not, various associations are possible. The value specified for 4CC can be any predefined value, and it is desirable that it be different from the value used for 4CC used for other purposes. Alternatively, to indicate a 3D region encompassed by another region, it may be possible to identify it using the flag value of VolumetricRegionItem shown in Figure 4. By associating and storing 3D regions encompassed by other 3D regions in this way, it becomes possible to indicate that the data is available for selective switching rather than selecting or displaying all annotation information associated with that 3D region during image file playback. This can be done, for example, by using viewpoint information or gaze information during playback.
[0135] The information processing device 100 may be configured to identify the target range in advance using item properties or other data structures, even for viewpoint position or line of sight direction. Such metadata can be stored, for example, as metadata intended to select (or display) annotation information associated with an internal 3D region item if the spatial position of the viewpoint specified by the playback device during playback is included in the 3D region indicated by the internal 3D region item included in another 3D region. Conversely, the information processing device 100 stores it as metadata intended not to select (display) annotation information associated with a wide area encompassing the region. On the other hand, if the spatial position of the viewpoint specified by the playback device during playback is not included in the 3D region indicated by the internal 3D region item, the metadata can be stored as metadata intended not to select (or display) annotation information associated with the internal 3D region item. Conversely, the information processing device 100 stores it as metadata intended to select (display) annotation information associated with a wide area encompassing the region. In other words, the information processing device 100 according to this embodiment may be configured to select (or display) annotation information associated with the lowest layer (the narrowest spatial range) of region information that indicates the location of the viewpoint position.
[0136] On the other hand, if the direction of line of sight is also considered in addition to the viewpoint position, the information processing device 100 may store it as metadata intended to select (or display) annotations associated with the video from a location where the viewpoint position is located. The information processing device 100 selects (or displays) annotation information associated with a 3D region item associated with a 3D image in the direction of line of sight from the top layer (the widest spatial range) region that contains the spatial position of the viewpoint specified by the playback device. In this case, if 3D sub-region information is not associated with the spatial position of the 3D image indicated by the viewpoint position, annotations directly associated with the 3D image are selected. If the spatial position is included in any of the 3D regions, annotation information associated with the interior of that 3D space or the 3D space itself is selected as the target for selection (or display). On the other hand, annotation information associated with other regions is not selected (or displayed). When the direction of line of sight is also considered, whether or not to select an internal region in the direction of line of sight from the viewpoint position (the target for displaying associated annotation information) can also be identified by whether or not a region in the 3D image indicated by a certain or larger 3D region is the target for display. If there is no target internal region, it can be selected in the same way as when only the viewpoint position is considered.
[0137] Furthermore, it is possible to configure the system to indicate that a certain subregion is included in another subregion, while also indicating that it will be selected by default (the region for which associated annotation information will be displayed). This can be done, for example, by using the flag value of VolumetricRegionItem to identify the relationship. Alternatively, an item property can be used to indicate that it is the default selected (or displayed) region. Furthermore, when selecting a region, the system can have data that selects (or displays) annotation information only if the region exceeds a certain threshold. This can also be specified using an item property. A 3D region designated as the default selected (or displayed) region may be selected (or displayed) regardless of the viewpoint or line of sight mentioned above, and may be selected (or displayed) if the region is included in the 3D image to be displayed, or if it is included in a certain amount of the image.
[0138] The information processing device 100 according to this embodiment can include information in its metadata indicating the display priority of which of the first annotation pieces of annotation information to display preferentially when both the first annotation piece of annotation information displayed in association with a first sub-region and the second annotation piece of annotation information displayed in association with a second sub-region included in the first sub-region satisfy the viewpoint conditions and are displayable. Setting such a display priority can be done in the same way as setting the priority of the viewpoint conditions in Embodiment 1. Here, for example, the information processing device 100 may set it so that the annotation piece of annotation information associated with the largest sub-region (since the first sub-region > the second sub-region, in this case the first sub-region) that satisfies the viewpoint conditions is displayed. Alternatively, the device may determine which of the two displayable annotation pieces to display according to user selection.
[0139] The information processing device 100 can also set a first viewpoint condition for the first annotation information and a second viewpoint condition for the second annotation information, and include them in the metadata. The viewpoint conditions here can be the same as those in Embodiments 1 and 2. For example, the information processing device 100 may generate a viewport based on the viewpoint position and line of sight direction, similar to Embodiment 2, and display the annotation information associated with the sub-region when a straight line from the viewpoint position to the reference point of each sub-region passes through the viewport. Alternatively, the information processing device 100 may display the annotation information associated with the sub-region when the angle between the straight line direction from the viewpoint position to the reference point of the sub-region and the line of sight direction is less than or equal to a predetermined threshold (e.g., 30°). Furthermore, the information processing device 100 may display the second annotation information if the second viewpoint condition is met when the viewpoint position is included in the first sub-region but outside the second sub-region, and not display the second annotation information if the second viewpoint condition is not met. Furthermore, if multiple pieces of annotation information exist, the information processing device 100 may set one or more of them to always be displayed as annotation information.
[0140] For example, when playing back a 3D image file, the information processing device 100 may display the second annotation information and not display the first annotation information if the viewpoint position is within the second sub-region. Alternatively, the information processing device 100 may display the first annotation information and not display the second annotation information if the viewpoint position is not within the second sub-region when playing back a 3D image file.
[0141] Next, with reference to Figures 34 to 36, an example of the file structure generated by the information processing device 100 according to this embodiment, and the structure of the metadata of such a file will be described. In this embodiment, the following description will be based on the assumption that an image file is generated that stores two 3D images and six 3D regions within the file data structure.
[0142] Figure 34 shows an example of a file structure generated by the information processing device 100 according to this embodiment. In the example in Figure 34, as shown in description 3403 corresponding to the 'mdat' box, descriptions 3430 and 3431 corresponding to Volumetric media Data are stored. Furthermore, in the example in Figure 34, descriptions 3432 and 3433 corresponding to VolumetricRegionItemData are stored in the image file. As shown in descriptions 3432 and 3433, the 3D region information data used here conforms to the definition shown in Figure 7, and each specifies a region of a rectangular parallelepiped. The region specified in description 3432 has the coordinates (x2, y2, z2) of the reference point in the coordinate space of the subregion, the shape size (sx2, sy2, sz2), and the components of the quaternion representation (qx2, qy2, qz2) specified. Similarly, the region specified in description 3433 has its coordinates (x8, y8, z8) of the reference point in the region's reference space, its shape size (sx8, sy8, sz8), and its quaternion representation components (qx8, qy8, qz8) specified.
[0143] Description 3401 corresponds to the 'ftyp' box, and in the example in Figure 34, 'gpc1' is described as the brand name (as major-brand) and 'mif1' as the compatible brand name (as compatible-brands).
[0144] Next, in description 3402, which corresponds to the 'meta' box, various metadata information describing the untimed data stored in the example output file is shown.
[0145] Description 3410 corresponds to the 'hdlr' box, and the handler type of the specified MetaBox('meta') is 'volv'. Description 3411 corresponds to the 'pitm' box, where 1 is stored as the item_ID, and the ID of the image to be displayed as the first priority image is specified.
[0146] Description 3412 corresponds to the 'iinf' box and shows item information (item ID (item_ID) or item type (item_type)) for each item. Description 3412 makes each item identifiable by its item ID and indicates what type of item the item identified by its item ID is. In the example in Figure 34, since 8 items are stored, entry_count is 8, and description 3412 contains 8 types of information, each with an item ID and item type specified. In the illustrated image file, the first and second pieces of information corresponding to descriptions 3440 and 3441 are G-PCC encoded image items of type 'gpe1'. Also, the third to eighth pieces of information corresponding to descriptions 3442 to 3447 are 3D region items of item type 'vran' that indicate a 3D region.
[0147] Here, we will explain the correspondence between each item in Figures 34, 35, and 36. Figure 36 is a schematic diagram of the metadata structure related to the 3D image items and 3D region items shown in the output file of Figure 34, as well as the annotation information associated with them. Figure 35 is an example of 3D image data corresponding to such an image file.
[0148] The 3D image corresponding to the G-PCC encoded image item described in description 3440 in Figure 34 corresponds to the rectangular prism 3500 shown by the solid line in Figure 35, and the 3D image item 3600 in Figure 36. Similarly, the 3D image corresponding to the G-PCC encoded image item described in description 3441 corresponds to the rectangular prism 3501 shown by the solid line in Figure 35, and the 3D image item 3601 in Figure 36. Next, the 3D subregion corresponding to the 3D region item described in description 3642 in Figure 34 corresponds to the rectangular prism 3510 shown by the dashed line in Figure 35, and the 3D region item 3610 in Figure 36. Likewise, the 3D subregion corresponding to the 3D region item described in description 3443 in Figure 34 corresponds to the rectangular prism 3520 shown by the dashed line in Figure 35, and the 3D region item 3620 in Figure 36. The 3D subregions corresponding to the 3D region item described in Figure 3444 correspond to the cuboid 3521 shown by the dashed line in Figure 35 and the 3D region item 3621 in Figure 36. The 3D subregions corresponding to the 3D region item described in Figure 3445 correspond to the cuboid 3530 shown by the dashed line in Figure 35 and the 3D region item 3630 in Figure 36. The 3D subregions corresponding to the 3D region item described in Figure 3446 correspond to the cuboid 1131 shown by the dashed line in Figure 35 and the 3D region item 3631 in Figure 36. The 3D subregions corresponding to the 3D region item described in Figure 3447 correspond to the cuboid 3540 shown by the dashed line in Figure 35 and the 3D region item 3640 in Figure 36.
[0149] As described above, the 3D region item shown in description 3443 is a subregion included in the 3D region item shown in description 3442, and is also a 3D region item that is directly associated with the 3D image item shown in description 3441. This can have effects such as allowing users to switch between images of the outside and inside of a building when 3D images acquired from outside and 3D images acquired from inside a building are stored in separate files.
[0150] Description 3413 corresponds to the 'iloc' box and describes the storage location and data size information for each item within the image file. For example, it indicates that the G-PCC encoded image item with item ID 1 is stored at offset 01 in the file and has a size of L1 bytes. In this way, by referring to description 3413, it is possible to determine the location of the data in the 'mdat' box.
[0151] Description 3414 corresponds to the 'iref' box and indicates the reference relationship (association) between each item. For the association between items shown in descriptions 3450 and 3451, the reference type is specified as 'cdsc', indicating a content description relationship. Description 3450 shows that the region information item with item ID 3 specified in from_item_ID references the G-PCC image item with item ID 1 specified in to_item_ID. This means that the 3D region information item with item ID 3 refers to a subregion within the G-PCC image item with item ID 1. This corresponds to the arrow (iref:cdsc) from 3D region item 3610 to 3D image item 3600 in Figure 36. Similarly, description 3451 shows that the 3D region information item with item ID 4 specified in from_item_ID references the G-PCC image item with item ID 2 specified in to_item_ID. As a result, the 3D region information item with item ID 4 represents a subregion within the G-PCC image item with item ID 2. This corresponds to the arrow (iref:cdsc) from 3D region item 3620 to 3D image item 3601 in Figure 36.
[0152] The associations between items shown in descriptions 3452, 3453, 3454, 3455, and 3456 have 'svrg' specified as the reference type, indicating a sub-region containment relationship (one sub-region is contained within the other sub-region). Description 3452 shows that a region information item with item ID 4, specified in from_item_ID, references a region information item with item ID 3, specified in to_item_ID. This indicates that the sub-region indicated by the 3D region information item with item ID 4 is a sub-region contained within the sub-region indicated by the region information item with item ID 3. This corresponds to the arrow (iref:svrg) from 3D region item 3620 to 3D region item 3610 in Figure 36.
[0153] Description 3453 indicates that the region information item with item ID 5 specified in from_item_ID references the region information item with item ID 3 specified in to_item_ID. This shows that the subregion indicated by the 3D region information item with item ID 5 is a subregion included in the subregion indicated by the region information item with item ID 3. This corresponds to the arrow (iref:svrg) from 3D region item 3621 to 3D region item 3610 in Figure 36.
[0154] Description 3454 indicates that the region information item with item ID 6 specified in `from_item_ID` references the region information item with item ID 4 specified in `to_item_ID`. This shows that the subregion indicated by the 3D region information item with item ID 6 is a subregion included in the subregion indicated by the region information item with item ID 4. This corresponds to the arrow (iref:svrg) from 3D region item 3630 to 3D region item 3620 in Figure 36.
[0155] Description 3455 indicates that the region information item with item ID 7 specified in from_item_ID references the region information item with item ID 4 specified in to_item_ID. This shows that the subregion indicated by the 3D region information item with item ID 7 is a subregion included in the subregion indicated by the region information item with item ID 4. This corresponds to the arrow (iref:svrg) from 3D region item 3630 to 3D region item 3620 in Figure 36.
[0156] Description 3456 indicates that the region information item with item ID 8 specified in `from_item_ID` references the region information item with item ID 6 specified in `to_item_ID`. This shows that the subregion indicated by the 3D region information item with item ID 8 is a subregion included in the subregion indicated by the region information item with item ID 6. This corresponds to the arrow (iref:svrg) from 3D region item 3640 to 3D region item 3630 in Figure 36.
[0157] By associating items of 3D domain information in this way, it becomes possible to show that one subdomain is included in another subdomain in the form of a pseudo-hierarchical structure (a structure that links spatially broad areas to spatially narrow areas) of associations between 3D domain items. This makes it possible to generate image files with metadata that allows for selective switching of annotation information associated with a subdomain that shows the interior of a 3D subdomain (a narrower space) within a 3D subdomain in a 3D image space.
[0158] Description 3415 corresponds to the 'iprp' box and includes Description 3420, which corresponds to the 'ipco' box, and Description 3421, which corresponds to the 'ipma' box. Description 3420 lists attribute information that may be used for each item or entity group as entry data. As shown in the figure, Description 3420 includes the first entry showing the G-PCC coding parameters, and the second and third entries showing the size of the 3D spatial region in Cartesian coordinates along the x, y, and z axes of the 3D image item. In addition, Description 3420 includes the fourth through ninth entries showing annotation information. Here, the annotation information has lang in American English (en-US) for all entries, and name is described as "map" for all entries. Furthermore, in the description section, entries 4 through 9 are listed in order as "X Shopping center," "X Shopping center A Building," "X Shopping center B Building," "X Shopping center A 1st floor," "X Shopping center A 2nd floor," and "Y Book Store X Shopping center," respectively, indicating the names of the locations on the map that the areas represent. Additionally, the tags section includes "AutoRecognition" to indicate that the area was automatically recognized.
[0159] The attribute information listed in description 3420 is associated with each item stored in the image file in the entry data of description 3421 corresponding to the 'ipma' box. In the example in Figure 34, the 3D image items with item IDs 1 and 2 are associated with a common 'gpcC' (property_index 1), indicating that they have the same G-PCC encoding parameters. The 3D image item with item ID 1 is associated with 'gpsr' (property_index 2), and the 3D image item with item ID 1 is associated with 'gpsr' (property_index 3), each specifying the size of the 3D spatial region. The 3D region information item with item ID 3 is associated with 'udes' (property_index 4), indicating annotation information associated with the subregion. The 3D region information item with item ID 4 is associated with 'udes' (property_index 5), indicating annotation information associated with the subregion. The 3D region information item with item ID 5 is associated with 'udes' (property_index is 6), and annotation information associated with the subregion is shown. The 3D region information item with item ID 6 is associated with 'udes' (property_index is 7), and annotation information associated with the subregion is shown. The 3D region information item with item ID 7 is associated with 'udes' (property_index is 8), and annotation information associated with the subregion is shown. The 3D region information item with item ID 8 is associated with 'udes' (property_index is 9), and annotation information associated with the subregion is shown. This corresponds to each 3D region item and the annotation information shown by the dashed lines in Figure 36.
[0160] The following describes the process of generating a 3D image file by the information processing device 100 according to this embodiment. Figure 37 is a flowchart showing an example of the process of generating a 3D image file executed by the information processing device 100 according to this embodiment. The process corresponding to the flowchart is realized by the CPU 101 reading the program stored in the ROM 102, expanding it in the RAM 103, and executing it, thereby operating each block. The process shown in Figure 37 may be executed, for example, in response to an operation start input by the user, or it may be started in response to image capture.
[0161] In S3701, the CPU 101 controls the imaging unit 104 or the image processing unit 105 to acquire 3D image data to be stored in a file. In S3702, the recognition processing unit 114 performs object detection processing within the 3D image. In S3703, the generation unit 113 generates region information relating to the partial region in the 3D image that indicates the object detected by the recognition processing unit 114. Since S3701 to S3703 are the same processes as S401 to S403 in Embodiment 1, a detailed explanation is omitted here.
[0162] In S3704, CPU101 determines whether to default select (replay) the annotation information related to the region information data generated in S3703. If default playback is selected, the process proceeds to S3705; otherwise, the process proceeds to S3706.
[0163] In S3705, the metadata processing unit 112 adds information indicating the default selected region to the generated region information data. S3705 can be executed, for example, by setting a specific bit (to 1) in the flags of the VolumetricRegionItem shown in Figure 4. In addition, S3705 may be executed by various predefined methods, such as associating properties with the region information item as item property information.
[0164] In S3706, CPU 101 determines whether the 3D region detected in S3702 is included in another detected 3D region. S3702 identifies whether the region to be processed is spatially included in another region. Here, it is determined whether one region contains the other region and whether it is a region that shows the internal structure of the other. In this example, the outer 3D region (a region that may contain other regions) is identified first, and then the 3D regions inside that 3D region are determined to be included in other 3D regions in turn. However, the process is not limited to this method as long as it is possible to similarly determine whether a given 3D region is included in (or contains) another 3D region. If it is determined that the 3D region detected in S3702 is included in another detected 3D region, the process proceeds to S3707; otherwise, the process proceeds to S3708.
[0165] In S3707, the metadata processing unit 112 generates metadata that associates the 3D region to be processed with other 3D region data that encompasses that 3D region. S3703 is performed by association using the item reference type 'svrg' shown in Figure 34.
[0166] In S3708, the metadata processing unit 112 associates annotation information with the 3D region to be processed. Here, it is assumed that text description information can be associated with the 3D region using item properties.
[0167] In S3709, CPU101 determines whether or not to perform detection of other objects (3D subregions). If detection is performed, the process returns to S3702; if detection is terminated, the process proceeds to S3710.
[0168] In S3710, the encoding / decoding unit 111 performs encoding processing on the 3D image data and saves the encoded data to the output buffer. The metadata processing unit 112 integrates the metadata created up to S3709 and the metadata necessary for decoding the encoded data, creates data with a 'meta' box structure, and saves it to the output buffer. In S3711, the metadata processing unit 112 combines the information in the 'ftyp' box related to the 3D image file, the information in the 'meta' box containing the final metadata, and the information in the 'mdat' box containing items such as encoded data and viewpoint condition information. The CPU 101 then writes the image file generated by storing the combined metadata and image data from RAM 103 to non-volatile memory 110 and saves it, ending the process shown in Figure 37. Since S3710 to S3711 are performed in the same way as S410 to S411, a detailed explanation is omitted here.
[0169] With this configuration, it is possible to generate a 3D image file that stores metadata corresponding to the 3D image, including a 3D image, first region information relating to a first subregion in the 3D image, second region information relating to a second subregion included in the first subregion, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. In particular, by adding metadata that indicates the inclusion relationships between subregions and storing the 3D regions in a pseudo-hierarchical structure (a structure that links spatially wide ranges to spatially narrow ranges), it is possible to generate an image file that can be used by selectively switching the annotation information to be displayed during playback processing.
[0170] In this embodiment, the image data stored in the image file is obtained by imaging by the imaging unit 104, but the image data used here is not particularly limited in this way. For example, the series of image data may be image data pre-stored in the ROM 102 or non-volatile memory 110, or image data received via the communication unit 108. In this case, the 3D image data may include an image file containing a single 3D still image. Furthermore, the image data may be image data encoded in an image file containing multiple 3D still image data, or it may be unencoded RAW image data.
[0171] Furthermore, although this embodiment has been described using a still image as a three-dimensional image, it is not particularly limited to this. For example, it is possible to define three-dimensional region information data using a three-dimensional moving image or a three-dimensional image sequence as a three-dimensional image, and similarly specify sub-regions contained within sub-regions according to type using a metadata structure for storing the three-dimensional moving image. Specifically, metadata related to a three-dimensional moving image or image sequence is described using a moov box, which is a metadata structure defined in ISOBMFF. A time-specified sequence of media data belonging to the presentation of the three-dimensional moving image data is represented in a trak box specified within the moov box. For example, a sequence of volume media frames, a sequence of sub-parts of volume media frames, or a sequence of time-specified metadata samples can be described as a time-specified sequence of media data. Within each track, each time unit of data is called a sample. This sample can be a frame or sub-part of a frame of volume media, video, audio, or time-specified metadata, or an image in an image sequence. Here, a sample is defined as all media data associated with the same presentation time within the track.
[0172] In this case, the 3D region information data is stored in the 'mdat' box, which stores the 3D volume media encoding data. The time-stamped sequence metadata describing the 3D region information data is configured as a track called a metadata track, which stores the time-stamped metadata. By associating the metadata track with the 3D video track using the tref box, it is possible to indicate the 3D region within the 3D video. Multiple 3D regions for different objects can be specified in a single 3D region metadata track, and each 3D region can be identified by an ID or similar within the region information data stored in 'mdat'. Samples can be grouped using a sample group structure, and by using a sample group description box, inclusion relationships and annotation information between 3D regions can be added.
[0173] This processing method allows for the individual identification of 3D region information not only in 3D still images but also in 3D video data, and further enables the identification of the spatial inclusion relationships of these regions. Furthermore, extensions using other video metadata structures may be implemented if similar processing is possible.
[0174] [Embodiment 5] Embodiment 5 describes the process of playing back a 3D image file generated by the information processing device 100 according to Embodiment 4. In this embodiment, the device for playing back the 3D image file is described as the information processing device 100, but an external device different from the information processing device 100 may be used as the playback device.
[0175] Figure 38 is a flowchart showing an example of the playback process performed by the information processing device 100 according to this embodiment. The process shown in Figure 38 can be realized by the CPU 101 reading the corresponding processing program stored in the ROM 102, loading it into the RAM 103, and executing it, thereby operating each block. The playback process shown in Figure 38 will be described as starting, for example, when the information processing device 100 is set to playback mode and an operation input related to the playback instruction of the image file is detected.
[0176] In S3801, CPU101 acquires the image file (target file) to be played back for which a playback command has been issued. In S3802, CPU101 acquires metadata and image data from the image file, and the metadata processing unit 112 analyzes the acquired metadata to understand the structure of the target file.
[0177] In S3803, the CPU 101 identifies a representative item based on the information in the 'pitm' box of the metadata and causes the encoding / decoding unit 111 to decode the encoded data 241 indicated by that representative item. The encoding / decoding unit 111 then obtains the corresponding encoded data from the metadata related to the image item designated as the representative image, performs the decoding process, and saves the decoded data to a buffer on RAM 103.
[0178] In S3804, the metadata processing unit 112 acquires 3D region data associated with the image to be played back, which is designated as the representative image. In S3805, the metadata processing unit 112 acquires information about the user's viewpoint position for displaying the 3D image during playback. Here, the metadata processing unit 112 can acquire the viewpoint position based on user input. Alternatively, for example, the metadata processing unit 112 may set the viewpoint position from information indicating where the user is in a pre-identified space.
[0179] In S3806, the metadata processing unit 112 determines whether or not there is 3D region data indicating a 3D region containing coordinates in the 3D space indicated by the acquired viewpoint position information. If no such data is found, the process proceeds to S3807; if it is found, the process proceeds to S3808.
[0180] In S3807, the metadata processing unit 112 selects (displays) annotation information for the highest-level (spatially widest) 3D region directly associated with the 3D image. Here, the metadata processing unit 112 displays representative image data stored in the buffer and outputs the associated region data along with the representative image data to the display unit 106 to display the image. If region annotation information corresponding to a sub-region is provided, it is displayed in a way that allows identification of the relationship with the region information. When S3807 is completed, the process proceeds to S3812.
[0181] In S3808, the metadata processing unit 112 determines whether or not to use gaze direction information for selecting annotation information (area information data). Whether or not to use gaze direction information is specified in advance in the settings of the information processing device 100. If gaze direction information is not used, the process proceeds to S3811; if it is used, the process proceeds to S3809.
[0182] In S3811, the metadata processing unit 112 selects (displays) annotations for the lowest layer (the narrowest spatial range) of the 3D region, which includes the coordinates in the 3D space indicated by the viewpoint position information. Here, for example, annotations associated with a spatial location where the specified position is located are selected (displayed). Then, similar to S3807, the region information and its associated annotation information are superimposed and displayed together with the representative image data stored in the buffer. If S3811 is completed, the process proceeds to S3812.
[0183] In S3809, the metadata processing unit 112 acquires information about the direction of the user's gaze from their viewpoint. The process of acquiring information about the direction of the user's gaze from their viewpoint may be performed based on user input, as in S3805, or it may be performed by detecting the direction the user is facing in a pre-identified space using a head-mounted display or the like.
[0184] In S3810, the metadata processing unit 112 selects (displays) annotation information for the area indicated by the 3D area information that is included in the display area within the line of sight from the viewpoint position for a certain amount or more. Then, as described above, the annotation information and the 3D image data are superimposed on each other. If there is no area included for a certain amount or more, the annotation information for the area indicated by the viewpoint position is selected (displayed) in the same way as in S3811. When S3811 is completed, the process proceeds to S3812.
[0185] In S3812, the metadata processing unit 112 determines whether the default-specified 3D region is stored (exists). If the default-specified region does not exist, the process shown in Figure 38 ends; otherwise, the process proceeds to S3813.
[0186] In S3813, the metadata processing unit 112 selects an additional default-specified 3D region, and then selects (displays) the 3D region data and associated annotation information before terminating the process. Regarding processing related to viewpoint position and line of sight, the selected (displayed) region data and annotation information may be switched each time the viewpoint position or line of sight information changes, or the data selected once may be retained without such switching. Furthermore, if 3D image data indicating the interior is associated with region data indicating the interior, and the interior region data and annotations are selected, the display may be switched from the representative image to the 3D image indicating the interior, or both the representative image and the interior image may be displayed.
[0187] With this configuration, it is possible to perform playback processing of a 3D image file that stores metadata corresponding to the 3D image, including a 3D image, first region information relating to a first subregion in the 3D image, second region information relating to a second subregion included in the first subregion, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. In particular, by storing the region information as metadata with a pseudo-hierarchical structure, it is possible to selectively switch and treat the regions and their annotation information as usable data when selecting (displaying) them. Furthermore, by making it possible to identify the default selected (displayed) region data and annotations that are not subject to switching, it becomes possible to separately identify the regions that are always selected.
[0188] In this embodiment, annotation information and region data information have been described as being displayed superimposed on the representative image, but the display of annotation information and region data information may be optional. In other words, whether or not to display this information may be set based on instructions such as UI operation.
[0189] The disclosures herein include the following information processing devices, information processing methods, and programs. (Item 1) An acquisition means for acquiring encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes region information relating to a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the 3D space of the 3D image, associated with the region information or the annotation information. A generation means for generating a 3D image file that stores the encoded data of the 3D image and the metadata, An information processing device equipped with the following features. (Item 2) The information processing device according to item 1, characterized in that the conditions for displaying the annotation information include a condition for specifying a first range in which the viewpoint position in the three-dimensional space is located. (Item 3) The information processing apparatus according to item 2, characterized in that the first range is a range set based on user input. (Item 4) The information processing apparatus according to any one of items 1 to 3, characterized in that the conditions for displaying the annotation information include conditions based on the partial region and the projection range of the three-dimensional image onto the display plane, which is set based on the line of sight direction in the three-dimensional space. (Item 5) The information processing device according to item 4, characterized in that the conditions for displaying the annotation information include conditions based on the positional relationship between a straight line from the viewpoint position to a reference point of the partial region and the projection range. (Item 6) The information processing device according to item 5, wherein the annotation information is displayed when a straight line from the viewpoint position to the reference point passes through the projection range. (Item 7) The information processing device according to item 2, characterized in that the conditions for displaying the annotation information further include conditions for specifying a second range in which the viewpoint position in the three-dimensional space is located. (Item 8) The information processing apparatus according to item 7, characterized in that the second range is a range in which the distance between the reference point of the sub-region in the three-dimensional space and the viewpoint position satisfies a predetermined condition. (Item 9) The information processing device according to item 7, characterized in that the second range is a range in which the distance between two points obtained by projecting the reference point of the sub-region in the three-dimensional space and the viewpoint position onto the xy-plane satisfies a predetermined condition. (Item 10) The aforementioned condition information includes first condition information indicating a first condition relating to the display of first annotation information, and second condition information indicating a second condition relating to the display of second annotation information. The information processing device according to any one of items 1 to 9, characterized in that the metadata acquired by the acquisition means further includes priority information indicating whether to display the first annotation information or the second annotation information when the first condition and the second condition are met. (Item 11) The information processing device according to any one of items 1 to 10, characterized in that the metadata acquired by the acquisition means further includes information on default annotation information to be displayed in the three-dimensional image when the conditions for displaying the annotation information are not met. (Item 12) The information processing device according to any one of items 1 to 11, characterized in that the metadata acquired by the acquisition means further includes information on the reference point of the subregion. (Item 13) An acquisition means for acquiring encoded data of a three-dimensional image, and metadata corresponding to the encoded data of the three-dimensional image, which includes first region information relating to a first subregion in the three-dimensional image, second region information relating to a second subregion included in the first subregion in the three-dimensional image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. A generation means for generating a 3D image file that stores the encoded data of the 3D image and the metadata, An information processing device equipped with the following features. (Item 14) The information processing device according to item 13, wherein the metadata further includes condition information indicating a first condition for displaying the first annotation information and a second condition for displaying the second annotation information, corresponding to the viewpoint position or line of sight direction in the three-dimensional space of the three-dimensional image file when the three-dimensional image file is played back. (Item 15) The information processing device according to item 14, characterized in that the condition information includes a condition that, when the three-dimensional image file is played back, the viewpoint position is included in the first subregion but not in the second subregion, and if the second condition is met, the second annotation information is displayed, and if the second condition is not met, the second annotation information is not displayed. (Item 16) The information processing device according to item 14 or 15, characterized in that the condition information includes a condition that displays first annotation information when a straight line from the viewpoint position to a reference point of the first partial region passes through the projection range of the three-dimensional image onto the display plane, which is set based on the line of sight direction in the three-dimensional space, and displays second annotation information when a straight line from the viewpoint position to a reference point of the second partial region passes through the projection range of the three-dimensional image onto the display plane. (Item 17) The information processing device according to item 13, characterized in that one or both of the first annotation information and the second annotation information are annotation information that is always displayed when the three-dimensional image file is played back. (Item 18) The information processing device according to any one of items 13 to 17, characterized in that the metadata further includes a third sub-region associated with the first sub-region and different from the second sub-region, and a third annotation information associated with the third sub-region. (Item 19) An acquisition means for acquiring a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes region information relating to a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the 3D space of the 3D image, associated with the region information or the annotation information. A playback means for playing back the aforementioned 3D image file, An information processing device characterized by comprising: (Item 20) An acquisition means for acquiring a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes first region information relating to a first subregion in the 3D image, second region information relating to a second subregion included in the first subregion in the 3D image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. A playback means for playing back the aforementioned 3D image file, An information processing device equipped with the following features. (Item 21) A step of acquiring encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes region information relating to a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the 3D space of the 3D image, associated with the region information or the annotation information. A step of generating a 3D image file that stores the encoded data of the 3D image and the metadata, An information processing method comprising: (Item 22) A step of obtaining encoded data of a three-dimensional image, and metadata corresponding to the encoded data of the three-dimensional image, which includes first region information relating to a first subregion in the three-dimensional image, second region information relating to a second subregion included in the first subregion in the three-dimensional image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. A step of generating a 3D image file that stores the encoded data of the 3D image and the metadata, An information processing method comprising: (Item 23) A step of obtaining a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes region information relating to a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the 3D space of the 3D image, associated with the region information or the annotation information. The process of playing back the aforementioned 3D image file, An information processing method characterized by comprising: (Item 24) A step of obtaining a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes first region information relating to a first subregion in the 3D image, second region information relating to a second subregion included in the first subregion in the 3D image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. The process of playing back the aforementioned 3D image file, An information processing method comprising: (Item 25) A program to cause a computer to function as one of the means of an information processing device described in any one of items 1 through 20.
[0190] (Other examples) The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0191] The invention is not limited to the embodiments described above, and various modifications and variations are possible without departing from the spirit and scope of the invention. Accordingly, claims are attached to disclose the scope of the invention. [Explanation of Symbols]
[0192] 101: CPU, 102: ROM, 103: RAM, 104: Imaging unit, 105: Image processing unit, 106: Display unit, 107: Operation input unit, 108: Communication unit, 109: System bus, 110: Non-volatile memory
Claims
1. An acquisition means for acquiring encoded data of a three-dimensional image, metadata corresponding to the encoded data of the three-dimensional image, which includes region information relating to a partial region in the three-dimensional image, annotation information associated with the partial region, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the three-dimensional space of the three-dimensional image, associated with the region information or the annotation information. A generation means for generating a 3D image file that stores the encoded data of the 3D image and the metadata, An information processing device equipped with the following features.
2. The information processing apparatus according to claim 1, characterized in that the conditions for displaying the annotation information include a condition for specifying a first range in which the viewpoint position in the three-dimensional space is located.
3. The information processing apparatus according to claim 2, characterized in that the first range is a range set based on user input.
4. The information processing apparatus according to claim 1, characterized in that the conditions for displaying the annotation information include conditions based on the partial region and the projection range of the three-dimensional image onto the display plane, which is set based on the line of sight direction in the three-dimensional space.
5. The information processing apparatus according to claim 4, characterized in that the conditions for displaying the annotation information include conditions based on the positional relationship between a straight line from the viewpoint position to a reference point of the partial region and the projection range.
6. The information processing device according to claim 5, wherein the annotation information is displayed when a straight line from the viewpoint position to the reference point passes through the projection range.
7. The information processing apparatus according to claim 2, characterized in that the conditions for displaying the annotation information further include conditions for specifying a second range in which the viewpoint position in the three-dimensional space is located.
8. The information processing apparatus according to claim 7, characterized in that the second range is a range in which the distance between the reference point of the sub-region in the three-dimensional space and the viewpoint position satisfies a predetermined condition.
9. The information processing apparatus according to claim 7, characterized in that the second range is a range in which the distance between two points obtained by projecting the reference point of the sub-region in the three-dimensional space and the viewpoint position onto the xy-plane satisfies a predetermined condition.
10. The aforementioned condition information includes first condition information indicating a first condition relating to the display of first annotation information, and second condition information indicating a second condition relating to the display of second annotation information. The information processing apparatus according to claim 1, wherein the metadata acquired by the acquisition means further includes priority information indicating whether to display the first annotation information or the second annotation information when the first condition and the second condition are met.
11. The information processing apparatus according to claim 1, characterized in that the metadata acquired by the acquisition means further includes information on default annotation information to be displayed in the three-dimensional image when the conditions for displaying the annotation information are not met.
12. The information processing apparatus according to claim 1, characterized in that the metadata acquired by the acquisition means further includes information on the reference point of the subregion.
13. An acquisition means for acquiring encoded data of a three-dimensional image, and metadata corresponding to the encoded data of the three-dimensional image, which includes first region information relating to a first subregion in the three-dimensional image, second region information relating to a second subregion included in the first subregion in the three-dimensional image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. A generation means for generating a 3D image file that stores the encoded data of the 3D image and the metadata, An information processing device equipped with the following features.
14. The information processing apparatus according to claim 13, characterized in that the metadata further includes condition information indicating a first condition for displaying the first annotation information and a second condition for displaying the second annotation information, corresponding to the viewpoint position or line of sight direction in the three-dimensional space of the three-dimensional image file when the three-dimensional image file is played back.
15. The information processing apparatus according to claim 14, characterized in that the condition information includes a condition that, when the three-dimensional image file is played back, the viewpoint position is included in the first partial region but not in the second partial region, and if the second condition is met, the second annotation information is displayed, and if the second condition is not met, the second annotation information is not displayed.
16. The information processing apparatus according to claim 14, characterized in that the condition information includes a condition that displays first annotation information when a straight line from the viewpoint position to a reference point of the first partial region passes through the projection range of the three-dimensional image onto the display plane, which is set based on the line of sight direction in the three-dimensional space, and displays second annotation information when a straight line from the viewpoint position to a reference point of the second partial region passes through the projection range of the three-dimensional image onto the display plane.
17. The information processing apparatus according to claim 13, characterized in that one or both of the first annotation information and the second annotation information are annotation information that is always displayed when the three-dimensional image file is played back.
18. The information processing apparatus according to claim 13, characterized in that the metadata further includes a third sub-region associated with the first sub-region and different from the second sub-region, and a third annotation information associated with the third sub-region.
19. An acquisition means for acquiring a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes region information relating to a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the 3D space of the 3D image, associated with the region information or the annotation information. A playback means for playing back the aforementioned three-dimensional image file, An information processing device characterized by comprising:
20. An acquisition means for acquiring a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes encoded data of a 3D image, first region information relating to a first subregion in the 3D image, second region information relating to a second subregion included in the first subregion in the 3D image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. A playback means for playing back the aforementioned three-dimensional image file, An information processing device equipped with the following features.
21. A step of acquiring encoded data of a three-dimensional image, metadata corresponding to the encoded data of the three-dimensional image, which includes region information relating to a sub-region in the three-dimensional image, annotation information associated with the sub-region, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the three-dimensional space of the three-dimensional image, associated with the region information or the annotation information. A step of generating a three-dimensional image file that stores the encoded data of the three-dimensional image and the metadata, An information processing method comprising:
22. A step of obtaining encoded data of a three-dimensional image, and metadata corresponding to the encoded data of the three-dimensional image, which includes first region information relating to a first subregion in the three-dimensional image, second region information relating to a second subregion included in the first subregion in the three-dimensional image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. A step of generating a three-dimensional image file that stores the encoded data of the three-dimensional image and the metadata, An information processing method comprising:
23. A step of obtaining a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes region information relating to a subregion of the 3D image, annotation information associated with the subregion, and condition information indicating conditions for displaying the annotation information according to the viewpoint position or line of sight direction in the 3D space of the 3D image, associated with the region information or the annotation information. The process of playing back the aforementioned three-dimensional image file, An information processing method characterized by comprising:
24. A step of obtaining a 3D image file that stores encoded data of a 3D image, metadata corresponding to the encoded data of the 3D image, which includes first region information relating to a first subregion in the 3D image, second region information relating to a second subregion included in the first subregion in the 3D image, first annotation information associated with the first subregion, and second annotation information associated with the second subregion. The process of playing back the aforementioned three-dimensional image file, An information processing method comprising:
25. A program for causing a computer to function as one of the means of an information processing device according to any one of claims 1 to 20.