Information processing device and method

The information processing device and method generate a scene description with low-latency controls to manage interaction playback delays and loads, ensuring consistent quality and alignment with content provider intentions.

WO2025206029A1PCT designated stage Publication Date: 2025-10-02SONY GROUP CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/012177
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-28
Filing Date
2025-03-26
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Conventional interaction playback methods in 3D environments often result in unacceptable delays and increased processing loads due to unpredictable trigger conditions, leading to a decrease in quality and difficulty in aligning with the content provider's intentions.

Method used

An information processing device and method that generates a scene description with low-latency processing information, including a maximum processing delay, to control interaction playback, ensuring actions are executed within a specified delay threshold, thereby maintaining intended quality and reducing unnecessary processing loads.

Benefits of technology

The solution ensures low-latency interaction playback that aligns with the content provider's intentions, minimizing delays and processing loads, thus enhancing the quality of interaction playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025012177_02102025_PF_FP_ABST
    Figure JP2025012177_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an information processing device and method that make it possible to suppress a reduction in quality of interaction replay and / or useless high load processing. The present disclosure generates a scene description representing a scene of a 3D space in which a 3D object is located, and stores, in the scene description, low-latency processing information including a maximum processing delay, for controlling a low-latency interaction relay. Further, the present disclosure performs low-latency interaction replay on the basis of the low-latency processing information stored in the scene description. The present disclosure can be applied to, for example, an information processing device or an information processing method.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device and method

[0001] The present disclosure relates to an information processing device and method, and more particularly to an information processing device and method that can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0002] Conventionally, there is the glTF (The GL Transmission Format) (registered trademark) 2.0, a scene description format for placing and rendering 3D (three-dimensional) objects in a three-dimensional space (see, for example, Non-Patent Document 1). Furthermore, a method for handling dynamic content in the time direction by extending glTF 2.0 has been proposed for MPEG-I scene description (MPEG (Moving Picture Experts Group)-I Scene Description) (see, for example, Non-Patent Document 2). Furthermore, in parallel with the standardization of compression and transmission technology for haptic media, the AMD2 version of MPEG-I scene description is also advancing standardization of technology extensions for handling haptic media in a 3D space and interaction technology extensions (see, for example, Non-Patent Document 3). Furthermore, in addition to audio and video media, which are components of 2D video content and 3DoF (Degree of Freedom) / 6DoF video content, standardization of compression and transmission technology for haptic media, which compresses haptic information, is also advancing (see, for example, Non-Patent Document 4).

[0003] In interaction playback, which outputs media when a trigger condition is met, a method of driving an output device after the trigger condition is met can result in an unacceptable delay in the output timing of the action from the trigger timing due to factors such as the startup time of the output device. To achieve low-latency interaction playback, one possible method is to predict the occurrence of an interaction and start up the device before the trigger condition is met.

[0004] However, even with such methods, predictions can be wrong, which can result in a decrease in the quality of interaction playback or an increase in unnecessary processing load. The tolerance for delays and erroneous presentations in interaction playback, as well as the tolerance for processing load, can depend on the content (scene and situation). Therefore, it is desirable to have interaction playback behave as close as possible to the intentions of the content provider (e.g., author). In other words, there is a need for a playback method that can bring the amount of delay in interaction playback, the frequency of erroneous presentations, the amount of processing load, and other factors closer to the intentions of the content provider (e.g., author).

[0005] Saurabh Bhatia, Patrick Cozzi, Alexey Knyazev, Tony Parisi, "Khronos glTF2.0", https: / / github.com / KhronosGroup / glTF / tree / master / specification / 2.0, June 9, 2017"Information technology - Coded representation of immersive media - Part 14: Scene Description", ISO / IEC DIS 23090-14:2021(E), ISO / IEC JTC 1 / SC 29 / WG 03 N00485, 137th MPEG meeting, January 2022, online"Revised text of ISO / IEC 23090-14 DAM 2: Support for Haptics, Augmented Reality, Avatars, Interactivity, MPEG-I Audio, and Lighting", ISO / IEC JTC 1 / SC 29 / WG 03 N01025, 2024-02-27"Preliminary Draft Text of FDIS ISO / IEC 23090-31 Haptics coding", ISO / IEC JTC 1 / SC 29 / WG 7 N765, 2023-11-10

[0006] However, with conventional methods, such interaction playback is performed solely by the device itself, making it difficult to reflect the intentions of the content provider (e.g., author) in the interaction playback. As a result, there is a risk that the interaction playback may behave differently from the intentions of the content provider (e.g., author). For example, there is a risk of incorrect presentation or delays of actions that are different from the intentions of the content provider (e.g., author). This could result in a decrease in the quality of the interaction playback or an unnecessary increase in the processing load.

[0007] The present disclosure has been made in light of such circumstances, and makes it possible to suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0008] An information processing device according to one aspect of the present technology includes a scene description generation unit that generates a scene description representing a scene in a 3D space in which a 3D object is placed, and stores, in the scene description, low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, wherein the low-latency interaction playback is a process in which interaction playback, which is triggered by the occurrence of an interaction in the 3D space and executes an action corresponding to the interaction, is executed with a delay lower than the maximum processing delay, and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

[0009] An information processing method according to one aspect of the present technology generates a scene description that represents a scene in a 3D space in which a 3D object is placed, and stores low-latency processing information in the scene description that includes a maximum processing delay and controls low-latency interaction playback, the low-latency interaction playback being a process in which interaction playback that executes an action corresponding to an interaction triggered by the occurrence of the interaction in the 3D space is executed with a delay lower than the maximum processing delay, and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of the action on an output device.

[0010] Another aspect of the present technology is an information processing device that includes a scene processing unit that selects an action that corresponds to an interaction in a 3D space and is capable of low-latency interaction playback, based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, and that performs the low-latency interaction playback for the selected action, based on low-latency processing information that is stored in a scene description that represents a scene in the 3D space in which a 3D object is placed, and that selects an action that is capable of low-latency interaction playback, and performs the low-latency interaction playback for the selected action, wherein the low-latency interaction playback is a process in which interaction playback that executes the action in response to the occurrence of the interaction as a trigger is executed with a delay lower than the maximum processing delay, and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

[0011] Another aspect of the present technology is an information processing method that selects an action that corresponds to an interaction in a 3D space and is capable of low-latency interaction playback, based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, and that is stored in a scene description that represents a scene in the 3D space in which a 3D object is placed, and performs the low-latency interaction playback for the selected action, wherein the low-latency interaction playback is a process in which interaction playback that executes the action in response to the occurrence of the interaction as a trigger is executed with a delay lower than the maximum processing delay, and the maximum processing delay indicates a maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

[0012] In an information processing device and method according to one aspect of the present technology, a scene description representing a scene in 3D space in which a 3D object is placed is generated, and low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback is stored in the scene description.

[0013] In another aspect of the information processing device and method of the present technology, an action that corresponds to an interaction in the 3D space and is capable of low-latency interaction playback is selected based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, and that is stored in a scene description that represents a scene in 3D space in which a 3D object is placed, and low-latency interaction playback is performed for the selected action.

[0014] 1 is a diagram illustrating an example of the main configuration of glTF2.0. A diagram illustrating an example of glTF objects and reference relationships. A diagram illustrating an example of a description of a scene description. A diagram illustrating a method of accessing binary data. A diagram illustrating an example of a description of a scene description. A diagram illustrating the relationship between a buffer object, a buffer view object, and an accessor object. A diagram illustrating an example of a description of a buffer object, a buffer view object, and an accessor object. A diagram illustrating an example of a configuration of a scene description object. A diagram illustrating an example of a description of a scene description. A diagram illustrating a method of extending an object. A diagram illustrating the configuration of client processing. A diagram illustrating an example of a configuration of an extension for handling timed metadata. A diagram illustrating an example of a description of a scene description. A diagram illustrating an example of a description of a scene description. A diagram illustrating an example of a configuration of an extension for handling timed metadata. A diagram illustrating an example of the main configuration of a client. A flowchart illustrating an example of the flow of client processing. A diagram illustrating an example of behavior information. A diagram illustrating trigger types. A diagram illustrating action types. A diagram illustrating behaviors. A flowchart illustrating an example of the flow of trigger activation processing. A diagram illustrating an example of an extension for haptic media. A diagram illustrating an example of a codec architecture. A diagram illustrating an example of an interaction playback method. FIG. 1 is a diagram showing an example of extended behavior information; FIG. 2 is a diagram showing an example of low latency processing information; FIG. 3 is a diagram showing an example of interaction playback with prediction; FIG. 4 is a diagram showing an example of extended behavior information; FIG. 5 is a diagram showing an example of predicted trigger activation control information; FIG. 6 is a diagram for explaining an example of how prediction occurs; FIG. 7 is a diagram showing an example of interaction playback with prediction; FIG. 8 is a diagram showing an example of extended behavior information; FIG. 9 is a diagram showing an example of predicted trigger activation offload control information; FIG. 10 is a diagram for explaining an example of how prediction occurs; A block diagram showing an example of the main configuration of a file generation device; A flowchart showing an example of the flow of a file generation process; A flowchart showing an example of the flow of a file generation process.1 is a flowchart showing an example of the flow of a file generation process. FIG. 2 is a block diagram showing an example of the main configuration of a distribution server. FIG. 3 is a flowchart showing an example of the flow of a content acquisition process. FIG. 4 is a flowchart showing an example of the flow of a distribution process. FIG. 5 is a block diagram showing an example of the main configuration of a client device and an output device. FIG. 6 is a flowchart showing an example of the flow of a playback process. FIG. 7 is a flowchart showing an example of the flow of a synchronous media playback process. FIG. 8 is a flowchart showing an example of the flow of an asynchronous media low latency interaction playback process. FIG. 9 is a flowchart showing an example of the flow of an asynchronous media low latency interaction playback process. FIG. 10 is a flowchart showing an example of the flow of an output process. FIG. 11 is a block diagram showing an example of the main configuration of a computer.

[0015] Below, modes for carrying out the present disclosure (hereinafter referred to as embodiments) will be described. The description will be made in the following order: 1. Literature, etc. supporting technical content and technical terminology 2. Interactive playback of asynchronous media 3. Low-latency interactive playback based on author's intentions 4. First embodiment (file generation device) 5. Second embodiment (distribution server) 6. Third embodiment (client device / output device) 7. Supplementary notes

[0016] <1. Literature, etc. supporting technical content and technical terminology> The scope of disclosure of the present technology includes not only the content described in the embodiments, but also the content described in the following non-patent documents and patent documents that were publicly known at the time of filing, as well as the content of other documents referenced in the following non-patent documents and patent documents.

[0017] Non-patent document 1: (as mentioned above) Non-patent document 2: (as mentioned above) Non-patent document 3: (as mentioned above) Non-patent document 4: (as mentioned above) Patent document 1: WO 2024 / 34336 Patent document 2: WO 2024 / 14526

[0018] In other words, the contents of the above-mentioned non-patent documents and patent documents, as well as the contents of other documents referenced in the above-mentioned non-patent documents and patent documents, are also used as the basis for determining the support requirements. For example, even if syntax and terminology such as glTF2.0 and its extensions described in the above-mentioned non-patent documents and patent documents are not directly defined in this disclosure, they are considered to be within the scope of this disclosure and meet the support requirements of the claims. Similarly, even if technical terms such as parsing, syntax, and semantics are not directly defined in this disclosure, they are considered to be within the scope of this disclosure and meet the support requirements of the claims.

[0019] <2. Interactive Playback of Asynchronous Media> <gltf2.0> Conventionally, as described in Non-Patent Document 1, for example, there is glTF (The GL Transmission Format) (registered trademark) 2.0, which is a scene description format for placing and rendering 3D (three-dimensional) objects within an area (e.g., three-dimensional space). As shown in FIG. 1, glTF2.0 is composed of a JSON format file (.glTF), a binary file (.bin), and an image file (.png, .jpg, etc.). The binary file stores binary data such as geometry and animation. The image file stores data such as texture.

[0020] A JSON format file is a scene description file written in JSON (JavaScript (registered trademark) Object Notation). A scene description is metadata that describes (a description of) a scene in 3D content. The description of this scene description defines what kind of scene it is. A scene description file is a file that stores such a scene description.

[0021] The description of the JSON format file consists of a list of key and value pairs. An example of the format is shown below: “KEY”:”VALUE”

[0022] Keys consist of strings, and values ​​consist of numbers, strings, booleans, arrays, objects, or null.

[0023] Additionally, multiple key-value pairs ("KEY":"VALUE") can be grouped together using {} (curly brackets). This grouping is also called a JSON object. An example of the format is shown below: "user":{"id":1, "name":"tanaka"}

[0024] In this example, a JSON object that combines the pair "id":1 and the pair "name":"tanaka" is defined as the value corresponding to the key (user).

[0025] Additionally, zero or more values ​​can be organized into an array using square brackets []. This array is also called a JSON array. For example, a JSON object can be applied as an element of this JSON array. An example of the format is shown below. test":["hoge", "fuga", "bar"] "users":[{"id":1, "name":"tanaka"},{"id":2,"name":"yamada"},{"id":3, "name":"sato"}]

[0026] Figure 2 shows the glTF objects that can be written at the top level of a JSON format file and the reference relationships they can have. The long circles in the tree structure shown in Figure 2 represent objects, and the arrows between those objects indicate the reference relationships. As shown in Figure 2, objects such as "scene", "node", "mesh", "camera", "skin", "material", and "texture" are written at the top level of a JSON format file.

[0027] An example of such a JSON format file (scene description) is shown in FIG. 3. The JSON format file 20 in FIG. 3 shows an example of a portion of the top-level description. In this JSON format file 20, all top-level objects 21 used are described at the top level. These top-level objects 21 are the glTF objects shown in FIG. 2. Furthermore, in the JSON format file 20, reference relationships between objects are indicated as indicated by arrows 22. More specifically, the reference relationships are indicated by specifying the index of an element in the array of the referencing object in the property of the higher-level object.

[0028] FIG. 4 is a diagram illustrating a method for accessing binary data. As shown in FIG. 4, binary data is stored in a buffer object. That is, information for accessing the binary data (e.g., a uniform resource identifier (URI)) is indicated in the buffer object. In a JSON format file, as shown in FIG. 4, objects such as a mesh, camera, and skin can access the buffer object via an accessor object and a bufferView object.

[0029] That is, for objects such as mesh, camera, and skin, the accessor object to be referenced is specified. An example of a description of a mesh object in a JSON format file is shown in Figure 5. For example, as shown in Figure 5, in a mesh object, vertex attributes such as NORMAL, POSITION, TANGENT, and TEXCORD_0 are defined as keys, and for each attribute, the accessor object to be referenced is specified as a value.

[0030] The relationship between buffer objects, buffer view objects, and accessor objects is shown in Figure 6. An example of how these objects are written in a JSON format file is shown in Figure 7.

[0031] 6, buffer object 41 is an object that stores information (such as a URI) for accessing binary data, which is actual data, and information indicating the data length (e.g., byte length) of that binary data. A in FIG. 7 shows an example of the description of buffer object 41. "bytelength":102040" shown in A in FIG. 7 indicates that the byte length of buffer object 41 is 102040 bytes, as shown in FIG. 6. Furthermore, "uri":"duck.bin" shown in A in FIG. 7 indicates that the URI of buffer object 41 is "duck.bin", as shown in FIG. 6.

[0032] 6, the buffer view object 42 is an object that stores information about a subset area of ​​binary data specified in the buffer object 41 (i.e., information about a partial area of ​​the buffer object 41). B of Fig. 7 shows an example of description of the buffer view object 42. As shown in Fig. 6 and B of Fig. 7, the buffer view object 42 stores information such as identification information of the buffer object 41 to which the buffer view object 42 belongs, an offset (e.g., a byte offset) indicating the position of the buffer view object 42 within the buffer object 41, and a length (e.g., a byte length) indicating the data length (e.g., byte length) of the buffer view object 42.

[0033] As shown in B of Figure 7, when there are multiple buffer view objects, information is written for each buffer view object (i.e., for each subset area). For example, information such as "buffer":0, "bytelength":25272, and "byteOffset":0 shown at the top of B of Figure 7 is information for the first buffer view object 42 (bufferView[0]) shown in the buffer object 41 in Figure 6. Furthermore, information such as "buffer":0, "bytelength":76768, and "byteOffset":25272 shown at the bottom of B of Figure 7 is information for the second buffer view object 42 (bufferView[1]) shown in the buffer object 41 in Figure 6.

[0034] "buffer":0" of the first buffer view object 42 (bufferView[0]) shown in B of FIG. 7 indicates that the identification information of the buffer object 41 to which the buffer view object 42 (bufferView[0]) belongs is "0" (Buffer[0]), as shown in FIG. 6. Also, "bytelength":25272" indicates that the byte length of the buffer view object 42 (bufferView[0]) is 25272 bytes. Furthermore, "byteOffset":0" indicates that the byte offset of the buffer view object 42 (bufferView[0]) is 0 bytes.

[0035] "buffer":0" of the second buffer view object 42 (bufferView[1]) shown in FIG. 7B indicates that the identification information of the buffer object 41 to which the buffer view object 42 (bufferView[0]) belongs is "0" (Buffer[0]), as shown in FIG. 6. Furthermore, "bytelength":76768" indicates that the byte length of the buffer view object 42 (bufferView[0]) is 76768 bytes. Furthermore, "byteOffset":25272" indicates that the byte offset of the buffer view object 42 (bufferView[0]) is 25272 bytes.

[0036] 6, the accessor object 43 is an object that stores information about how to interpret data in the buffer view object 42. C in Fig. 7 shows an example of the description of the accessor object 43. As shown in Fig. 6 and C in Fig. 7, the accessor object 43 stores information such as identification information of the buffer view object 42 to which the accessor object 43 belongs, an offset (e.g., a byte offset) indicating the position of the buffer view object 42 within the buffer object 41, the component type of the buffer view object 42, the number of data items stored in the buffer view object 42, and the type of data items stored in the buffer view object 42. This information is described for each buffer view object.

[0037] In the example of C in FIG. 7, information such as "bufferView":0, "byteOffset":0, "componentType":5126, "count":2106," and "type":"VEC3" is shown. As shown in FIG. 6, "bufferView":0 indicates that the identification information of the buffer view object 42 to which the accessor object 43 belongs is "0" (bufferView[0]). Furthermore, "byteOffset":0" indicates that the byte offset of the buffer view object 42 (bufferView[0]) is 0 bytes. Furthermore, "componentType":5126" indicates that the component type is FLOAT type (OpenGL macro constant). Furthermore, "count":2106" indicates that 2106 pieces of data are stored in the buffer view object 42 (bufferView[0]). Furthermore, "type":"VEC3" indicates that the data (type) stored in the buffer view object 42 (bufferView[0]) is a three-dimensional vector.

[0038] All accesses to data other than images are defined by reference to this accessor object 43 (by specifying the index of the accessor).

[0039] Next, we will explain how to specify a point cloud 3D object in such a glTF 2.0-compliant scene description (JSON format file). A point cloud is 3D content that represents a three-dimensional structure (a three-dimensional object) as a collection of numerous points. Point cloud data consists of position information (also called geometry) and attribute information (also called attributes) for each point. Attributes can contain any information. For example, attributes may include color information, reflectance information, normal information, etc. for each point. As such, point clouds have a relatively simple data structure, and by using a sufficient number of points, they can represent any three-dimensional structure with sufficient accuracy.

[0040] When a point cloud does not change over time (also called static), 3D objects are specified using the mesh.primitives object of glTF2.0. Figure 8 shows an example of the configuration of objects in a scene description when the point cloud is static. Figure 9 shows an example of how the scene description is written.

[0041] As shown in Figure 9, the mode of the primitives object is set to 0, indicating that the data is treated as points in a point cloud. As shown in Figures 8 and 9, the POSITION property of the attributes object in mesh.primitives specifies an accessor to a buffer that stores the position information of the points. Similarly, the COLOR property of the attributes object specifies an accessor to a buffer that stores the color information of the points. The buffer and bufferView may be one (the data may be stored in one file).

[0042] Next, we will explain how to extend such scene description objects. Each object in glTF 2.0 can store newly defined objects within an extension object. Figure 10 shows an example of how to specify a newly defined object (ExtensionExample). As shown in Figure 10, when using a newly defined extension, the extension object name (ExtensionExample in the example of Figure 10) is written in "extensionUsed" and "extensionRequired." This indicates that this extension is an extension that will be used or an extension that is required to be loaded.

[0043] <Client Processing> MPEG-I Scene Description is a standard for controlling the playback of 6DoF visuals using scene descriptions compliant with glTF 2.0. Here, "visual" refers to visual information such as images (information transmitted visually), and 6DoF visuals refer to visuals that correspond to the six degrees of freedom (6DoF) of movement (so-called free viewpoints) of the viewer (receiver of visual information). In other words, MPEG-I Scene Description controls the playback of 6DoF visuals (i.e., the playback of visual scenes) using scene descriptions that represent visual scenes. Here, "visual scene" refers to a visual scene. A visual scene is composed of visual objects arranged in a region (e.g., three-dimensional space). A visual object is an object that exists within a region and is composed of visuals. In other words, a playback device using MPEG-I Scene Description plays back 6DoF visuals and reconstructs the visual scene indicated by the scene description.

[0044] This MPEG-I scene description represents a visual scene composed of high-resolution visual objects. In this specification, a scene description conforming to the MPEG-I scene description may also be referred to as an MPEG-I scene description.

[0045] Next, we will explain how a client device processes this MPEG-I scene description. The client device acquires the scene description, acquires 3D object data based on the scene description, and generates a display image using the scene description and 3D object data.

[0046] As described in Non-Patent Document 2, in a client device, a presentation engine, a media access function, and the like perform processing. For example, as shown in FIG. 11 , a presentation engine 51 of a client device 50 acquires the initial value of a scene description and information for updating the scene description (hereinafter also referred to as update information), and generates a scene description for the processing time. The presentation engine 51 then analyzes the scene description and identifies the media (video, audio, etc.) to be played. The presentation engine 51 then requests a media access function 52 to acquire the media via a media access API (Application Program Interface). The presentation engine 51 also sets up pipeline processing, specifies buffers, and the like.

[0047] The media access function 52 acquires various media data requested by the presentation engine 51 from the cloud, local storage, etc. The media access function 52 supplies the acquired various media data (encoded data) to a pipeline 53.

[0048] The pipeline 53 decodes the various media data (encoded data) supplied thereto through pipeline processing, and supplies the decoded results to a buffer 54. The buffer 54 holds the various media data supplied thereto.

[0049] The presentation engine 51 performs rendering and the like using various media data stored in the buffer 54 .

[0050] <Application of Timed Media> In recent years, the application of timed media as 3D object content by extending glTF2.0 in MPEG-I Scene Description has been studied, as shown in, for example, Non-Patent Document 2. Timed media is media data that changes along the time axis, such as a moving image in a two-dimensional image.

[0051] glTF was only applicable to still image data as media data (3D object content). In other words, glTF did not support moving image media data. To move a 3D object, animation (a method of switching between still images along a time axis) was applied.

[0052] MPEG-I Scene Description will use glTF 2.0, use JSON format files as scene descriptions, and furthermore, it is being considered to extend glTF so that it can handle timed media (e.g., video data) as media data. For example, the following extensions will be made to handle timed media:

[0053] Fig. 12 is a diagram illustrating an extension for handling timed media. In the example of Fig. 12, an MPEG media object (MPEG_media) is a glTF extension, and is an object that specifies attributes of MPEG media such as video data, for example, uri, track, startTime, etc.

[0054] 12, an MPEG texture video object (MPEG_texture_video) is provided as an extension object (extension) of the texture object (texture). The MPEG texture video object stores information about an accessor corresponding to a buffer object to be accessed. In other words, the MPEG texture video object is an object that specifies the index of an accessor corresponding to a buffer in which the texture media specified by the MPEG media object (MPEG_media) is decoded and stored.

[0055] 13 is a diagram showing an example of the description of an MPEG media object (MPEG_media) and an MPEG texture video object (MPEG_texture_video) in a scene description to explain extensions for handling timed media. In the example of Fig. 13, the MPEG texture video object (MPEG_texture_video) is set as an extension object (extensions) of the texture object (texture) in the second line from the top, as shown below. The accessor index ("2" in this example) is then specified as the value of that MPEG video texture object.

[0056] "texture":[{"sampler":0, "source":1, "extensions":{"MPEG_texture_video ":"accessor":2}}],

[0057] 13, an MPEG media object (MPEG_media) is set as a glTF extension object (extensions) on lines 7 to 16 from the top, as shown below: The value of the MPEG media object stores various information about the MPEG media object, such as the encoding and URI of the MPEG media object.

[0058] "MPEG_media":{ "media":[ {"name":"source_1", "startTime":9.0, "loop":"true", "controls":"false", "alternatives":[{"mimeType":"video / mp4;codecs=\"avc1.42E01E\"", "uri":"video1.mp4", "tracks":[{"track":""#track_ID=1"}]}]} ]}

[0059] Furthermore, each frame of data is decoded and stored sequentially in a buffer. However, since the position of each frame fluctuates, a mechanism is provided in the scene description to store this fluctuating information so that a renderer can read the data. For example, as shown in FIG. 12 , an MPEG buffer circular object (MPEG_buffer_circular) is provided as an extension object of a buffer object. The MPEG buffer circular object stores information for dynamically storing data in the buffer object. For example, information such as access information to MPEG media and information indicating the number of frames is stored in the MPEG buffer circular object.

[0060] 12, an MPEG accessor-timed object (MPEG_accessor_timed) is provided as an extension object (extension) of the accessor object (accessor). In this case, since the media data is video, the buffer view object (bufferView) referenced in the time direction may change (its position may fluctuate). Therefore, information indicating the buffer view object referenced is stored in this MPEG accessor-timed object. For example, the MPEG accessor-timed object stores information indicating a reference to a buffer view object (bufferView) in which a timed accessor information header (timedAccessor information header) is written. The timed accessor information header is header information that stores, for example, dynamically changing accessor objects and information within the buffer view object.

[0061] 14 is a diagram showing an example of the description of an MPEG buffer circular object (MPEG_buffer_circular) and an MPEG accessor timed object (MPEG_accessor_timed) in a scene description to explain extensions for handling timed media. In the example of Fig. 14, the MPEG accessor timed object (MPEG_accessor_timed) is set as an extension object (extensions) of the accessor object (accessors) in the fifth line from the top, as shown below. Then, parameters such as the index of the buffer view object ("1" in this example), the recommended update rate (suggestedUpdateRate), and immutable information (immutable) and their values ​​are specified as the value of the MPEG accessor timed object.

[0062] "MPEG_accessor_timed":{"bufferView":1, "suggestedUpdateRate":25.0, "immutable":1,"}

[0063] 14, an MPEG buffer circular object (MPEG_buffer_circular) is set as an extension object (extensions) of a buffer object (buffer) on the 13th line from the top, as shown below: Then, parameters such as a buffer frame count (count) and access information to MPEG media (media) and their values ​​are specified as values ​​of the MPEG buffer circular object.

[0064] "MPEG_buffer_circular":{"count":5, "media":0}

[0065] Fig. 15 is a diagram for explaining extensions for handling timed media, showing examples of the relationship between an MPEG accessor timed object, an MPEG buffer circular object, an accessor object, a buffer view object, and a buffer object.

[0066] As described above, the MPEG buffer circular object of the buffer object stores information necessary for storing time-varying data in the buffer area indicated by the buffer object, such as the buffer frame count (count), access information to the MPEG media (media), etc.

[0067] As described above, the MPEG accessor-timed object of the accessor object stores information about the buffer view object it references, such as the index of the buffer view object (bufferView), the recommended update rate (suggestedUpdateRate), immutable information (immutable), etc. The MPEG accessor-timed object also stores information about the buffer view object in which the timed accessor information header it references is stored. The timed accessor information header can store a timestamp delta (timestamp_delta), update data for the accessor object, update data for the buffer view object, etc.

[0068] <Client Processing When Using MPEG_texture_video> A scene description is spatial layout information for arranging one or more 3D objects in 3D space. The contents of this scene description can be updated along the time axis. In other words, the layout of 3D objects can be updated over time. This section describes the client processing performed on the client device in this case.

[0069] Fig. 16 shows an example of the main configuration of a client device related to client processing, and Fig. 17 is a flowchart showing an example of the flow of the client processing. As shown in Fig. 16, the client device has a presentation engine (hereinafter also referred to as PE) 51, a media access function (hereinafter also referred to as MAF) 52, a pipeline 53, and a buffer 54. The presentation engine (PE) 51 has a glTF analysis unit 63 and a rendering processing unit 64.

[0070] A presentation engine (PE) 51 causes a media access function 52 to acquire media, acquires the data via a buffer 54, and performs display-related processing. Specifically, the processing is performed, for example, in the following manner.

[0071] When client processing begins, the glTF analysis unit 63 of the presentation engine (PE) 51 begins PE processing as shown in the example of Figure 17, and in step S21, obtains the SD(glTF) file 62, which is a scene description file, and parses the scene description.

[0072] In step S22, the glTF parser 63 checks the media associated with the 3D object (texture), the buffer in which the media will be stored after processing, and the accessor. In step S23, the glTF parser 63 notifies the media access function 52 of this information as a file acquisition request.

[0073] The media access function (MAF) 52 starts MAF processing as in the example of Fig. 17, and receives a notification of this in step S11. In step S12, the media access function 52 receives media (3D object file (mp4)) based on the notification.

[0074] In step S13, the media access function 52 decodes the acquired media (3D object file (mp4)). In step S14, the media access function 52 stores the media data obtained by the decoded decoding in the buffer 54 based on a notification from the presentation engine (PE 51).

[0075] In step S24, the rendering processing unit 64 of the presentation engine 51 reads (acquires) the data at an appropriate timing from the buffer 54. In step S25, the rendering processing unit 64 performs rendering using the acquired data to generate an image for display.

[0076] The media access function 52 repeats the processes of steps S13 and S14 to execute these processes for each time (each frame). The rendering processing unit 64 of the presentation engine 51 repeats the processes of steps S24 and S25 to execute these processes for each time (each frame). When the processes for all frames have been completed, the media access function 52 ends the MAF process, and the presentation engine 51 ends the PE process. In other words, the client process ends.

[0077] <Support for Interactivity in MPEG-I Scene Description> In recent years, progress has been made in standardizing compression and transmission technologies for haptic media, which compress haptic information, in addition to audio and video media, which are components of 2D video content and 3DoF (Degree of Freedom) / 6DoF video content, as described in Non-Patent Document 4. In parallel with this standardization of compression and transmission technologies for haptic media, Non-Patent Document 3 discloses technical extensions for MPEG-I Scene Description AMD2 to handle haptic media in a 3D space and technical extensions for interaction.

[0078] Haptic media can be categorized into two models: a synchronous model in which it plays (vibrates) in sync with audio or video media, and an interaction model (e.g., events such as touching, moving, or bumping) between the viewer and the 3D video object. In the interaction model, when a trigger condition is met due to an event occurring in the scene, an action corresponding to that event is executed. For example, when an object collision is triggered, the haptic device outputs vibrations corresponding to the collision event.

[0079] Non-Patent Document 3 defines interactive media processing. Interactive processing is an interaction-type process in which media processing (action) is executed when a certain execution condition is met as a trigger. A playback method that applies such interactive processing is also called interaction playback. When a trigger is provided to a node, an action is returned. In other words, a trigger is an event that triggers some form of interactivity, and an action is interaction-type feedback in response to that trigger. For example, when a collision between 3D objects is detected in a (virtual) three-dimensional space, feedback (an interaction-type action) such as animation, audio, or haptics (e.g., vibration) is returned.

[0080] In MPEG-I scene descriptions, information about such triggers (execution conditions) is defined as trigger information, and the action to be executed when the condition is met is defined as action information. In other words, trigger information defines the conditions that trigger an interaction, such as "contact" or "proximity." Action information defines the scene changes that occur as a result of the trigger, such as "movement" or "transformation" of a visual object. Behavior information defines interactive processing by associating (pairing) trigger information with action information.

[0081] In Non-Patent Document 3, interactivity is supported at the scene level and the node level through the definition of two extended functions, MPEG_scene_interactivity and MPEG_node_interactivity.

[0082] In MPEG_scene_interactivity, the semantics of behavior, trigger, and action are defined as shown in FIG.

[0083] A behavior defines what type of interactivity is allowed at runtime for a dedicated virtual object corresponding to a glTF node. A behavior has the ability to associate one or more triggers with one or more actions. A trigger defines the execution conditions that must be met before an action is executed. In other words, a trigger indicates the execution conditions for an interactive process. An action defines how that operation affects the scene. In other words, an action indicates the processing content of the interactive process. By associating such triggers with actions, a behavior expresses interactive processing (what processing is executed under what conditions).

[0084] The trigger types (i.e., execution condition types) are defined as shown in FIG. 19. The collision type (TRIGGER_COLLISION = 0) is a trigger type that activates an action when objects in a scene collide with each other (Collision Trigger). The proximity type (TRIGGER_PROXIMITY) is a trigger type that activates an action based on the distance between the virtual scene and the avatar (Proximity Trigger). The user input type (TRIGGER_USER_INPUT) is a trigger type that activates an action based on user interaction such as hand gestures (User_Input_Trigger). The visibility type (TRIGGER_VISIBILITY) is a trigger type that activates an action based on the viewing frustum (angle of view) (Visibility Trigger).

[0085] Additionally, the action types (i.e., processing content types) shown in FIG. 20 are defined. The activate type (ACTION_ACTIVATE = 0) is an action type activated by an application on a node. The transform type (ACTION_TRANSFORM) is an action type that applies a transformation matrix to a node (i.e., deforms the node). The block type (ACTION_BLOCK) is an action type that prevents a node from being deformed. The animation type (ACTION_ANIMATION) is an action type that plays an animation. The media type (ACTION_MEDIA) is an action type that plays media. The manipulate type (ACTION_MANIPULATE) is an action type that performs an operation that causes physical changes to a node. The set material type (ACTION_SET_MATERIAL) is an action type that sets a material for a node. The haptic type (ACTION_HAPTIC) is an action type that sets haptics (tactile sensations) for a node. The set avatar type (ACTION_SET_AVATAR) is an action type for setting an avatar for a node.

[0086] In a behavior, parameters such as those shown in FIG. 21 can be set. For example, the parameter "trggers" defines the index of the trigger in the trigger array that is considered in this behavior. The parameter "actions" defines the index of the action in the action array that is considered in this behavior. The parameter "triggersCombinationControl" defines the set of logical operations to apply to each trigger in the trigger array. The parameter "TriggerActivationControl" indicates when a trigger combination is activated to launch an action. The parameter "actionsControl" defines how the defined action is executed. The parameter "interruptAction" defines the index of the action in the action array that is executed if the behavior is still running and has not been defined in a newly received scene update. The parameter "priority" defines the priority associated with the behavior. When multiple behaviors affect the same node at the same time, the behavior with the highest priority value is processed. Behaviors with lower priority are not processed. For behaviors with the same priority, applications must apply their own criteria.

[0087] Furthermore, the processing model for activating a trigger is specified, for example, as shown in the flowchart of FIG. 22. When the trigger activation process starts, a trigger to be processed is selected in step S41, and it is determined in step S42 whether it has been expanded at the node level. If it is determined that it has been expanded at the node level, then in step S43, the conditions are checked using the scene and node parameters. When the process of step S43 is completed, the process proceeds to step S45. Also, if it is determined in step S42 that it has not been expanded at the node level, then the conditions are checked using the scene parameters in step S44. When the process of step S44 is completed, the process proceeds to step S45.

[0088] In step S45, it is determined whether the condition is satisfied. If it is determined that the condition is satisfied, the trigger is activated in step S46. That is, if the status satisfies the triggersActivationControl value, the action corresponding to the trigger is activated. When the processing of step S46 ends, the processing proceeds to step S47. On the other hand, if it is determined that the condition is not satisfied in step S45, the processing of step S46 is skipped and the processing proceeds to step S47.

[0089] In step S47, it is determined whether all triggers have been processed, and if it is determined that an unprocessed trigger exists, the process returns to step S41. In this way, the processes from step S41 to step S47 are executed for each trigger, and if it is determined in step S47 that all triggers have been processed, the trigger activation process ends.

[0090] To support haptic media in the scene description standard, two gLTF extensions are defined: "MPEG_haptic" and "MPEG_haptic_material," as shown in Figure 23. "MPEG_haptic" contains an array that defines all haptic objects in the glTF file-level extension. "MPEG_haptic_material" contains an array that defines all texture-based haptic data in the glTF file-level extension. Additionally, the mesh-level extension contains a single reference to the same array in the glTF file-level extension.

[0091] Non-Patent Document 4 discloses a standard that defines a haptic stream for compressing and transmitting a haptic signal (wav) and a description of the haptic signal (ivs, ahap). For example, the codec architecture shown in FIG. 24 is disclosed.

[0092] <Delay in Interaction Playback> In such interaction playback, it may take time for media to be output for various reasons. For example, there is the processing time required to convert the prepared media into output data (data for vibrating the output device) in accordance with the output device. In this specification, such output data is also referred to as "action data." There is also the transmission time required to transmit the generated action data to the output device. Furthermore, to reduce power consumption, the output device may be turned off (stopped or in a sleep state) when not in use and activated only when necessary. In such cases, before outputting the action data, the output device must be activated and ready to output, which requires time. In this specification, this time is also referred to as the "device startup time." Performing these processes after an event actually occurs (after a trigger condition is satisfied) could result in an unacceptably long delay in the output of the action data due to the reasons described above. For example, there is a risk that the time between the occurrence of a collision and the start of vibration of the vibration device could be unacceptably long.

[0093] To suppress such an increase in delay and play back interactions with low delay, it is possible to predict the occurrence of an event and perform processing related to the output of action data based on that prediction. However, even with such a method, the prediction may still be wrong, which could result in a decrease in the quality of the interaction playback or an unnecessary increase in the processing load.

[0094] For example, if an event is predicted, it is conceivable to start outputting action data before the event actually occurs, according to a known delay time. However, if the prediction is incorrect (if there is a mismatch between the predicted behavior and the actual behavior), an erroneous action presentation (incorrect action data output) may occur, such as a haptic presentation being provided even when there is no contact with a virtual object, which may reduce the quality of the interaction playback.

[0095] Another possible method is to launch an output device before an event actually occurs when the event is predicted. However, if the prediction is incorrect, there is a risk of erroneous action presentation. Even if the output of action data is stopped when the prediction is incorrect, the output device may be launched unnecessarily, which may unnecessarily increase the processing load. In the first place, controlling the launch of an output device at an appropriate timing based on such a prediction is difficult. For example, if the output device is launched too early relative to the event occurrence, there is a risk of unnecessarily increasing power consumption. In other words, there is a risk of unnecessarily increasing the processing load. Conversely, if the output device is launched too late relative to the event occurrence, there is a risk of a large delay in the output of action data relative to the event occurrence, which may degrade the quality of the interaction playback.

[0096] In general, the earlier an event occurrence is predicted (i.e., the longer the period between the predicted event occurrence and the actual event occurrence), the earlier the pre-processing (such as outputting action data and starting up an output device) can be started, thereby suppressing an increase in delay time. However, since the prediction is made further into the future, the prediction accuracy may decrease and the prediction may be incorrect. Therefore, as described above, there is a risk of incorrect action presentation or unnecessary processing, which may reduce the quality of the interaction playback or increase unnecessary processing load. Conversely, the later an event occurrence is predicted (i.e., the shorter the period between the predicted event occurrence and the actual event occurrence), the closer the future will be predicted, thereby improving the prediction accuracy. However, since the start timing of the pre-processing is delayed, there is a risk of an increase in delay time. Therefore, as described above, there is a risk of an increase in the delay in the output timing of action data and a decrease in the quality of the interaction playback.

[0097] The tolerance for delays and erroneous presentations in interaction playback, as well as the tolerance for processing load, may depend on the content (situation or situation). For example, there may be situations or situations where erroneous presentations of actions or long delays are acceptable, while there may be situations or situations where these are completely unacceptable. There may also be situations or situations where an increased processing load is acceptable, while there may be situations or situations where a lower processing load is required. In other words, the preferred method of interaction playback (especially in terms of prediction accuracy and control of delay amount) may depend on the content (situation or situation). Therefore, it is desirable to have interaction playback behave as closely as possible to the intentions of the content provider (e.g., author). That is, there has been a need for a playback method that ensures that the amount of delay, frequency of erroneous presentations, processing load, etc. in interaction playback are as closely as possible to the intentions of the content provider (e.g., author).

[0098] However, with conventional methods, such interaction playback is performed solely by the device, making it difficult to reflect the intentions of the content provider (e.g., author) in the interaction playback. As a result, there is a risk that the interaction playback may behave differently from the intentions of the content provider (e.g., author). For example, there is a risk of incorrect presentation or delay of actions that are different from the intentions of the content provider (e.g., author). Therefore, as described above, conventional methods may result in a decrease in the quality of interaction playback or an unnecessary increase in processing load.

[0099] <3. Low-latency interaction playback based on the author's intentions> <Method 1> Therefore, low-latency processing information is stored in the scene description (Method 1), as shown in the top row of the table in Fig. 25. For example, as shown in Fig. 26, the behavior information defined in Non-Patent Document 3 is extended to define low-latency processing information (LowLatencyProcessing).

[0100] In this specification, a playback method that applies interactive processing, i.e., a method of playing media (executing an action) triggered by the occurrence of an event (the satisfaction of a trigger condition), is also referred to as "interactive playback." In contrast, a method of playing media in synchronization with the time axis is also referred to as "synchronous playback." Media that is played synchronously is also referred to as "synchronous media," and media that is played interactively is also referred to as "asynchronous media."

[0101] The types of synchronous media and asynchronous media may be any. For example, they may be video media, audio media, haptic media, or other types of media. The types of synchronous media and asynchronous media may be the same or different. Furthermore, the number of media constituting the synchronous media and asynchronous media may be any number, and may be the same or different. Furthermore, the number of types of media constituting the synchronous media and asynchronous media may be any number, and may be the same or different.

[0102] Note that an "action" refers to the behavior of an output device when outputting interactively played asynchronous media, and "action data" refers to the data output by the output device, i.e., the data for outputting the asynchronous media. For example, if the asynchronous media is haptic media, the action data may be vibration data that vibrates the haptic device, which is the output device.

[0103] In this specification, the time from the occurrence of an event to the output of action data during interaction playback processing delay is also referred to as the "processing delay." As described above, this processing delay occurs for various reasons, such as the processing time required to convert asynchronous media into action data, the transmission time required to transmit action data to an output device, and the startup time of the output device. The length of this processing delay is also referred to as the "amount of processing delay." The longest time allowed for this amount of processing delay is also referred to as the "maximum delay time." Interaction playback that outputs action data within this maximum delay time (i.e., interaction playback that is performed so that the amount of processing delay is equal to or less than the maximum delay time) is also referred to as "low-latency interaction playback."

[0104] The low-latency processing information described above is information used to control low-latency interaction playback. This low-latency processing information is stored (described) in the scene description. In other words, the media provider provides this low-latency processing information, and the media player performs low-latency interaction playback based on the low-latency processing information.

[0105] For example, the first information processing device is provided with a scene description generation unit that generates a scene description representing a scene of a 3D space in which a 3D object is placed, and stores low-latency processing information in the scene description that includes a maximum processing delay and controls low-latency interaction playback.

[0106] In addition, the first information processing method executed by the first information processing device generates a scene description representing a scene in 3D space in which a 3D object is placed, and stores low-latency processing information in the scene description that includes a maximum processing delay and controls low-latency interaction playback.

[0107] The first program also causes the first information processing device to execute a process of generating a scene description representing a scene in 3D space in which the 3D object is placed, and storing low-latency processing information, which includes a maximum processing delay and controls low-latency interaction playback, in the scene description.

[0108] The second information processing device may further include a scene processing unit that performs low-latency interaction playback based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, the low-latency processing information being stored in a scene description that describes a scene in a 3D space in which a 3D object is placed. For example, the scene processing unit may select an action that can be performed with low-latency interaction playback, and perform low-latency interaction playback for the selected action.

[0109] In addition, the second information processing method executed by the second information processing device performs low-latency interaction playback based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, and that is stored in a scene description that represents a scene in 3D space in which the 3D object is placed.

[0110] In addition, the second program causes the second information processing device to execute processing for performing low-latency interaction playback based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, which is stored in a scene description that represents a scene in 3D space in which the 3D object is placed.

[0111] In these terms, low-latency interaction playback refers to a process in which an interaction in 3D space is triggered by an interaction and an action corresponding to the interaction is executed with less latency than the maximum processing latency, and the maximum processing latency refers to the maximum allowable time from when an interaction occurs until the action data of the action is output on an output device.

[0112] In this way, the low-latency processing information can include the intentions of the content author. For example, this low-latency processing information can be set (by the author, etc.) according to the content of the content. In other words, the content provider (author, etc.) can control low-latency interaction playback. In other words, the intentions of the content provider (author, etc.) can be reflected in low-latency interaction playback. In other words, low-latency interaction playback can be achieved in accordance with the intentions of the content provider (author, etc.), and at least one of a reduction in the quality of interaction playback and unnecessary high-load processing can be suppressed.

[0113] <Low-Latency Processing Information> The low-latency processing information may include any information. For example, as shown in FIG. 27 , the parameter “Maximum Processing Delay” may be included. This “Maximum Processing Delay” is a parameter that specifies the maximum delay time. In other words, the “Maximum Processing Delay” is information that indicates the maximum allowable time from when an interaction occurs until the action data of the action is output on an output device. For example, a content provider (such as an author) may set this “Maximum Processing Delay” based on the content of the content. For example, in a first information processing device, a scene description generation unit may store this “Maximum Processing Delay” in the scene description as low-latency processing information. Furthermore, in a second information processing device, a scene processing unit may perform low-latency interaction playback using the “Maximum Processing Delay.” For example, the scene processing unit may select an action that can be played back with low-latency interaction based on the “Maximum Processing Delay” and perform low-latency interaction playback for the selected action. In this way, by providing this "maximum processing delay" as low-latency processing information, the playback device can perform low-latency interaction playback in accordance with the intentions of the content provider (author, etc.). In other words, it is possible to prevent situations where interaction playback is performed with an amount of processing delay that is not intended by the content provider (author, etc.), thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0114] 27 , the low-latency processing information may include a parameter “Processing Policy.” This “Processing Policy” is a parameter (information) indicating a processing policy (processing behavior) for low-latency interaction playback. For example, this “Processing Policy” may include information indicating whether to prioritize action presentation or suppression of erroneous action presentation in low-latency interaction playback. For example, in a first information processing device, a scene description generation unit may store this “Processing Policy” in the scene description as low-latency processing information. Furthermore, in a second information processing device, a scene processing unit may perform interaction playback in accordance with this “Processing Policy.” In this way, by providing this “Processing Policy” as low-latency processing information, the playback device can perform low-latency interaction playback with a processing policy that is in line with the intention of the content provider (e.g., author), such as whether to prioritize action presentation or suppression of erroneous action presentation. In other words, for example, it is possible to prevent situations in which low-latency interaction playback behaves in a way that is not intended by the content provider (author, etc.), thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0115] As shown in FIG. 27 , the low-latency processing information may also include a parameter “Processing Priority.” This “Processing Priority” is information indicating the priority of processing for actions to be played back in low-latency interaction. For example, the “Processing Priority” may indicate the priority (the order in which processing execution is prioritized) for each action included in the behavior information. For example, in a first information processing device, a scene description generation unit may store this “Processing Priority” in the scene description as low-latency processing information. Furthermore, in a second information processing device, a scene processing unit may preferentially process actions with higher priorities indicated by the “Processing Priority.” In this way, by providing this “Processing Priority” as low-latency processing information, the playback device can execute actions in a priority order consistent with the intent of the content provider (e.g., author). In other words, for example, it is possible to prevent a situation in which an unintended action is executed preferentially over an action intended by the content provider (e.g., author), thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0116] Also, as shown in FIG. 27 , the low-latency processing information may include a parameter “low-latency processing flag (LowLatency Processing).” This “low-latency processing flag” is flag information indicating whether low-latency interaction playback is necessary. For example, a playback device may perform low-latency interaction playback when the “low-latency processing flag” is true, and may not perform low-latency interaction playback when the “low-latency processing flag” is false. For example, in a first information processing device, a scene description generation unit may store this “low-latency processing flag” in the scene description as low-latency processing information. Also, in a second information processing device, a scene processing unit may perform interaction playback based on the “low-latency processing flag.” For example, the scene processing unit may perform low-latency interaction playback when the “low-latency processing flag” is true. Also, the scene processing unit may prohibit low-latency interaction playback when the “low-latency processing flag” is false. In this way, by providing this "low-latency processing flag" as low-latency processing information, the playback device can control whether or not to execute low-latency interaction playback according to the intention of the content provider (author, etc.). In other words, it is possible to prevent situations where low-latency interaction playback is performed even though the content provider (author, etc.) does not intend it to, and to prevent at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0117] 27 , the low-latency processing information may also include a parameter “application actions (Actions to be processed).” The “application actions” are information indicating actions (among the actions included in the behavior information) for which low-latency interaction playback is performed. In other words, by using the “application actions,” low-latency interaction playback can be applied to only some of the actions included in the behavior information. For example, in a first information processing device, a scene description generation unit may store the “application actions” in the scene description as low-latency processing information. In a second information processing device, a scene processing unit may select and execute actions that are capable of low-latency interaction playback from among the actions specified by the “application actions.” In this way, by providing the “application actions” as low-latency processing information, the playback device can select actions for low-latency interaction playback in accordance with the intention of the content provider (e.g., author). In other words, it is possible to prevent situations where actions unintended by the content provider (author, etc.) are played back in low-latency interaction playback, thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0118] Instead of this "application action", the necessity of low latency processing may be indicated by flag information in the action structure.

[0119] <Supplementary Note> Low-latency interaction playback may be performed by any method that can keep the processing delay below the maximum delay time. For example, the method described in Non-Patent Document 4 may be used, or an action may be executed before an event occurs based on the distance between objects and a known processing delay, or a future event may be predicted and an action may be executed before the event actually occurs based on the prediction.

[0120] The low-delay processing information may be stored anywhere in the scene description, for example, it may be defined as an extension of the trigger information or the action information. The low-delay processing information may also be stored in a location other than the behavior information. For example, the low-delay processing information may be defined as an extension of MPEG_media.

[0121] <Method 1-1> When making such predictions (detecting the occurrence of future events), the predictions may be made by a playback device that acquires and processes media. For example, as shown in Fig. 28, the playback device (scene processing unit) may predict the probability of an event occurring (probability prediction), determine a trigger condition based on the occurrence probability (trigger determination), and, if the trigger condition is met, convert the media into action data and supply (output) it to an output device (action execution). The output device (device processing unit) may then be activated, acquire the action data, and output it (action output).

[0122] When applying such prediction in the case where Method 1 is applied, predictive trigger activation control information may be stored in the scene description (Method 1-1), as shown in the second row from the top of the table in FIG. 25 . For example, as shown in FIG. 29 , the behavior information defined in Non-Patent Document 3 may be expanded to further define predictive trigger activation control information (PredictiveTriggerActivationControl). That is, low-latency processing information and predictive trigger activation control information may be defined as behavior information. This predictive trigger activation control information is information used by a processing unit that generates action data (e.g., the scene processing unit in FIG. 28 ) to control low-latency interaction playback in which interactions are predicted. That is, a media provider may provide this predictive trigger activation control information, and a media player may perform low-latency interaction playback as shown in FIG. 28 based on the predictive trigger activation control information. For example, in a first information processing device, a scene description generation unit may store the predictive trigger activation control information in the scene description. In the second information processing device, a scene processing unit may perform low-latency interaction playback based on predicted trigger activation control information stored in the scene description. The low-latency interaction playback may include generating action data, deriving an interaction occurrence probability, determining a trigger condition, and providing the action data to an output device when the trigger condition is met.

[0123] The predictive trigger activation control information is generated by the content provider. Therefore, the predictive trigger activation control information can reflect the intentions of the content author. For example, this predictive trigger activation control information may be set (by the author, etc.) according to the content. In other words, by applying such predictive trigger activation control information, the content provider (author, etc.) can control low-latency interaction playback. Therefore, as described above, even when the playback device predicts event occurrence, low-latency interaction playback according to the intentions of the content provider (author, etc.) can be achieved, thereby suppressing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0124] <Predictive Trigger Activation Control Information> The predictive trigger activation control information may include any information. For example, as shown in FIG. 30 , it may include a parameter “predictive trigger activation control flag (PredictiveTriggerActivationControl).” This “predictive trigger activation control flag” is flag information indicating whether a threshold determination of the probability of an event occurrence is used as a trigger condition. That is, the “predictive trigger activation control information” may include a “predictive trigger activation control flag” indicating whether a processing unit that generates action data derives the probability of an interaction as a prediction of the interaction and performs threshold determination on the probability of the interaction. For example, if the “predictive trigger activation control flag” is true, the playback device may use a threshold determination of the probability of an event occurrence as a trigger condition (i.e., may predict the event occurrence). Alternatively, if the “predictive trigger activation control flag” is false, the playback device may not use a threshold determination of the probability of an event occurrence as a trigger condition (i.e., may prohibit the prediction of the event occurrence). For example, in the first information processing device, the scene description generation unit may store the “predictive trigger activation control flag” in the scene description as predictive trigger activation control information. In the second information processing apparatus, the scene processing unit may derive the probability of an interaction occurring when the "prediction trigger activation control flag" is true.

[0125] In this way, by providing this "prediction trigger activation control flag" as prediction trigger activation control information, the playback device can control the prediction of event occurrence according to the intention of the content provider (author, etc.). In other words, it is possible to prevent situations where an interaction is predicted even though the content provider (author, etc.) did not intend it to, and to prevent at least one of a decrease in the quality of interaction playback and unnecessary high-load processing.

[0126] For example, as shown in FIG. 30 , a parameter "Predictive Trigger Activation Control Threshold" may be included. This "Predictive Trigger Activation Control Threshold" is information indicating a threshold applied to threshold determination of the probability of an event occurrence. In other words, the predictive trigger activation control information may include a "Predictive Trigger Activation Control Threshold" indicating a threshold used to determine the threshold of the probability of an interaction occurrence. For example, in a first information processing device, a scene description generation unit may store this "Predictive Trigger Activation Control Threshold" in the scene description as predictive trigger activation control information. Furthermore, in a second information processing device, a scene processing unit may use the "Predictive Trigger Activation Control Threshold" to determine the threshold of the probability of an interaction occurrence as a trigger condition determination.

[0127] In this way, by providing this "prediction trigger activation control threshold" as the prediction trigger activation control information, the playback device can control the prediction of event occurrence according to the intention of the content provider (author, etc.). In other words, it is possible to prevent a situation in which it is determined that a trigger condition has been met at a stage not intended by the content provider (author, etc.), thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0128] <Prediction Determination Processing Method> Note that any method may be used by the scene processing unit to predict an interaction (event occurrence). For example, as described above, the scene processing unit may predict the probability of an interaction occurring, determine the occurrence probability based on a threshold using a "prediction trigger activation control threshold" (a threshold set by the content author, etc.), and determine whether or not the trigger condition is met based on the determination result. For example, if the occurrence probability is equal to or greater than a threshold, it may be determined that an event will occur.

[0129] For example, if the maximum delay time is 0.6 seconds and the threshold is set to 90%, the scene processing unit predicts the probability of an interaction occurring, as shown in FIG. 31 , and executes an action at time t, where the prediction is calculated and the probability after t+0.6 seconds is calculated to be 90%. That is, action data is supplied to the device processing unit. The device processing unit acquires the action data and outputs it at a timing that matches the startup characteristics of the device. For example, if the device startup time is 0.4 seconds, the action data is held for 0.2 seconds (after adjusting the timing) and then output.

[0130] <Supplementary Information> Note that the prediction determination result may be an erroneous determination, for example, when an event is predicted to occur but does not actually occur. In such a case, scene processing may be continued assuming that the trigger condition is met (an event has occurred), or may be continued assuming that the trigger condition is not met (an event has not occurred). Furthermore, when such an erroneous determination occurs, flag information specifying whether the trigger condition is met or not may be included in the prediction trigger activation control information in order to continue scene processing.

[0131] Furthermore, since prediction algorithms are implementation-dependent, there is a possibility that predictions may vary depending on the device. Therefore, the prediction trigger activation control information may include "recommended parameters" that specify recommended parameters to be applied to deriving probability data. Any parameters may be specified as these "recommended parameters." For example, the "recommended parameters" may include speed, distance, acceleration, or viewing direction. Furthermore, the "recommended parameters" may specify any number of parameters. For example, in a first information processing device, a scene description generation unit may store these "recommended parameters" in the scene description as prediction trigger activation control information. Furthermore, in a second information processing device, a scene processing unit may derive the probability of an interaction occurring using the "recommended parameters."

[0132] Although the low-latency processing information and the predictive trigger activation control information have been described above as different pieces of information, they may be combined into the same processing information. The predictive trigger activation control information may be stored anywhere in the scene description, and may be defined, for example, as an extension of the trigger information or as an extension of the action information. The predictive trigger activation control information may also be stored in a location other than the behavior information. For example, the predictive trigger activation control information may be defined as an extension of MPEG_media.

[0133] <Method 1-2> Alternatively, as shown in FIG. 32 , the output device (device processing unit) may perform trigger determination. That is, the playback device (scene processing unit) may predict the probability of an event occurring (probability prediction), convert media into action data, and supply the occurrence probability and action data to the output device (action pre-execution). The output device (device processing unit) may then be started, acquire the occurrence probability and action data, determine whether a trigger condition exists based on the occurrence probability (trigger determination), and output the action data if the trigger condition is met (action output).

[0134] When applying such prediction in the case where Method 1 is applied, predictive trigger activation offload control information may be stored in the scene description, as shown in the bottom row of the table in FIG. 25 (Method 1-2). For example, as shown in FIG. 33, the behavior information defined in Non-Patent Document 3 may be expanded to further define predictive trigger activation offload control information (PredictiveTriggerActivationOffloadControl). That is, low-latency processing information and predictive trigger activation offload control information may be defined as behavior information. This predictive trigger activation offload control information is information used to control low-latency interaction playback in which an output device predicts an interaction. That is, a media provider may provide this predictive trigger activation offload control information, and a media player may perform low-latency interaction playback as shown in FIG. 32 based on the predictive trigger activation offload control information. For example, in a first information processing device, a scene description generation unit may store predictive trigger activation offload control information in a scene description. In addition, in the second information processing device, a scene processing unit may perform low-latency interaction playback based on predictive trigger activation offload control information stored in the scene description. Furthermore, as the low-latency interaction playback, the scene processing unit may generate action data, derive an interaction occurrence probability, and provide the occurrence probability data, the action data, and the predictive trigger activation offload control information to an output device.

[0135] The predictive trigger activation offload control information is generated by the content provider. Therefore, the predictive trigger activation offload control information can reflect the intent of the content author. For example, this predictive trigger activation offload control information may be set (by the author, etc.) according to the content. That is, by applying such predictive trigger activation offload control information, the content provider (author, etc.) can control low-latency interaction playback. Therefore, as described above, even when the output device predicts event occurrence, low-latency interaction playback according to the intent of the content provider (author, etc.) can be achieved, thereby suppressing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0136] <Predictive Trigger Activation Offload Control Information> The predictive trigger activation offload control information may include any information. For example, as shown in FIG. 34 , a parameter “Predictive Trigger Activation Offload Control Flag (PredictiveTriggerActivationOffload)” may be included. This “Predictive Trigger Activation Offload Control Flag” is flag information indicating whether or not the output device is to execute a trigger determination (determine a threshold value for the occurrence probability). That is, the predictive trigger activation offload control information may include a “Predictive Trigger Activation Offload Control Flag” indicating whether or not the output device is to perform a threshold value determination on the occurrence probability of an interaction as a predicted interaction. For example, if the “Predictive Trigger Activation Offload Control Flag” is true, the output device may use a threshold value determination on the occurrence probability of an event as a trigger condition (i.e., may predict the occurrence of an event). Alternatively, if the “Predictive Trigger Activation Offload Control Flag” is false, the output device may not use a threshold value determination on the occurrence probability of an event as a trigger condition (i.e., may prohibit the prediction of an event occurrence). For example, in the first information processing device, the scene description generation unit may store this "prediction trigger activation offload control flag" in the scene description as prediction trigger activation offload control information. Also, in the second information processing device, the scene processing unit may derive the probability of an interaction occurring when the "prediction trigger activation offload control flag" is true.

[0137] In this way, by providing this "prediction-trigger-activation offload control flag" as the prediction-trigger-activation offload control information, the output device can control the prediction of event occurrence according to the intention of the content provider (author, etc.). In other words, it is possible to prevent situations where an interaction is predicted even though the content provider (author, etc.) did not intend it to, and to prevent at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0138] For example, as shown in FIG. 34 , a parameter "Predictive Trigger Activation Offload Control Threshold (PredictiveTriggerActivationOffloadThreshold)" may be included. This "Predictive Trigger Activation Offload Control Threshold" is information indicating a threshold applied to threshold determination of the probability of an event occurrence. In other words, the predictive trigger activation offload control information may include a "Predictive Trigger Activation Offload Control Threshold" indicating a threshold used to determine the threshold of the probability of an interaction occurrence. For example, in a first information processing device, a scene description generation unit may store this "Predictive Trigger Activation Offload Control Threshold" in the scene description as the predictive trigger activation offload control information. Furthermore, in a second information processing device, a scene processing unit may supply the "Predictive Trigger Activation Offload Control Threshold" to an output device to cause the output device to perform threshold determination of the probability of an interaction occurrence. In other words, a device processing unit of the output device may use the "Predictive Trigger Activation Offload Control Threshold" to determine the threshold of the probability of an interaction occurrence.

[0139] In this way, by providing this "prediction-trigger-activation offload control threshold" as the prediction-trigger-activation offload control information, the playback device can control the prediction of event occurrence according to the intention of the content provider (author, etc.). In other words, it is possible to prevent, for example, a situation in which it is determined that a trigger condition has been met at a stage not intended by the content provider (author, etc.), thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0140] The predictive trigger activation offload control information may also include format information indicating a format of interaction occurrence probability data to be provided to the output device. For example, in a first information processing device, a scene description generation unit may store this format information in the scene description as predictive trigger activation offload control information. In a second information processing device, a scene processing unit may derive an occurrence probability based on the format information and generate occurrence probability data. The scene processing unit may also supply the format information as predictive trigger activation offload control information to the output device, causing it to perform a threshold determination of the interaction occurrence probability. In other words, the device processing unit of the output device may analyze the occurrence probability data based on the format information and perform a threshold determination.

[0141] In this way, by providing this format information as the prediction-trigger-activation offload control information, the playback device generates occurrence probability data according to the intentions of the content provider (e.g., author), and the output device can correctly interpret the occurrence probability data and make threshold determinations. In other words, it is possible to prevent situations such as the occurrence probability data being updated in a format different from the intentions of the content provider (e.g., author) or the occurrence probability data being misinterpreted, thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0142] For example, as shown in FIG. 34 , the format information (i.e., predictive trigger activation offload control information) may include a “probability data frame rate (ProbabilityDataFrameRate).” This “probability data frame rate” is information indicating the update frequency of the occurrence probability data. For example, in a first information processing device, a scene description generation unit may store this “probability data frame rate” in the scene description as format information (i.e., predictive trigger activation offload control information). Furthermore, in a second information processing device, a scene processing unit may derive the occurrence probability at a frequency corresponding to the “probability data frame rate” (i.e., update the occurrence probability data at a frequency corresponding to the “probability data frame rate”). Furthermore, the scene processing unit may provide the “probability data frame rate” as format information (i.e., predictive trigger activation offload control information) to an output device to perform a threshold determination of the interaction occurrence probability. In other words, a device processing unit of the output device may analyze the occurrence probability data based on the “probability data frame rate.”

[0143] In this way, by providing this "probability data frame rate" as format information (i.e., prediction trigger activation offload control information), the playback device generates occurrence probability data according to the intentions of the content provider (e.g., author), and the output device can correctly interpret the occurrence probability data and make threshold determinations. In other words, it is possible to prevent situations such as the occurrence probability data being updated at a frequency different from the intentions of the content provider (e.g., author) or being misinterpreted, thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0144] For example, as shown in FIG. 34 , the format information (i.e., predictive trigger activation offload control information) may include an “intra-frame probability data duration (ProbabilityDataDurationInFrame).” This “intra-frame probability data duration” is information indicating the time length corresponding to one frame of occurrence probability data. For example, in a first information processing device, a scene description generation unit may store this “intra-frame probability data duration” in the scene description as format information (i.e., predictive trigger activation offload control information). Furthermore, in a second information processing device, a scene processing unit may generate occurrence probability data in a format corresponding to the “intra-frame probability data duration.” Furthermore, the scene processing unit may provide the “intra-frame probability data duration” as format information (i.e., predictive trigger activation offload control information) to an output device, causing the output device to perform a threshold determination of the occurrence probability of an interaction. In other words, a device processing unit of the output device may analyze the occurrence probability data based on the “intra-frame probability data duration.”

[0145] In this way, by providing this "intra-frame probability data duration" as format information (i.e., prediction trigger activation offload control information), the playback device generates occurrence probability data according to the intentions of the content provider (e.g., author), and the output device can correctly interpret the occurrence probability data and make threshold determinations. In other words, it is possible to prevent situations such as generating occurrence probability data in a format different from the intentions of the content provider (e.g., author) or misinterpreting the occurrence probability data, thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0146] For example, as shown in FIG. 34 , the format information (i.e., predictive trigger activation offload control information) may include an “intra-frame probability data rate (ProbabilityDataRateInFrame).” This “intra-frame probability data rate” is information indicating the time length corresponding to one occurrence probability data. For example, in a first information processing device, a scene description generation unit may store this “intra-frame probability data rate” in the scene description as format information (i.e., predictive trigger activation offload control information). Furthermore, in a second information processing device, a scene processing unit may generate occurrence probability data in a format corresponding to the “intra-frame probability data rate.” Furthermore, the scene processing unit may supply the “intra-frame probability data rate” as format information (i.e., predictive trigger activation offload control information) to an output device, causing the output device to perform a threshold determination of the occurrence probability of an interaction. In other words, a device processing unit of the output device may analyze the occurrence probability data based on the “intra-frame probability data rate.”

[0147] In this way, by providing this "intra-frame probability data rate" as format information (i.e., prediction trigger activation offload control information), the playback device generates occurrence probability data according to the intentions of the content provider (e.g., author), and the output device can correctly interpret the occurrence probability data and make threshold determinations. In other words, it is possible to prevent situations such as generating occurrence probability data in a format different from the intentions of the content provider (e.g., author) or misinterpreting the occurrence probability data, thereby preventing at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0148] <Prediction Determination Processing Method> Note that the scene processing unit may use any method for predicting an interaction (event occurrence). For example, as described above, the scene processing unit may predict the probability of an interaction and convert media into action data. For example, if the maximum delay time is 0.6 seconds and the threshold is set to 90%, the scene processing unit may predict the probability of an interaction and generate occurrence probability data in a format such as that shown in FIG. 35. The scene processing unit may supply the occurrence probability data to an output device (device processing unit) along with the action data and predicted trigger activation offload control information.

[0149] The output device (device processing unit) may compare the occurrence probability data with a threshold based on the predicted trigger activation offload control information, and output action data if it is determined that the trigger condition is met. For example, if the device delay time is 0.4 seconds and the threshold is set to 90%, the device processing unit may output the action data received at time t when the prediction is calculated and the probability after t + 0.4 seconds is calculated to be 90%.

[0150] <Supplementary Note> Note that the above-described method 1-2 can be used in combination with the above-described method 1-1. That is, the low-latency processing information, the predictive trigger activation control information, and the predictive trigger activation offload control information may be stored in the scene description.

[0151] Furthermore, when applying method 1-2, as in the case of applying method 1-1, if the prediction determination result is an erroneous determination, scene processing may be continued assuming that the trigger condition is met (an event has occurred), or may be continued assuming that the trigger condition is not met (an event has not occurred). Furthermore, when such an erroneous determination occurs, in order to continue scene processing, flag information specifying whether or not the trigger condition is met may be included in the prediction trigger activation offload control information.

[0152] Furthermore, when applying Method 1-2, as in the case of applying Method 1-1, the predictive trigger activation offload control information may include "recommended parameters" that specify parameters recommended as parameters to be applied to derive occurrence probability data. Any parameters may be specified as these "recommended parameters." For example, the "recommended parameters" may include speed, distance, acceleration, or viewing direction. The "recommended parameters" may specify any number of parameters. For example, in the first information processing device, the scene description generation unit may store these "recommended parameters" in the scene description as predictive trigger activation offload control information. In the second information processing device, the scene processing unit may derive the occurrence probability of an interaction using the "recommended parameters."

[0153] Although the low-latency processing information and the predictive trigger activation offload control information have been described above as different pieces of information, they may be combined into the same processing information. Similarly, the predictive trigger activation control information and the predictive trigger activation offload control information may be combined into the same processing information. That is, the low-latency processing information, the predictive trigger activation control information, and the predictive trigger activation offload control information may be combined into the same processing information. Furthermore, the predictive trigger activation offload control information may be stored anywhere in the scene description, and may be defined, for example, as an extension of the trigger information or as an extension of the action information. Furthermore, the predictive trigger activation offload control information may be stored outside of the behavior information. For example, the predictive trigger activation offload control information may be defined as an extension of MPEG_media.

[0154] 4. First Embodiment File Generation Device The present technology described above may be applied to any device. FIG. 36 is a block diagram showing an example of the configuration of a file generation device, which is one aspect of an information processing device to which the present technology is applied. The file generation device 300 shown in FIG. 36 is a device that generates and outputs distribution media for content that represents a scene using synchronous media and asynchronous media. For example, the file generation device 300 may acquire synchronous media, asynchronous media, scene information, etc. that constitute a scene and use them to generate a scene description. The file generation device 300 may also encode the synchronous media, asynchronous media, and scene description, respectively, to create files, and output the files as distribution media.

[0155] Note that Fig. 36 shows the main processing units, data flows, etc., and does not necessarily include everything shown in Fig. 36. In other words, in file generation device 300, there may be processing units that are not shown as blocks in Fig. 36, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 36.

[0156] As shown in FIG. 36, the file generation device 300 includes an SD generation unit 311 , an encoding unit 312 , a file generation unit 313 , a storage unit 314 , and a supply unit 315 .

[0157] The SD generation unit 311 executes processing related to the generation of a scene description. For example, the SD generation unit 311 may acquire synchronous media, asynchronous media, and scene information supplied from outside the file generation device 300. The SD generation unit 311 may then generate a scene description using the acquired synchronous media, asynchronous media, and scene information. The SD generation unit 311 may then supply the generated scene description to the encoding unit 312 (SD encoding unit 321).

[0158] Synchronous media and asynchronous media are media played in a content scene. Synchronous media is media that is played synchronously, while asynchronous media is media that is played interactively. These may be any type of media, such as video, audio, or haptics. Scene information is information used to generate a scene description and is composed of information for expressing a scene. Scene information is generated, for example, by a content author. Scene information may include any information related to a scene. For example, it may include information indicating the placement and movement of objects. It may also include information related to the playback of synchronous media and asynchronous media. Information necessary for generating low-latency processing information, predictive trigger activation control information, and predictive trigger activation offload control information is supplied as this scene information. The scene description may be in any format, and may be generated in accordance with a standard described in, for example, Non-Patent Document 3.

[0159] The encoding unit 312 executes processing related to encoding. As shown in Fig. 36 , the encoding unit 312 includes an SD encoding unit 321, a synchronous media encoding unit 322, and an asynchronous media encoding unit 323.

[0160] The SD encoding unit 321 executes processing related to encoding of a scene description. For example, the SD encoding unit 321 may acquire a scene description supplied from the SD generation unit 311. The SD encoding unit 321 may encode the scene description and generate a bitstream. Any encoding method may be used. The SD encoding unit 321 may supply the generated bitstream (encoded data) of the scene description to the file generation unit 313 (SD file generation unit 331).

[0161] The synchronized media encoder 322 executes processing related to encoding of synchronized media. For example, the synchronized media encoder 322 may acquire synchronized media supplied from outside the file generation device 300. This synchronized media is the same as that acquired by the SD generation unit 311 (i.e., the synchronized media used to generate the scene description). The synchronized media encoder 322 may encode the synchronized media to generate a bitstream. Any encoding method may be used. The synchronized media encoder 322 may supply the generated synchronized media bitstream (encoded data) to the file generation unit 313 (synchronized media file generation unit 332).

[0162] The asynchronous media encoder 323 performs processing related to encoding of asynchronous media. For example, the asynchronous media encoder 323 may acquire asynchronous media supplied from outside the file generation device 300. This asynchronous media is the same as that acquired by the SD generation unit 311 (i.e., the asynchronous media used to generate the scene description). The asynchronous media encoder 323 may encode the asynchronous media to generate a bitstream. Any encoding method may be used. The asynchronous media encoder 323 may supply the generated asynchronous media bitstream (encoded data) to the file generation unit 313 (asynchronous media file generation unit 333).

[0163] The file generation unit 313 executes processing related to file generation. As shown in FIG. 36 , the file generation unit 313 includes an SD file generation unit 331, a synchronous media file generation unit 332, and an asynchronous media file generation unit 333.

[0164] The SD file generation unit 331 executes processing related to the generation of a scene description file that is a file of a scene description. For example, the SD file generation unit 331 may acquire a bitstream of the scene description supplied from the SD encoding unit 321. The SD file generation unit 331 may generate a scene description file that stores the bitstream. The file format of this scene description file may be any format. The SD file generation unit 331 may supply the generated scene description file to the storage unit 314.

[0165] The synchronized media file generator 332 executes processing related to the generation of synchronized media files that are files of synchronized media. For example, the synchronized media file generator 332 may acquire a synchronized media bitstream supplied from the synchronized media encoder 322. The synchronized media file generator 332 may generate a synchronized media file that stores the bitstream. The synchronized media file may have any file format. The synchronized media file generator 332 may supply the generated synchronized media file to the storage unit 314.

[0166] The asynchronous media file generator 333 executes processing related to the generation of asynchronous media files that are files of asynchronous media. For example, the asynchronous media file generator 333 may acquire an asynchronous media bitstream supplied from the asynchronous media encoder 323. The asynchronous media file generator 333 may generate an asynchronous media file that stores the bitstream. The asynchronous media file may have any file format. The asynchronous media file generator 333 may supply the generated asynchronous media file to the storage unit 314.

[0167] The storage unit 314 has any storage medium and performs processing related to the storage of data using the storage medium. For example, the storage unit 314 may acquire and store a scene description file supplied from the SD file generation unit 331. The storage unit 314 may acquire and store a synchronized media file supplied from the synchronized media file generation unit 332. The storage unit 314 may acquire and store an asynchronous media file supplied from the asynchronous media file generation unit 333. The storage unit 314 may also read a scene description file stored in its own storage medium and supply it to the supply unit 315. The storage unit 314 may also read a synchronized media file stored in its own storage medium and supply it to the supply unit 315. The storage unit 314 may also read an asynchronous media file stored in its own storage medium and supply it to the supply unit 315.

[0168] The supply unit 315 has a communication unit that communicates with other devices and transmits and receives information, and executes processing related to the supply of files. For example, the supply unit 315 may read a scene description file from the storage unit 314 and supply it to other devices (e.g., a distribution server, a playback device, etc.). The supply unit 315 may read a synchronized media file from the storage unit 314 and supply it to other devices (e.g., a distribution server, a playback device, etc.). The supply unit 315 may read an asynchronous media file from the storage unit 314 and supply it to other devices (e.g., a distribution server, a playback device, etc.).

[0169] In the file generation device 300 configured as above, the present technology described above in <3. Low-latency interaction playback based on author's intentions> may be applied.

[0170] For example, the file generation device 300 may be used as a first information processing device, and the above-described method 1 may be applied. That is, the SD generation unit 311 may generate a scene description representing a scene in 3D space in which a 3D object is placed, and store low-latency processing information, including a maximum processing delay, for controlling low-latency interaction playback in the scene description. In other words, the SD generation unit 311 may also be referred to as a scene description generation unit. Here, "low-latency interaction playback" refers to a process in which interaction playback, triggered by the occurrence of an interaction in 3D space, executes an action corresponding to that interaction, with a delay shorter than the maximum processing delay. The parameter "maximum processing delay" indicates the maximum allowable time from the occurrence of the interaction to the output of the action data of the action on an output device.

[0171] The low-latency processing information may include a parameter "processing policy" that indicates a processing policy for low-latency interaction playback. For example, the parameter "processing policy" may include information that indicates whether to prioritize action presentation or suppression of erroneous action presentation in low-latency interaction playback. The low-latency processing information may also include a parameter "processing priority" that indicates the priority of processing for actions to be played back in low-latency interaction playback. The low-latency processing information may also include flag information "low-latency processing flag" that indicates whether low-latency interaction playback is necessary. The low-latency processing information may also include an applied action that indicates an action for which low-latency interaction playback is performed. The low-latency processing information may also include any combination of these pieces of information.

[0172] Alternatively, the file generation device 300 may be used as a first information processing device, and the above-described method 1-1 may be applied. That is, the SD generation unit 311 may store predictive trigger activation control information in the scene description. Here, the predictive trigger activation control information is information used by the processing unit generating action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation control information may include flag information, a "predictive trigger activation control flag." This "predictive trigger activation control flag" indicates whether the processing unit generating action data derives the probability of an interaction as a predicted interaction and performs threshold determination on the probability. The predictive trigger activation control information may also include a "predictive trigger activation control threshold" parameter indicating a threshold used to determine the threshold value of the interaction occurrence probability. The predictive trigger activation control information may also include a "recommended parameter" parameter indicating a parameter recommended as a parameter to be used for prediction. The predictive trigger activation control information may also include any combination of these pieces of information.

[0173] Alternatively, the file generation device 300 may be a first information processing device, and the above-described method 1-2 may be applied. That is, the SD generation unit 311 may store predictive trigger activation offload control information in the scene description. Here, the predictive trigger activation offload control information is information used by an output device that outputs action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation offload control information may also include flag information ("predictive trigger activation offload control flag") indicating whether the output device determines the probability of an interaction as a predicted interaction by comparing the probability with a threshold. The predictive trigger activation offload control information may also include a parameter ("predictive trigger activation offload control threshold") indicating a threshold used to determine the probability of an interaction. The predictive trigger activation offload control information may also include format information indicating the format of the interaction occurrence probability data provided to the output device. For example, the format information may include a parameter ("probability data frame rate") indicating the update frequency of the occurrence probability data. The format information may also include a parameter "intra-frame probability data duration" indicating the time length corresponding to one frame's worth of occurrence probability data. The format information may also include a parameter "intra-frame probability data rate" indicating the time length corresponding to one piece of occurrence probability data. The predictive trigger activation offload control information may also include a parameter "recommended parameter" indicating a parameter recommended as a parameter to be used in predicting an interaction. The predictive trigger activation offload control information may also include any combination of these pieces of information.

[0174] With this configuration, the file generation device 300 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the file generation device 300 can achieve low-latency interaction playback in accordance with the intentions of the content provider (such as the author), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0175] <Flow of file generation process 1> An example of the flow of file generation process executed by the file generation device 300 when applying method 1 of the present technology described above in <3. Low-latency interaction playback based on the author's intention> will be described with reference to the flowchart in FIG. 37 .

[0176] When the file generation process begins, in step S301, the SD generation unit 311 of the file generation device 300 generates a scene description using scene information, synchronous media, asynchronous media, etc., and stores low-latency processing information. This low-latency processing information includes at least a maximum processing delay. That is, the SD generation unit 311 generates a scene description representing a scene in 3D space in which 3D objects are placed, and stores low-latency processing information, including a maximum processing delay and controlling low-latency interaction playback, in the scene description. Here, "low-latency interaction playback" refers to interaction playback, in which an interaction in 3D space is triggered and an action corresponding to that interaction is executed with less latency than the maximum processing delay. The parameter "maximum processing delay" indicates the maximum allowable time from the occurrence of the interaction to the output of the action data of the action on an output device.

[0177] The low-latency processing information may include a parameter "processing policy" that indicates a processing policy for low-latency interaction playback. For example, the parameter "processing policy" may include information that indicates whether to prioritize action presentation or suppression of erroneous action presentation in low-latency interaction playback. The low-latency processing information may also include a parameter "processing priority" that indicates the priority of processing for actions to be played back in low-latency interaction playback. The low-latency processing information may also include flag information "low-latency processing flag" that indicates whether low-latency interaction playback is necessary. The low-latency processing information may also include an applied action that indicates an action for which low-latency interaction playback is performed. The low-latency processing information may also include any combination of these pieces of information.

[0178] In step S302, the SD encoding unit 321 encodes the scene description to generate a bitstream. In step S303, the SD file generation unit 331 generates a scene description file and stores the bitstream of the scene description. In step S304, the storage unit 314 stores the scene description file.

[0179] In step S305, the synchronized media encoding unit 322 encodes the synchronized media to generate a bitstream. In step S306, the synchronized media file generation unit 332 generates a synchronized media file and stores the synchronized media bitstream. In step S307, the storage unit 314 stores the synchronized media file.

[0180] In step S308, the asynchronous media encoder 323 encodes the asynchronous media to generate a bitstream. In step S309, the asynchronous media file generator 333 generates an asynchronous media file and stores the asynchronous media bitstream. In step S310, the storage unit 314 stores the asynchronous media file.

[0181] In step S311, the supply unit 315 reads the scene description file and supplies it to another device (for example, a distribution server, a playback device, etc.). The supply unit 315 also reads the synchronized media file and supplies it to another device (for example, a distribution server, a playback device, etc.). The supply unit 315 also reads the asynchronous media file and supplies it to another device (for example, a distribution server, a playback device, etc.). When the processing of step S311 ends, the file generation processing ends.

[0182] By performing each process in this manner, the file generation device 300 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. In other words, the file generation device 300 can achieve low-latency interaction playback in accordance with the intentions of the content provider (such as the author), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0183] <Flow of file generation process 2> An example of the flow of the file generation process executed by the file generation device 300 when applying method 1-1 of the present technology described above in <3. Low-latency interaction playback based on the author's intention> will be described with reference to the flowchart in FIG. 38 .

[0184] When the file generation process starts, in step S331, the SD generation unit 311 of the file generation device 300 generates a scene description using scene information, synchronous media, asynchronous media, etc., and stores low-latency processing information and predictive trigger activation control information. The predictive trigger activation control information is information used by the processing unit generating action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation control information may include flag information, such as a "predictive trigger activation control flag." This "predictive trigger activation control flag" indicates whether the processing unit generating action data derives the probability of an interaction as a predicted interaction and determines whether the probability of the interaction is to be evaluated against a threshold. The predictive trigger activation control information may also include a "predictive trigger activation control threshold" parameter indicating a threshold used to determine the probability of the interaction. The predictive trigger activation control information may also include a "recommended parameter" parameter indicating a recommended parameter to be used for the prediction. The predictive trigger activation control information may also include any combination of these pieces of information.

[0185] The processes of steps S332 to S341 are executed in the same manner as the processes of steps S302 to S311 in Fig. 37. When the process of step S341 ends, the file generation process ends.

[0186] By performing each process in this manner, the file generation device 300 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. In other words, the file generation device 300 can achieve low-latency interaction playback in accordance with the intentions of the content provider (such as the author), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0187] <Flow of file generation process 3> An example of the flow of file generation process executed by the file generation device 300 when applying method 1-2 of the present technology described above in <3. Low-latency interaction playback based on the author's intention> will be described with reference to the flowchart in FIG. 39 .

[0188] When the file generation process starts, in step S361, the SD generation unit 311 of the file generation device 300 generates a scene description using scene information, synchronous media, asynchronous media, etc., and stores low-latency processing information and predictive trigger activation offload control information. The predictive trigger activation offload control information is information used by an output device that outputs action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation offload control information may include a "predictive trigger activation offload control flag" indicating whether the output device determines the probability of an interaction as a predicted interaction by a threshold. The predictive trigger activation offload control information may also include a "predictive trigger activation offload control threshold" parameter indicating a threshold used to determine the probability of an interaction. The predictive trigger activation offload control information may also include format information indicating the format of the interaction occurrence probability data provided to the output device. For example, the format information may include a "probability data frame rate" parameter indicating the update frequency of the occurrence probability data. The format information may also include a parameter "intra-frame probability data duration" indicating the time length corresponding to one frame's worth of occurrence probability data. The format information may also include a parameter "intra-frame probability data rate" indicating the time length corresponding to one piece of occurrence probability data. The predictive trigger activation offload control information may also include a parameter "recommended parameter" indicating a parameter recommended as a parameter to be used in predicting an interaction. The predictive trigger activation offload control information may also include any combination of these pieces of information.

[0189] The processes of steps S362 to S371 are executed in the same manner as the processes of steps S302 to S311 in Fig. 37. When the process of step S371 ends, the file generation process ends.

[0190] By performing each process in this manner, the file generation device 300 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. In other words, the file generation device 300 can achieve low-latency interaction playback in accordance with the intentions of the content provider (such as the author), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0191] 5. Second Embodiment Distribution Server The present technology described above may be applied to any device. FIG. 40 is a block diagram showing an example of the configuration of a distribution server, which is one aspect of an information processing device to which the present technology is applied. The distribution server 400 shown in FIG. 40 is a device that executes processing related to content distribution. For example, the distribution server 400 may acquire, store, and manage files for content distribution (scene description files, synchronous media files, asynchronous media files, etc.) generated by the file generation device 300 or the like. Alternatively, the distribution server 400 may manage these files stored externally. The distribution server 400 may supply a scene description file requested by a client device (playback device) or the like. Furthermore, the distribution server 400 may supply synchronous media or asynchronous media requested based on the scene description to the client device (playback device).

[0192] Note that Fig. 40 shows the main processing units, data flows, etc., and does not necessarily show everything. In other words, in distribution server 00, there may be processing units that are not shown as blocks in Fig. 40, and there may be processing and data flows that are not shown as arrows, etc. in Fig. 40.

[0193] As shown in FIG. 40, the distribution server 400 includes an acquisition unit 411 , a storage unit 412 , and a distribution unit 413 .

[0194] The acquisition unit 411 executes processing related to the acquisition of content files. For example, the acquisition unit 411 may acquire a scene description file supplied from an external device such as the file generation device 300, and supply the file to the storage unit 412. The acquisition unit 411 may acquire a synchronous media file supplied from an external device such as the file generation device 300, and supply the file to the storage unit 412. The acquisition unit 411 may acquire an asynchronous media file supplied from an external device such as the file generation device 300, and supply the file to the storage unit 412.

[0195] The storage unit 412 has any storage medium and performs processing related to the storage of data using the storage medium. For example, the storage unit 412 may acquire and store a scene description file supplied from the acquisition unit 411. The storage unit 412 may acquire and store a synchronized media file supplied from the acquisition unit 411. The storage unit 412 may acquire and store an asynchronous media file supplied from the acquisition unit 411.

[0196] The distribution unit 413 includes a communication unit that communicates with other devices and transmits and receives information, and performs processing related to content distribution. As shown in FIG. 40 , the distribution unit 413 includes an SD distribution unit 421, a synchronized media distribution unit 422, and an asynchronous media distribution unit 423. The SD distribution unit 421 performs processing related to the distribution of scene description files. For example, the SD distribution unit 421 may read and provide a scene description file in response to a request from an external device such as a client device. The synchronized media distribution unit 422 performs processing related to the distribution of synchronized media files. For example, the synchronized media distribution unit 422 may read and provide a synchronized media file requested from an external device such as a client device based on a scene description. The asynchronous media distribution unit 423 performs processing related to the distribution of asynchronous media files. For example, the asynchronous media distribution unit 423 may read and provide asynchronous media files requested from an external device such as a client device based on a scene description.

[0197] In the distribution server 400 configured as above, the present technology described above in <3. Low-latency interaction playback based on author's intentions> may be applied.

[0198] For example, the distribution server 400 may be a third information processing device, and the above-described method 1 may be applied. That is, the SD distribution unit 421 may provide, in response to a request, a scene description that describes a scene in 3D space in which a 3D object is placed and that stores low-latency processing information. In other words, the SD distribution unit 421 may also be considered a scene description supply unit. Here, "low-latency interaction playback" refers to a process in which interaction playback, triggered by the occurrence of an interaction in 3D space, executes an action corresponding to that interaction with less latency than the maximum processing latency. The parameter "maximum processing latency" indicates the maximum allowable time from the occurrence of the interaction to the output of the action data of the action on an output device.

[0199] The low-latency processing information may include a parameter "processing policy" that indicates a processing policy for low-latency interaction playback. For example, the parameter "processing policy" may include information that indicates whether to prioritize action presentation or suppression of erroneous action presentation in low-latency interaction playback. The low-latency processing information may also include a parameter "processing priority" that indicates the priority of processing for actions to be played back in low-latency interaction playback. The low-latency processing information may also include flag information "low-latency processing flag" that indicates whether low-latency interaction playback is necessary. The low-latency processing information may also include an applied action that indicates an action for which low-latency interaction playback is performed. The low-latency processing information may also include any combination of these pieces of information.

[0200] Alternatively, the distribution server 400 may be a third information processing device, and the above-described method 1-1 may be applied. That is, the SD distribution unit 421 may further provide a scene description including predictive trigger activation control information. Here, the predictive trigger activation control information is information used by the processing unit generating action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation control information may include flag information, such as a "predictive trigger activation control flag." The "predictive trigger activation control flag" indicates whether the processing unit generating action data derives the probability of an interaction as a predicted interaction and performs threshold determination on the probability. The predictive trigger activation control information may also include a "predictive trigger activation control threshold" parameter indicating a threshold used to determine the threshold value of the interaction occurrence probability. The predictive trigger activation control information may also include a "recommended parameter" parameter indicating a parameter recommended as a parameter to be used for the prediction. The predictive trigger activation control information may also include any combination of these pieces of information.

[0201] Alternatively, the distribution server 400 may be a third information processing device, and the above-described method 1-2 may be applied. That is, the SD distribution unit 421 may further provide a scene description including predictive trigger activation offload control information. Here, the predictive trigger activation offload control information is information used by an output device that outputs action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation offload control information may include flag information (predictive trigger activation offload control flag) indicating whether the output device determines the probability of an interaction as a predicted interaction by comparing the probability with a threshold. The predictive trigger activation offload control information may also include a parameter (predictive trigger activation offload control threshold) indicating a threshold used to determine the probability of an interaction. The predictive trigger activation offload control information may also include format information indicating the format of the interaction occurrence probability data provided to the output device. For example, the format information may include a parameter (probability data frame rate) indicating the update frequency of the occurrence probability data. The format information may also include a parameter "intra-frame probability data duration" indicating the time length corresponding to one frame's worth of occurrence probability data. The format information may also include a parameter "intra-frame probability data rate" indicating the time length corresponding to one piece of occurrence probability data. The predictive trigger activation offload control information may also include a parameter "recommended parameter" indicating a parameter recommended as a parameter to be used in predicting an interaction. The predictive trigger activation offload control information may also include any combination of these pieces of information.

[0202] With this configuration, the distribution server 400 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the distribution server 400 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0203] <Flow of Content Acquisition Process> An example of the flow of content acquisition process executed by distribution server 400 will be described with reference to the flowchart of FIG.

[0204] When the content acquisition process starts, in step S401, the acquisition unit 411 of the distribution server 400 acquires a content distribution file uploaded from an external device (e.g., the file generation device 300, etc.). For example, the acquisition unit 411 may acquire a scene description file as the distribution file. Alternatively, the acquisition unit 411 may acquire a synchronized media file as the distribution file. Alternatively, the acquisition unit 411 may acquire an asynchronous media file as the distribution file.

[0205] In step S402, the storage unit 412 stores the acquired distribution files. For example, the storage unit 412 may store the acquired scene description files. The storage unit 412 may also store the acquired synchronized media files. The storage unit 412 may also store the acquired asynchronous media files.

[0206] In step S403, the acquisition unit 411 determines whether or not to end the content acquisition. If it is determined not to end, the process returns to step S401. If it is determined to end the content acquisition, the content acquisition process ends.

[0207] <Flow of Distribution Processing> An example of the flow of distribution processing executed by distribution server 400 will be described with reference to the flowchart of FIG.

[0208] When the distribution process starts, the SD distribution unit 421 determines in step S421 whether a scene description file has been requested from an external device (for example, a client device (playback device)). If it is determined that a scene description file has been requested, the process proceeds to step S422.

[0209] In step S422, the SD distribution unit 421 reads the requested scene description file and supplies it to the requestor (e.g., a client device). When the processing of step S422 ends, the processing proceeds to step S423. Also, if it is determined in step S421 that a scene description file has not been requested, the processing proceeds to step S423.

[0210] In step S423, the synchronized media distribution unit 422 determines whether a synchronized media file has been requested from outside (for example, a client device (playback device) or the like). If it is determined that a synchronized media file has been requested, the process proceeds to step S424. In step S424, the synchronized media distribution unit 422 reads the requested synchronized media file and provides it to the requestor (for example, a client device). When the process of step S424 ends, the process proceeds to step S425. Also, if it is determined in step S423 that a synchronized media file has not been requested, the process proceeds to step S425.

[0211] In step S425, the asynchronous media delivery unit 423 determines whether an asynchronous media file has been requested from outside (for example, a client device (playback device) or the like). If it is determined that an asynchronous media file has been requested, the process proceeds to step S426. In step S426, the asynchronous media delivery unit 423 reads the requested asynchronous media file and supplies it to the requestor (for example, a client device). When the process of step S426 ends, the process proceeds to step S427. Also, if it is determined in step S425 that an asynchronous media file has not been requested, the process proceeds to step S427.

[0212] In step S427, the distribution unit 413 determines whether or not to end the distribution. If it is determined not to end, the process returns to step S421. If it is determined to end the distribution, the distribution process ends.

[0213] In this distribution process, the method 1 of the present technology described above in <3. Low-latency interaction playback based on the author's intentions> may be applied. That is, in step S422, the SD distribution unit 421 may provide, in response to a request, a scene description that describes a 3D space scene in which a 3D object is placed and stores low-latency processing information. Here, the "low-latency processing information" includes a maximum processing delay and controls low-latency interaction playback. Furthermore, "low-latency interaction playback" refers to a process in which interaction playback, triggered by the occurrence of an interaction in 3D space, executes an action corresponding to that interaction, with a delay shorter than the maximum processing delay. Furthermore, the parameter "maximum processing delay" indicates the maximum allowable time from the occurrence of the interaction to the output of the action data of the action on the output device.

[0214] The low-latency processing information may include a parameter "processing policy" that indicates a processing policy for low-latency interaction playback. For example, the parameter "processing policy" may include information that indicates whether to prioritize action presentation or suppression of erroneous action presentation in low-latency interaction playback. The low-latency processing information may also include a parameter "processing priority" that indicates the priority of processing for actions to be played back in low-latency interaction playback. The low-latency processing information may also include flag information "low-latency processing flag" that indicates whether low-latency interaction playback is necessary. The low-latency processing information may also include an applied action that indicates an action for which low-latency interaction playback is performed. The low-latency processing information may also include any combination of these pieces of information.

[0215] The SD distribution unit 421 may also provide a scene description that stores low-latency processing information and predictive trigger activation control information. The predictive trigger activation control information is information used by the processing unit that generates action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation control information may also include flag information, such as a "predictive trigger activation control flag." The "predictive trigger activation control flag" indicates whether the processing unit that generates action data derives the probability of an interaction as a predicted interaction and determines whether the probability of the interaction is to be evaluated against a threshold. The predictive trigger activation control information may also include a "predictive trigger activation control threshold" parameter indicating a threshold used to determine the probability of the interaction. The predictive trigger activation control information may also include a "recommended parameter" parameter indicating a recommended parameter to be used for the prediction. The predictive trigger activation control information may also include any combination of these pieces of information.

[0216] The SD distribution unit 421 may also provide a scene description that stores low-latency processing information and predictive trigger activation offload control information. The predictive trigger activation offload control information is information used by an output device that outputs action data to control low-latency interaction playback in which an interaction is predicted. The predictive trigger activation offload control information may include flag information (a "predictive trigger activation offload control flag") indicating whether the output device determines whether the probability of an interaction is predicted and uses a threshold value to determine the probability of the interaction. The predictive trigger activation offload control information may also include a "predictive trigger activation offload control threshold" parameter indicating a threshold value used to determine the threshold value for the probability of the interaction. The predictive trigger activation offload control information may also include format information indicating a format for the interaction probability data provided to the output device. For example, the format information may include a "probability data frame rate" parameter indicating an update frequency for the probability data. The format information may also include a "intra-frame probability data duration" parameter indicating the duration corresponding to one frame of probability data. The format information may also include an "intra-frame probability data rate" parameter indicating the time length corresponding to one occurrence probability data. The predictive trigger activation offload control information may also include a "recommended parameter" parameter indicating a parameter recommended as a parameter to be used in predicting an interaction. The predictive trigger activation offload control information may also include any combination of these pieces of information.

[0217] By performing each process in this manner, the distribution server 400 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the distribution server 400 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0218] 6. Third Embodiment Client Device / Output Device The present technology described above may be applied to any device. FIG. 43 is a block diagram showing an example of the configuration of a client device and an output device, each of which is an aspect of an information processing device to which the present technology is applied. The client device 500 shown in FIG. 43 is a device that executes processing related to content playback. For example, the client device 500 may acquire a synchronous media file or an asynchronous media file from an external device (e.g., the file generation device 300 or the distribution server 400), decode and reconstruct the file, and output the file to the output device 600. In this case, the client device 500 may play such content based on a scene description. For example, the client device 500 may acquire a scene description file from an external device (e.g., the file generation device 300 or the distribution server 400), decode a bitstream stored in the scene description file, generate a scene description, and use the scene description to control content playback.

[0219] 43 is a device that executes processing related to the output of synchronous media or asynchronous media. For example, the output device 600 may acquire and output output data of synchronous media or asynchronous media supplied from the client device 500.

[0220] Note that Fig. 43 shows the main processing units, data flows, etc., and is not necessarily all that is shown in Fig. 43. In other words, in the client apparatus 500 or the output device 600, there may be processing units that are not shown as blocks in Fig. 43, or there may be processing or data flows that are not shown as arrows, etc. in Fig. 43.

[0221] As shown in FIG. 43, the client device 500 has a presentation engine (PE) 511 and a media access function (MAF) 512.

[0222] The presentation engine (PE) 511 executes processing related to playback control and scene reconstruction based on the scene description. As shown in Fig. 43 , the presentation engine (PE) 511 includes an SD acquisition unit 521, an SD decoding unit 522, a control unit 523, and a scene processing unit 524.

[0223] The SD acquisition unit 521 executes processing related to acquisition of a scene description file. For example, the SD acquisition unit 521 may request a scene description file from an external device (such as the file generation device 300 or the distribution server 400), acquire the requested scene description file supplied in response, and supply the file to the SD decoding unit 522.

[0224] The SD decoding unit 522 executes processing related to decoding of the scene description. For example, the SD decoding unit 522 may acquire a scene description file supplied from the SD acquisition unit 521. The SD decoding unit 522 may extract and decode a bitstream stored in the scene description file to generate (restore) a scene description. Note that this decoding method may be any method as long as it corresponds to the encoding method applied when encoding the scene description. The SD decoding unit 522 may supply the generated scene description to the control unit 523.

[0225] The control unit 523 executes processing related to content playback control. For example, the control unit 523 may acquire a scene description supplied from the SD decoding unit 522. Based on the scene description, the control unit 523 may control the media access function (MAF) 512 to control the acquisition of synchronous media and asynchronous media. Furthermore, based on the scene description, the control unit 523 may control the scene processing unit 524 to control processing related to scene reconstruction. For example, the control unit 523 may control the reconstruction of synchronous media and asynchronous media.

[0226] The scene processing unit 524 executes processing related to scene reconstruction under the control of the control unit 523. As shown in FIG. 43, the scene processing unit 524 has a synchronous media scene processing unit 531 and an asynchronous media scene processing unit 532.

[0227] The synchronized media scene processing unit 531 performs processing related to the reconstruction of synchronized media under the control of the control unit 523. For example, the synchronized media scene processing unit 531 may acquire synchronized media generated (restored) by the media access function (MAF) 512 via a buffer (not shown). The synchronized media scene processing unit 531 may render the synchronized media and generate output data. The synchronized media scene processing unit 531 may supply the generated output data of the synchronized media to the output device 600 (synchronized media device processing unit 621).

[0228] The asynchronous media scene processing unit 532 performs processing related to the reconstruction of asynchronous media under the control of the control unit 523. For example, the asynchronous media scene processing unit 532 may acquire the asynchronous media generated (restored) by the media access function (MAF) 512 via a buffer (not shown). The asynchronous media scene processing unit 532 may render the asynchronous media and generate output data. The asynchronous media scene processing unit 532 may supply the generated output data of the asynchronous media to the output device 600 (asynchronous media device processing unit 622).

[0229] The media access function (MAF) 512 executes processes related to acquisition and decoding of synchronized media files and asynchronous media files under the control of the presentation engine 511 (control unit 523). As shown in Fig. 43, the media access function (MAF) 512 has a media acquisition unit 541 and a media decoding unit 542.

[0230] The media acquisition unit 541 executes processing related to acquisition of media files under the control of the presentation engine 511 (control unit 523). As shown in FIG. 43, the media acquisition unit 541 includes a synchronous media acquisition unit 551 and an asynchronous media acquisition unit 552.

[0231] The synchronized media acquisition unit 551 executes processing related to acquisition of synchronized media files under the control of the presentation engine 511 (control unit 523). For example, the synchronized media acquisition unit 551 may request a synchronized media file specified based on a scene description from an external device (such as the file generation device 300 or the distribution server 400) and acquire the requested synchronized media file provided in response. The synchronized media acquisition unit 551 may provide the acquired synchronized media file to the media decoding unit 542 (synchronized media decoding unit 561).

[0232] The asynchronous media acquisition unit 552 executes processing related to the acquisition of asynchronous media files under the control of the presentation engine 511 (control unit 523). For example, the asynchronous media acquisition unit 552 may request an asynchronous media file specified based on a scene description from an external device (such as the file generation device 300 or the distribution server 400) and acquire the requested asynchronous media file provided in response. The asynchronous media acquisition unit 552 may provide the acquired asynchronous media file to the media decoding unit 542 (asynchronous media decoding unit 562).

[0233] The media decoding unit 542 executes processing related to media decoding. As shown in FIG. 43, the media decoding unit 542 includes a synchronous media decoding unit 561 and an asynchronous media decoding unit 562.

[0234] The synchronized media decoding unit 561 executes processing related to the decoding of synchronized media. For example, the synchronized media decoding unit 561 may acquire a synchronized media file supplied from the synchronized media acquisition unit 551. The synchronized media decoding unit 561 may extract and decode a bitstream stored in the acquired synchronized media file to generate (restore) synchronized media. Note that this decoding method may be any method as long as it corresponds to the encoding method applied when encoding the synchronized media. The synchronized media decoding unit 561 may supply the generated (restored) synchronized media to the scene processing unit 524 (synchronized media scene processing unit 531) via a buffer (not shown).

[0235] The asynchronous media decoding unit 562 performs processing related to the decoding of asynchronous media. For example, the asynchronous media decoding unit 562 may acquire an asynchronous media file supplied from the asynchronous media acquisition unit 552. The asynchronous media decoding unit 562 may extract and decode a bitstream stored in the acquired asynchronous media file to generate (restore) asynchronous media. Note that this decoding method may be any method as long as it corresponds to the encoding method applied when encoding the asynchronous media. The asynchronous media decoding unit 562 may supply the generated (restored) asynchronous media to the scene processing unit 524 (asynchronous media scene processing unit 532) via a buffer (not shown).

[0236] 43, the output device 600 has a device processing unit 611. The device processing unit 611 performs processing related to media output. As shown in FIG. 43, the device processing unit 611 has a synchronous media device processing unit 621 and an asynchronous media device processing unit 622.

[0237] The synchronized media device processing unit 621 executes processing related to the output of synchronized media. For example, the synchronized media device processing unit 621 may acquire output data of synchronized media supplied from the scene processing unit 524 (synchronized media scene processing unit 531) of the client device 500. The synchronized media device processing unit 621 may activate the output function of the synchronized media of the output device 600 and output the acquired output data.

[0238] The asynchronous media device processing unit 622 performs processing related to the output of asynchronous media. For example, the asynchronous media device processing unit 622 may acquire asynchronous media output data (action data) supplied from the scene processing unit 524 (asynchronous media scene processing unit 532) of the client device 500. The asynchronous media device processing unit 622 may activate the asynchronous media output function of the output device 600 and output the acquired output data.

[0239] In the client device 500 and the output device 600 configured as described above, the present technology described above in <3. Low-latency interaction playback based on the author's intention> may be applied.

[0240] For example, the client device 500 may be a second information processing device, and the above-described method 1 may be applied. That is, the asynchronous media scene processing unit 532 may perform low-latency interaction playback based on low-latency processing information, including a maximum processing delay, that controls low-latency interaction playback and is stored in a scene description that describes a scene in 3D space in which a 3D object is placed. In other words, the asynchronous media scene processing unit 532 may also be referred to as a scene processing unit. Here, "low-latency interaction playback" refers to a process in which interaction playback, triggered by the occurrence of an interaction in 3D space, executes an action corresponding to the interaction, with a delay shorter than the maximum processing delay. Furthermore, the parameter "maximum processing delay" indicates the maximum allowable time from the occurrence of the interaction to the output of the action data of the action on an output device. For example, the asynchronous media scene processing unit 532 may select an action capable of low-latency interaction playback and perform low-latency interaction playback for the selected action.

[0241] The low-latency processing information may include a parameter "processing policy" indicating a processing policy for low-latency interaction playback. The asynchronous media scene processing unit 532 may then perform low-latency interaction playback in accordance with the processing policy. For example, the parameter "processing policy" may include information indicating whether to prioritize action presentation or suppression of erroneous action presentation in low-latency interaction playback. The low-latency processing information may also include a parameter "processing priority" indicating the processing priority for actions to be played back with low-latency interaction. The asynchronous media scene processing unit 532 may then prioritize processing of actions with higher priorities. The low-latency processing information may also include flag information "low-latency processing flag" indicating whether low-latency interaction playback is necessary. The asynchronous media scene processing unit 532 may then perform low-latency interaction playback if the low-latency processing flag is true. The low-latency processing information may also include an application action indicating an action for which low-latency interaction playback is to be performed. The asynchronous media scene processing unit 532 may then select an action to process from among the actions specified by the applied action. The low-latency processing information may also include any combination of these pieces of information.

[0242] Alternatively, the client device 500 may be used as a second information processing device, and the above-described method 1-1 may be applied. That is, the asynchronous media scene processing unit 532 may perform low-latency interaction playback based on predictive trigger activation control information stored in the scene description. Here, the predictive trigger activation control information is information used by a processing unit that generates action data to control low-latency interaction playback in which an interaction is predicted. For example, the asynchronous media scene processing unit 532 may perform the low-latency interaction playback by generating action data, deriving the probability of occurrence of a previous interaction, determining a trigger condition, and providing the action data to an output device when the trigger condition is met.

[0243] The predictive trigger activation control information may include flag information "predictive trigger activation control flag." This "predictive trigger activation control flag" indicates whether the processing unit that generates the action data derives the occurrence probability of an interaction as a prediction of the interaction and determines whether the occurrence probability is a threshold value. If the predictive trigger activation control flag is true, the asynchronous media scene processing unit 532 may derive the occurrence probability of the interaction. The predictive trigger activation control information may also include a parameter "predictive trigger activation control threshold" indicating a threshold value used to determine the threshold value of the occurrence probability of the interaction. The asynchronous media scene processing unit 532 may then use the predictive trigger activation control threshold value to determine the occurrence probability as a trigger condition. The predictive trigger activation control information may also include a parameter "recommended parameter" indicating a parameter recommended as a parameter to be used for prediction. The asynchronous media scene processing unit 532 may then use the recommended parameter to derive the occurrence probability. The predictive trigger activation control information may also include any combination of these pieces of information.

[0244] Alternatively, the client device 500 may be used as a second information processing device, and the above-described method 1-2 may be applied. That is, the asynchronous media scene processing unit 532 may perform low-latency interaction playback based on predictive trigger activation offload control information stored in the scene description. Here, the predictive trigger activation offload control information is information used by an output device that outputs action data to control low-latency interaction playback in which an interaction is predicted. For example, the asynchronous media scene processing unit 532 may perform low-latency interaction playback by generating action data, deriving an interaction occurrence probability, and providing the occurrence probability data, action data, and predictive trigger activation offload control information to the output device.

[0245] The predictive trigger activation offload control information may also include flag information "predictive trigger activation offload control flag" indicating whether the output device will predict an interaction and determine the probability of the interaction based on a threshold. If the predictive trigger activation offload control flag is true, the asynchronous media scene processing unit 532 may derive the probability of the interaction. The predictive trigger activation offload control information may also include a parameter "predictive trigger activation offload control threshold" indicating a threshold used to determine the threshold for the probability of the interaction. The predictive trigger activation offload control information may also include format information indicating a format of the interaction occurrence probability data provided to the output device. The asynchronous media scene processing unit 532 may then derive the interaction occurrence probability based on the format information. For example, the format information may include a parameter "probability data frame rate" indicating an update frequency for the occurrence probability data. The asynchronous media scene processing unit 532 may then derive the interaction occurrence probability at a frequency corresponding to the probability data frame rate. The format information may also include an "intra-frame probability data duration" parameter indicating a time length corresponding to one frame's worth of occurrence probability data. The asynchronous media scene processing unit 532 may then generate occurrence probability data in a format corresponding to the intra-frame probability data duration. The format information may also include an "intra-frame probability data rate" parameter indicating a time length corresponding to one occurrence probability data. The asynchronous media scene processing unit 532 may then generate occurrence probability data in a format corresponding to the intra-frame probability data rate. The predictive trigger activation offload control information may also include a "recommended parameter" parameter indicating a parameter recommended as a parameter to be used in predicting an interaction. The asynchronous media scene processing unit 532 may then use the recommended parameter to derive the occurrence probability of an interaction. The predictive trigger activation offload control information may also include any combination of these pieces of information.

[0246] With this configuration, the client device 500 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the client device 500 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0247] For example, the output device 600 may be a fourth information processing device, and the above-described method 1-2 may be applied. That is, the asynchronous media device processing unit 622 may acquire occurrence probability data indicating the probability of an interaction, action data indicating an action corresponding to the interaction, and predictive trigger activation offload control information, perform a threshold-based judgment on the occurrence probability data based on the predictive trigger activation offload control information, and output the action data when a trigger condition is met. In other words, the asynchronous media device processing unit 622 may also be considered a device processing unit. Here, the "predictive trigger activation offload control information" is information used to control low-latency interaction playback. Furthermore, "low-latency interaction playback" refers to interaction playback, in which an action corresponding to an interaction is executed in response to the occurrence of an interaction in 3D space, with a delay shorter than the maximum processing delay. Furthermore, the parameter "maximum processing delay" indicates the maximum allowable time from the occurrence of the interaction to the output of the action data of the action on the output device.

[0248] With this configuration, the output device 600 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the output device 600 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0249] <Flow of playback process> An example of the flow of playback process executed by the client device 500 when applying method 1 of the present technology described above in <3. Low-latency interaction playback based on the author's intention> will be described with reference to the flowchart in Fig. 44 .

[0250] When the playback process is started, the SD acquisition unit 521 acquires a scene description file from an external device such as the file generation device 300 or the distribution server 400 in step S501.

[0251] In step S502, the SD decoding unit 522 extracts a bitstream from the acquired scene description file, decodes the bitstream, and generates (restores) a scene description.

[0252] In step S503, the control unit 523 analyzes the generated scene description, and controls the media access function (MAF) 512 and the scene processing unit 524 based on the analysis result of the scene description.

[0253] In step S504, the synchronized media acquisition unit 551, synchronized media decoding unit 561, synchronized media scene processing unit 531, and synchronized media device processing unit 621 execute synchronized media playback processing and play the synchronized media under the control of the control unit 523. When the processing of step S504 ends, the processing proceeds to step S506.

[0254] Furthermore, the processing of step S505 is executed in parallel with the processing of step S504. For example, in step S505, the asynchronous media acquisition unit 552, the asynchronous media decoding unit 562, the asynchronous media scene processing unit 532, and the asynchronous media device processing unit 622 execute asynchronous media low-latency interaction playback processing under the control of the control unit 523, and perform low-latency interaction playback of the asynchronous media. When the processing of step S505 ends, the process proceeds to step S506.

[0255] In step S506, the control unit 523 determines whether or not to end the playback process. If it is determined not to end, the process returns to steps S504 and S505. If it is determined in step S506 that the playback process is to end, the playback process ends.

[0256] <Synchronized Media Playback Processing Flow> Next, an example of the flow of synchronized media playback processing executed in step S504 of FIG. 44 will be described with reference to the flowchart of FIG.

[0257] When the synchronized media playback process is started, the synchronized media acquisition unit 551 acquires a synchronized media file in step S521.

[0258] In step S522, the synchronized media decoding unit 561 extracts the bit stream stored in the synchronized media file, decodes the bit stream, and generates (restores) the synchronized media.

[0259] In step S523, the synchronized media scene processing unit 531 renders the generated (restored) synchronized media and generates output data of the synchronized media.

[0260] Then, in step S524, the synchronized media scene processing unit 531 supplies the generated output data to the output device 600 (synchronized media device processing unit 621) for output. When the processing of step S524 ends, the synchronized media playback processing ends, and the processing returns to FIG.

[0261] <Flow 1 of asynchronous media low-latency interaction playback processing> Next, an example of the flow of asynchronous media low-latency interaction playback processing when applying method 1 of the present technology described above in <3. Low-latency interaction playback based on author's intentions>, which is executed in step S505 of Fig. 44, will be described with reference to the flowchart of Fig. 46.

[0262] When the asynchronous media low-latency interaction playback process is started, in step S541, the asynchronous media acquisition unit 552 acquires asynchronous media files that can be output with low latency, at a timing that will allow for output, in accordance with the control of the control unit 523, i.e., based on the low-latency processing information stored in the scene description. The control unit 523 can determine whether the asynchronous media files can be output with low latency, and the timing that will allow for output, by analyzing the low-latency processing information.

[0263] In step S542, the asynchronous media decoding unit 562 extracts a bitstream from the asynchronous media file in accordance with the control of the control unit 523, i.e., based on the low-delay processing information stored in the scene description, and decodes the bitstream in time for output to generate (restore) the asynchronous media. The control unit 523 can determine the timing for output, etc. by analyzing the low-delay processing information.

[0264] In step S543, the asynchronous media scene processing unit 532, under the control of the control unit 523, i.e., based on the low-latency processing information stored in the scene description, renders the generated asynchronous media and generates action data at a timing that will allow for output. The control unit 523 can determine the timing that will allow for output by analyzing the low-latency processing information. Then, in step S544, under the control of the control unit 523, i.e., based on the low-latency processing information stored in the scene description, the asynchronous media scene processing unit 532 supplies the action data to the output device 600 (asynchronous media device processing unit 622) and causes it to be output with low latency in response to the occurrence of an interaction.

[0265] In other words, the asynchronous media scene processing unit 532 performs low-latency interaction playback based on low-latency processing information, including a maximum processing delay, that controls low-latency interaction playback and is stored in a scene description that describes a scene in 3D space in which 3D objects are placed. Here, "low-latency interaction playback" refers to interaction playback that is triggered by the occurrence of an interaction in 3D space and executes an action corresponding to that interaction, with a delay shorter than the maximum processing delay. The parameter "maximum processing delay" indicates the maximum allowable time from the occurrence of the interaction to the output of the action data for the action on the output device.

[0266] For example, the asynchronous media scene processing unit 532 may select an action for which low-latency interaction playback is possible, and perform low-latency interaction playback for the selected action.

[0267] The low-latency processing information may include a parameter "processing policy" indicating a processing policy for low-latency interaction playback. The asynchronous media scene processing unit 532 may then perform low-latency interaction playback in accordance with the parameter "processing policy." For example, the parameter "processing policy" may include information indicating whether to prioritize action presentation or suppression of erroneous action presentation in low-latency interaction playback. The low-latency processing information may also include a parameter "processing priority" indicating the priority of processing for actions to be played back with low-latency interaction. The asynchronous media scene processing unit 532 may then prioritize processing of actions with higher priorities based on the parameter "processing priority." The low-latency processing information may also include flag information "low-latency processing flag" indicating whether low-latency interaction playback is necessary. The asynchronous media scene processing unit 532 may then perform low-latency interaction playback if the low-latency processing flag is true. The low-latency processing information may also include an application action indicating an action for which low-latency interaction playback is to be performed. The asynchronous media scene processing unit 532 may then select an action to process from among the actions specified by the parameter “application action.” The low-latency processing information may also include any combination of these pieces of information.

[0268] When the processing in step S544 is completed, the asynchronous media low-delay interaction playback processing is completed, and the processing returns to FIG.

[0269] By executing each process in this manner, the client device 500 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the client device 500 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0270] <Flow of asynchronous media low-latency interaction playback processing 2> Next, an example of the flow of asynchronous media low-latency interaction playback processing when applying method 1-1 of the present technology described above in <3. Low-latency interaction playback based on author's intentions> will be described with reference to the flowchart in Fig. 47 .

[0271] The processes of steps S561 to S563 are executed in the same manner as the processes of steps S541 to S543 in FIG. 46. In step S564, the asynchronous media scene processing unit 532 derives a predicted probability in accordance with control by the control unit 523, i.e., based on the low-latency processing information stored in the scene description. In step S565, the asynchronous media scene processing unit 532 determines a trigger condition based on the predicted trigger activation control information stored in the scene description. In step S566, the asynchronous media scene processing unit 532 supplies the action data for which the trigger condition is met to the output device at a timing that allows low-latency output, taking into account the low-latency processing information stored in the scene description, the device startup time, etc., and causes the action data to be output.

[0272] That is, the asynchronous media scene processing unit 532 may perform low-latency interaction playback based on predictive trigger activation control information stored in the scene description. Here, the "predictive trigger activation control information" is information used by the processing unit that generates action data to control low-latency interaction playback in which an interaction is predicted. For example, the asynchronous media scene processing unit 532 may perform low-latency interaction playback by generating action data, deriving the probability of an interaction occurring, determining whether a trigger condition exists, and providing the action data to the output device 600 when the trigger condition exists.

[0273] The predictive trigger activation control information may include a predictive trigger activation control flag indicating whether the processing unit generating the action data derives an interaction occurrence probability as a prediction of an interaction and determines whether the occurrence probability is a threshold value. If the predictive trigger activation control flag is true, the asynchronous media scene processing unit 532 may derive the interaction occurrence probability. The predictive trigger activation control information may also include a predictive trigger activation control threshold value indicating a threshold value used to determine the interaction occurrence probability. The asynchronous media scene processing unit 532 may then use the predictive trigger activation control threshold value to determine the interaction occurrence probability as a trigger condition. The predictive trigger activation control information may also include a "recommended parameter" parameter indicating a parameter recommended as a parameter used for prediction. The asynchronous media scene processing unit 532 may then use the recommended parameter value to derive the occurrence probability. The predictive trigger activation control information may also include any combination of these pieces of information.

[0274] When the processing in step S566 is completed, the asynchronous media low-delay interaction playback processing is completed, and the processing returns to FIG.

[0275] By executing each process in this manner, the client device 500 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the client device 500 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0276] <Output Process Flow 1> Next, an example of the flow of the output process executed by the output device 600 in response to this asynchronous media low-latency interaction playback process will be described with reference to the flowchart in FIG.

[0277] When the output process starts, the asynchronous media device processing unit 622 of the output device 600 acquires action data supplied from the client device 500 (asynchronous media scene processing unit 532) in step S581.

[0278] In step S582, the asynchronous media device processing unit 622 activates the device (the function of the output device 600 to output asynchronous media).

[0279] When the device is ready to be driven, the asynchronous media device processing unit 622 outputs the acquired action data (asynchronous media output data) in step S583.

[0280] When the process of step S583 ends, the output process ends.

[0281] <Flow of asynchronous media low-latency interaction playback processing 3> Next, an example of the flow of asynchronous media low-latency interaction playback processing when applying method 1-2 of the present technology described above in <3. Low-latency interaction playback based on author's intentions> will be described with reference to the flowchart in Fig. 49 .

[0282] The processes of steps S601 to S604 are executed in the same manner as the processes of steps S561 to S564 in Fig. 47. In step S605, the asynchronous media scene processing unit 532 supplies the predicted trigger activation offload control information stored in the scene description, the derived predicted probability data, and the generated action data to the output device in accordance with the control of the control unit 523, that is, at a timing that takes into account the low-latency processing information stored in the scene description, the device startup time, etc.

[0283] That is, the asynchronous media scene processing unit 532 may perform low-latency interaction playback based on predictive trigger activation offload control information stored in the scene description. Here, the "predictive trigger activation offload control information" is information used by the output device 600 to control low-latency interaction playback in which interactions are predicted. For example, the asynchronous media scene processing unit 532 may perform low-latency interaction playback by generating action data, deriving the probability of an interaction occurring, and providing the occurrence probability data, action data, and predictive trigger activation offload control information to the output device 600.

[0284] The predictive trigger activation offload control information may include flag information "predictive trigger activation offload control flag" indicating whether the output device 600 will predict an interaction and determine the probability of the interaction occurrence against a threshold. If the predictive trigger activation offload control flag is true, the asynchronous media scene processing unit 532 may derive the probability of the interaction occurrence. The predictive trigger activation offload control information may also include a parameter "predictive trigger activation offload control threshold" indicating a threshold used to determine the threshold for the probability of the interaction occurrence. The predictive trigger activation offload control information may also include format information indicating a format of the interaction occurrence probability data provided to the output device. The asynchronous media scene processing unit 532 may then derive the interaction occurrence probability based on the format information. For example, the format information may include a parameter "probability data frame rate" indicating an update frequency of the occurrence probability data. The asynchronous media scene processing unit 532 may then derive the interaction occurrence probability at a frequency corresponding to the probability data frame rate. The format information may also include an "intra-frame probability data duration" parameter indicating a time length corresponding to one frame's worth of occurrence probability data. The asynchronous media scene processing unit 532 may then generate occurrence probability data in a format corresponding to the intra-frame probability data duration. The format information may also include an "intra-frame probability data rate" parameter indicating a time length corresponding to one occurrence probability data. The asynchronous media scene processing unit 532 may then generate occurrence probability data in a format corresponding to the intra-frame probability data rate. The predictive trigger activation offload control information may also include a "recommended parameter" parameter indicating a parameter recommended as a parameter to be used in predicting an interaction. The asynchronous media scene processing unit 532 may then use the recommended parameter to derive the occurrence probability of an interaction. The predictive trigger activation offload control information may also include any combination of these pieces of information.

[0285] When the processing in step S605 is completed, the asynchronous media low-delay interaction playback processing is completed, and the processing returns to FIG.

[0286] By executing each process in this manner, the client device 500 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the client device 500 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0287] <Output Processing Flow 2> Next, an example of the flow of the output processing executed by the output device 600 in response to this asynchronous media low-latency interaction playback processing (that is, an example of the flow of the output processing when applying Method 1-2 of the present technology described above in <3. Low-latency interaction playback based on the author's intention>) will be described with reference to the flowchart in FIG. 50 .

[0288] When the output process starts, in step S621, the asynchronous media device processing unit 622 of the output device 600 acquires predicted trigger activation offload control information, predicted probability data, and action data supplied from the client device 500 (asynchronous media scene processing unit 532).

[0289] In step S622, the asynchronous media device processing unit 622 determines the trigger condition using the predicted probability data based on the predicted trigger activation offload control information.

[0290] If the trigger condition is met, the asynchronous media device processing section 622 activates the device (the function of the output device 600 to output asynchronous media) in step S623.

[0291] When the device is ready to be driven, the asynchronous media device processing unit 622 outputs the action data (asynchronous media output data) for which the trigger condition is met in step S624.

[0292] When the process of step S624 ends, the output process ends.

[0293] In other words, the asynchronous media device processing unit 622 may acquire occurrence probability data indicating the probability of an interaction occurring, action data indicating an action for that interaction, and predicted trigger activation offload control information, judge the occurrence probability data against a threshold based on the predicted trigger activation offload control information, and output the action data if the trigger condition is met.

[0294] By executing each process in this manner, the output device 600 can achieve the same effect as described above in <3. Low-latency interaction playback based on the author's intentions>. That is, the output device 600 can achieve low-latency interaction playback in accordance with the intentions of the content provider (author, etc.), and can suppress at least one of a reduction in the quality of interaction playback and unnecessary high-load processing.

[0295] <7. Supplementary Notes> <Computer> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that make up the software are installed on a computer. Here, the term "computer" includes computers built into dedicated hardware, and general-purpose personal computers, for example, that can execute various functions by installing various programs.

[0296] FIG. 51 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.

[0297] In a computer 1900 shown in FIG. 51, a CPU (Central Processing Unit) 1901, a ROM (Read Only Memory) 1902, and a RAM (Random Access Memory) 1903 are interconnected via a bus 1904.

[0298] An input / output interface 1910 is also connected to the bus 1904. An input unit 1911, an output unit 1912, a storage unit 1913, a communication unit 1914, and a drive 1915 are connected to the input / output interface 1910.

[0299] The input unit 1911 includes, for example, a keyboard, a mouse, a microphone, a touch panel, and an input terminal. The output unit 1912 includes, for example, a display, a speaker, and an output terminal. The storage unit 1913 includes, for example, a hard disk, a RAM disk, and a non-volatile memory. The communication unit 1914 includes, for example, a network interface. The drive 1915 drives removable media 1921 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.

[0300] In a computer configured as described above, the CPU 1901 performs the above-described series of processes by, for example, loading a program stored in the storage unit 1913 into the RAM 1903 via the input / output interface 1910 and the bus 1904 and executing the program. The RAM 1903 also stores data and the like necessary for the CPU 1901 to execute various processes.

[0301] The program executed by the computer can be applied by recording it on, for example, removable media 1921 such as package media. In this case, the program can be installed in storage unit 1913 via input / output interface 1910 by attaching removable media 1921 to drive 1915.

[0302] This program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, digital satellite broadcasting, etc. In this case, the program can be received by the communication unit 1914 and installed in the storage unit 1913.

[0303] Alternatively, this program can be installed in advance in the ROM 1902 or the storage unit 1913 .

[0304] <Applicable Targets of the Present Technology> The present technology can be applied to any encoding / decoding method.

[0305] Furthermore, the present technology can be applied to any configuration, for example, various electronic devices.

[0306] Furthermore, for example, the present technology can also be implemented as part of an apparatus, such as a processor (e.g., a video processor) as a system LSI (Large Scale Integration), a module using multiple processors (e.g., a video module), a unit using multiple modules (e.g., a video unit), or a set in which other functions are added to a unit (e.g., a video set).

[0307] Furthermore, for example, the present technology can also be applied to a network system configured with multiple devices. For example, the present technology may be implemented as cloud computing in which multiple devices share and collaborate on processing via a network. For example, the present technology may be implemented in a cloud service that provides image (video)-related services to any terminal, such as a computer, an AV (Audio Visual) device, a portable information processing terminal, or an IoT (Internet of Things) device.

[0308] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.

[0309] <Fields and uses to which this technology can be applied> Systems, devices, processing units, etc. to which this technology is applied can be used in any field, for example, transportation, medical care, crime prevention, agriculture, livestock farming, mining, beauty, factories, home appliances, weather, nature monitoring, etc. In addition, the uses thereof are also arbitrary.

[0310] For example, the present technology can be applied to systems and devices used to provide viewing content, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for transportation, such as monitoring traffic conditions and controlling automatic driving. Furthermore, for example, the present technology can also be applied to systems and devices used for security. Furthermore, for example, the present technology can also be applied to systems and devices used for automatic control of machines, etc. Furthermore, for example, the present technology can also be applied to systems and devices used for agriculture and livestock farming. Furthermore, for example, the present technology can also be applied to systems and devices used to monitor natural conditions, such as volcanoes, forests, and oceans, and wildlife. Furthermore, for example, the present technology can also be applied to systems and devices used for sports.

[0311] <Others> In this specification, a "flag" refers to information for identifying multiple states, and includes not only information used to identify two states, true (1) or false (0), but also information capable of identifying three or more states. Therefore, the value that this "flag" can take may be, for example, two values, 1 / 0, or three or more values. That is, the number of bits constituting this "flag" is arbitrary, and may be one bit or multiple bits. Furthermore, identification information (including flags) can be included not only in a bitstream, but also in a bitstream that includes differential information of the identification information relative to certain reference information. Therefore, in this specification, "flag" and "identification information" encompass not only the information itself, but also differential information relative to the reference information.

[0312] Furthermore, various information (e.g., metadata) related to the coded data (bitstream) may be transmitted or recorded in any form as long as it is associated with the coded data. Here, the term "associate" means, for example, making one piece of data available (linked) when processing the other piece of data. That is, data associated with each other may be combined into one piece of data or may be stored as separate pieces of data. For example, information associated with coded data (image) may be transmitted over a transmission path separate from that of the coded data (image). Furthermore, for example, information associated with coded data (image) may be recorded on a recording medium separate from that of the coded data (image) (or on a different recording area of ​​the same recording medium). Note that this "association" may refer not to the entire data, but to only a portion of the data. For example, an image and information corresponding to that image may be associated with each other in any unit, such as multiple frames, one frame, or a portion of a frame.

[0313] In this specification, terms such as "composite," "multiplex," "add," "integrate," "include," "store," "embed," "insert," and the like refer to combining multiple items into one, such as combining encoded data and metadata into one piece of data, and refer to one method of "associating" as described above.

[0314] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.

[0315] For example, a configuration described as one device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, configurations described above as multiple devices (or processing units) may be combined and configured as one device (or processing unit). Of course, configurations other than those described above may be added to the configuration of each device (or each processing unit). Furthermore, as long as the configuration and operation of the entire system are substantially the same, part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0316] Furthermore, for example, the above-described program may be executed in any device, as long as the device has the necessary functions (functional blocks, etc.) and is able to obtain the necessary information.

[0317] Also, for example, each step of a single flowchart may be executed by a single device, or may be shared and executed by multiple devices. Furthermore, when a single step includes multiple processes, the multiple processes may be executed by a single device, or may be shared and executed by multiple devices. In other words, multiple processes included in a single step can be executed as multiple step processes. Conversely, processes described as multiple steps can be executed collectively as a single step.

[0318] For example, the steps of a program executed by a computer may be executed in chronological order in the order described herein, or may be executed in parallel or individually at the required timing, such as when a call is made. In other words, as long as no contradiction occurs, the steps may be executed in an order different from the order described above. Furthermore, the steps of this program may be executed in parallel with the processing of another program, or may be executed in combination with the processing of another program.

[0319] Furthermore, for example, multiple technologies related to the present technology can be implemented independently and independently, as long as no contradiction occurs. Of course, any multiple technologies can also be implemented in combination. For example, part or all of the present technology described in any embodiment can be implemented in combination with part or all of the present technology described in another embodiment. Furthermore, part or all of any of the above-described present technologies can be implemented in combination with other technologies not described above.

[0320] Note that the present technology can also be configured as follows. (1) An information processing device including a scene description generation unit that generates a scene description representing a scene in a 3D space in which a 3D object is placed, and stores, in the scene description, low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, wherein the low-latency interaction playback is a process in which interaction playback, triggered by the occurrence of an interaction in the 3D space, executes an action corresponding to the interaction, with a delay shorter than the maximum processing delay, and the maximum processing delay indicates a maximum allowable time from the occurrence of the interaction to the output of action data of the action at an output device. (2) The information processing device described in (1), wherein the low-latency processing information further includes information indicating a processing policy for the low-latency interaction playback. (3) The information processing device described in (2), wherein the processing policy includes information indicating whether to prioritize presentation of the action or suppression of erroneous presentation of the action in the low-latency interaction playback. (4) The information processing device according to any one of (1) to (3), wherein the low-latency processing information further includes a processing priority indicating a priority of a process related to the action to be played back in low-latency interaction. (5) The information processing device according to any one of (1) to (4), wherein the low-latency processing information further includes a low-latency processing flag indicating whether the low-latency interaction playback is necessary. (6) The information processing device according to any one of (1) to (5), wherein the low-latency processing information further includes an applied action indicating the action for which the low-latency interaction playback is performed. (7) The information processing device according to any one of (1) to (6), wherein the scene description generation unit further stores predictive trigger activation control information in the scene description, and the predictive trigger activation control information is information used for controlling the low-latency interaction playback in which the processing unit that generates the action data predicts the interaction.(8) The information processing device according to (7), wherein the predictive trigger activation control information includes a predictive trigger activation control flag indicating whether the processing unit derives a probability of occurrence of the interaction as the prediction and performs threshold judgment on the occurrence probability. (9) The information processing device according to (7) or (8), wherein the predictive trigger activation control information includes a predictive trigger activation control threshold indicating a threshold used for threshold judgment on the probability of occurrence of the interaction. (10) The information processing device according to any of (7) to (9), wherein the predictive trigger activation control information includes recommended parameters recommended as parameters used for the prediction. (11) The information processing device according to any of (1) to (10), wherein the scene description generation unit further stores predictive trigger activation offload control information in the scene description, and the predictive trigger activation offload control information is information used to control the low-latency interaction playback in which the output device predicts the interaction. (12) The information processing device according to (11), wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control flag indicating whether the output device will perform a threshold determination on the probability of the interaction occurrence as the prediction. (13) The information processing device according to (11) or (12), wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control threshold indicating a threshold used for threshold determination of the probability of the interaction occurrence. (14) The information processing device according to any of (11) to (13), wherein the predictive trigger activation offload control information includes format information indicating a format of occurrence probability data of the interaction to be provided to the output device. (15) The information processing device according to (14), wherein the format information includes a probability data frame rate indicating an update frequency of the occurrence probability data. (16) The information processing device according to (14) or (15), wherein the format information includes an intra-frame probability data duration indicating a time length corresponding to one frame of the occurrence probability data.(17) The information processing device according to any one of (14) to (16), wherein the format information includes an intra-frame probability data rate indicating a time length corresponding to one of the occurrence probability data. (18) The information processing device according to any one of (11) to (17), wherein the prediction trigger activation offload control information includes recommended parameters that are recommended as parameters to be used for the prediction. (19) An information processing method comprising: generating a scene description that represents a scene of a 3D space in which a 3D object is placed; storing, in the scene description, low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, wherein the low-latency interaction playback is processing that executes an action corresponding to an interaction using the occurrence of an interaction in the 3D space as a trigger, with a delay shorter than the maximum processing delay; and the maximum processing delay indicates a maximum allowable time from the occurrence of the interaction to the output of the action on an output device. (20) A program for causing a computer to execute a process of generating a scene description representing a scene in a 3D space in which a 3D object is placed, and storing low-latency processing information in the scene description that includes a maximum processing delay and controls low-latency interaction playback, wherein the low-latency interaction playback is a process in which an interaction occurring in the 3D space is used as a trigger to execute an action corresponding to the interaction, with a delay shorter than the maximum processing delay, and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of the action on an output device.

[0321] (31) An information processing device comprising: a scene description supply unit that supplies, in response to a request, a scene description that represents a scene in a 3D space in which a 3D object is placed and that stores low-latency processing information, wherein the low-latency processing information includes a maximum processing delay and controls low-latency interaction playback, wherein the low-latency interaction playback is a process in which interaction playback, triggered by the occurrence of an interaction in the 3D space, executes an action corresponding to the interaction, with a delay shorter than the maximum processing delay, and the maximum processing delay indicates a maximum allowable time from the occurrence of the interaction to the output of action data of the action at an output device. (32) The information processing device according to (31), wherein the low-latency processing information further includes information indicating a processing policy for the low-latency interaction playback. (33) The information processing device according to (32), wherein the processing policy includes information indicating whether to prioritize presentation of the action or suppression of erroneous presentation of the action in the low-latency interaction playback. (34) The information processing device according to any one of (31) to (33), wherein the low-latency processing information further includes a processing priority indicating a priority of a process related to the action to be played back in low-latency interaction. (35) The information processing device according to any one of (31) to (34), wherein the low-latency processing information further includes a low-latency processing flag indicating whether the low-latency interaction playback is necessary. (36) The information processing device according to any one of (31) to (35), wherein the low-latency processing information further includes an applied action indicating the action for which the low-latency interaction playback is performed. (37) The information processing device according to any one of (31) to (36), wherein the scene description further stores predictive trigger activation control information, and the predictive trigger activation control information is information used for controlling the low-latency interaction playback in which a processing unit that generates the action data predicts the interaction.(38) The information processing device according to (37), wherein the predictive trigger activation control information includes a predictive trigger activation control flag indicating whether the processing unit derives an occurrence probability of the interaction as the prediction and performs threshold judgment on the occurrence probability. (39) The information processing device according to (37) or (38), wherein the predictive trigger activation control information includes a predictive trigger activation control threshold indicating a threshold used for threshold judgment of the occurrence probability of the interaction. (40) The information processing device according to any of (37) to (39), wherein the predictive trigger activation control information includes recommended parameters recommended as parameters used for the prediction. (41) The information processing device according to any of (31) to (40), wherein the scene description further stores predictive trigger activation offload control information, and the predictive trigger activation offload control information is information used to control the low-latency interaction playback in which the output device predicts the interaction. (42) The information processing device according to (41), wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control flag indicating whether the output device will perform threshold determination on the probability of the interaction as the prediction. (43) The information processing device according to (41) or (42), wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control threshold indicating a threshold used for threshold determination of the probability of the interaction. (44) The information processing device according to any of (41) to (43), wherein the predictive trigger activation offload control information includes format information indicating a format of occurrence probability data of the interaction to be provided to the output device. (45) The information processing device according to (44), wherein the format information includes a probability data frame rate indicating an update frequency of the occurrence probability data. (46) The information processing device according to (44) or (45), wherein the format information includes an intra-frame probability data duration indicating a time length corresponding to one frame of the occurrence probability data.(47) The information processing device according to any one of (44) to (46), wherein the format information includes an intra-frame probability data rate indicating a time length corresponding to one of the occurrence probability data. (48) The information processing device according to any one of (41) to (47), wherein the prediction trigger activation offload control information includes recommended parameters recommended as parameters to be used for the prediction. (49) An information processing method comprising: supplying, in response to a request, a scene description representing a scene in a 3D space in which a 3D object is placed and storing low-latency processing information; the low-latency processing information including a maximum processing delay; and controlling low-latency interaction playback, wherein interaction playback is processing in which an occurrence of an interaction in the 3D space is used as a trigger to execute an action corresponding to the interaction, with a delay shorter than the maximum processing delay; and the maximum processing delay indicates a maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device. (50) A program for causing a computer to execute a process that, upon request, provides a scene description that represents a scene in a 3D space in which a 3D object is placed and stores low-latency processing information, the low-latency processing information including a maximum processing delay and controls low-latency interaction playback, the low-latency interaction playback being a process in which an interaction occurring in the 3D space is used as a trigger to execute an action corresponding to the interaction, with a delay shorter than the maximum processing delay, and the maximum processing delay indicates a maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

[0322] (61) An information processing device comprising: a scene processing unit that performs low-latency interaction playback based on low-latency processing information including a maximum processing delay and controlling low-latency interaction playback, the low-latency interaction playback being stored in a scene description that represents a scene in a 3D space in which a 3D object is placed, wherein the low-latency interaction playback is a process in which interaction playback, triggered by the occurrence of an interaction in the 3D space, is performed with a delay shorter than the maximum processing delay, and the maximum processing delay indicates a maximum allowable time from the occurrence of the interaction to the output of action data of the action at an output device. (62) The information processing device described in (61), wherein the scene processing unit selects the action for which low-latency interaction playback is possible, and performs the low-latency interaction playback for the selected action. (63) The information processing device described in (62), wherein the low-latency processing information further includes information indicating a processing policy for the low-latency interaction playback, and the scene processing unit performs the low-latency interaction playback in accordance with the processing policy. (64) The information processing device according to (63), wherein the processing policy includes information indicating whether to prioritize presentation of the action or suppression of erroneous presentation of the action in the low-latency interaction playback. (65) The information processing device according to any of (62) to (64), wherein the low-latency processing information further includes a processing priority indicating a priority of processing for the action to be played back in the low-latency interaction, and the scene processing unit preferentially processes the action having a higher priority. (66) The information processing device according to any of (62) to (65), wherein the low-latency processing information further includes a low-latency processing flag indicating whether the low-latency interaction playback is necessary, and the scene processing unit performs the low-latency interaction playback when the low-latency processing flag is true.(67) The information processing device according to any one of (62) to (66), wherein the low-latency processing information further includes an applied action indicating the action for which the low-latency interaction playback is performed, and the scene processing unit selects from the actions specified by the applied action. (68) The information processing device according to any one of (62) to (67), wherein the scene processing unit performs the low-latency interaction playback based on predictive trigger activation control information stored in the scene description, and the predictive trigger activation control information is information used for controlling the low-latency interaction playback in which a processing unit that generates the action data predicts the interaction. (69) The information processing device according to (68), wherein the scene processing unit performs the low-latency interaction playback by generating the action data, deriving the probability of the interaction occurring, determining a trigger condition, and providing the action data to the output device when the trigger condition is met. (70) The information processing device according to (69), wherein the predictive trigger activation control information includes a predictive trigger activation control flag indicating whether the processing unit derives the occurrence probability as the prediction and makes a threshold decision on the occurrence probability, and the scene processing unit derives the occurrence probability when the predictive trigger activation control flag is true. (71) The information processing device according to (69) or (70), wherein the predictive trigger activation control information includes a predictive trigger activation control threshold indicating a threshold used for threshold decision on the occurrence probability, and the scene processing unit makes a threshold decision on the occurrence probability using the predictive trigger activation control threshold as a decision on the trigger condition. (72) The information processing device according to any of (69) to (71), wherein the predictive trigger activation control information includes recommended parameters recommended as parameters used in the prediction, and the scene processing unit derives the occurrence probability using the recommended parameters.(73) The information processing device according to any one of (62) to (72), wherein the scene processing unit performs the low-latency interaction playback based on predictive trigger activation offload control information stored in the scene description, and the predictive trigger activation offload control information is information used for controlling the low-latency interaction playback in which the output device predicts the interaction. (74) The information processing device according to (73), wherein the scene processing unit executes, as the low-latency interaction playback, generating the action data, deriving an occurrence probability of the interaction, and providing occurrence probability data, the action data, and the predictive trigger activation offload control information to the output device. (75) The information processing device according to (74), wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control flag indicating whether the output device will determine, as the prediction, the occurrence probability against a threshold, and the scene processing unit derives the occurrence probability when the predictive trigger activation offload control flag is true. (76) The information processing device according to (74) or (75), wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control threshold indicating a threshold used for threshold determination of the occurrence probability. (77) The information processing device according to any of (74) to (76), wherein the predictive trigger activation offload control information includes format information indicating a format of the occurrence probability data to be provided to the output device, and the scene processing unit derives the occurrence probability based on the format information. (78) The information processing device according to (77), wherein the format information includes a probability data frame rate indicating an update frequency of the occurrence probability data, and the scene processing unit derives the occurrence probability at a frequency corresponding to the probability data frame rate. (79) The information processing device according to (77) or (78), wherein the format information includes an intra-frame probability data duration indicating a time length corresponding to one frame of the occurrence probability data, and the scene processing unit generates the occurrence probability data in the format corresponding to the intra-frame probability data duration.(80) The information processing device according to any one of (77) to (79), wherein the format information includes an intra-frame probability data rate indicating a time length corresponding to one of the occurrence probability data, and the scene processing unit generates the occurrence probability data in the format according to the intra-frame probability data rate. (81) The information processing device according to any one of (74) to (80), wherein the prediction trigger activation offload control information includes recommended parameters recommended as parameters to be used for the prediction, and the scene processing unit derives the occurrence probability using the recommended parameters. (82) An information processing method, comprising: performing low-latency interaction playback based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, the low-latency interaction playback being stored in a scene description that represents a scene in a 3D space in which a 3D object is placed; the low-latency interaction playback is a process in which interaction playback, which is triggered by the occurrence of an interaction in the 3D space and executes an action corresponding to the interaction, is executed with a delay lower than the maximum processing delay; and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device. (83) A program for causing a computer to execute a process in which low-latency interaction playback is performed based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, the low-latency interaction playback being stored in a scene description that represents a scene in a 3D space in which a 3D object is placed, the low-latency interaction playback being a process in which interaction playback that executes an action corresponding to an interaction, triggered by the occurrence of the interaction in the 3D space, is executed with a delay lower than the maximum processing delay, and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

[0323] (91) An information processing device comprising: a device processing unit that acquires occurrence probability data indicating the probability of an interaction occurring, action data indicating an action for the interaction, and predictive trigger activation offload control information, determines a threshold value for the occurrence probability data based on the predictive trigger activation offload control information, and outputs the action data when a trigger condition is met; the predictive trigger activation offload control information is information used to control low-latency interaction playback; the low-latency interaction playback is a process in which interaction playback, which executes the action triggered by the occurrence of the interaction, is executed with a delay shorter than a maximum processing delay; and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of the action data. (92) An information processing method comprising: acquiring occurrence probability data indicating the probability of an interaction occurring, action data indicating an action for the interaction, and predictive trigger activation offload control information; determining a threshold for the occurrence probability data based on the predictive trigger activation offload control information; and outputting the action data when a trigger condition is met; the predictive trigger activation offload control information is information used to control low-latency interaction playback; the low-latency interaction playback is a process in which interaction playback, which executes the action triggered by the occurrence of the interaction, is executed with a delay shorter than a maximum processing delay; and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of the action data.(93) A program for causing a computer to execute a process of acquiring occurrence probability data indicating the probability of an interaction occurring, action data indicating an action for the interaction, and predictive trigger activation offload control information, making a threshold judgment on the occurrence probability data based on the predictive trigger activation offload control information, and outputting the action data when a trigger condition is met, wherein the predictive trigger activation offload control information is information used to control low-latency interaction playback, and the low-latency interaction playback is a process in which interaction playback, which executes the action triggered by the occurrence of the interaction, is executed with a delay shorter than a maximum processing delay, and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of the action data.

[0324] 300 File generation device, 311 SD generation unit, 312 Encoding unit, 313 File generation unit, 314 Storage unit, 315 Supply unit, 321 SD encoding unit, 322 Synchronous media encoding unit, 323 Asynchronous media encoding unit, 331 SD file generation unit, 332 Synchronous media file generation unit, 333 Asynchronous media file generation unit, 400 Distribution server, 411 Acquisition unit, 412 Storage unit, 413 Distribution unit, 421 SD distribution unit, 422 Synchronous media distribution unit, 423 Asynchronous media distribution unit, 500 Client device, 511 PE, 512 MAF, 521 SD acquisition unit, 522 SD decoding unit, 523 Control unit, 524 Scene processing unit, 531 Synchronous media scene processing unit, 532 Asynchronous media scene processing unit, 541 Media acquisition unit, 542 media decoding unit, 551 synchronous media acquisition unit, 552 asynchronous media acquisition unit, 561 synchronous media decoding unit, 562 asynchronous media decoding unit, 600 output device, 611 device processing unit, 621 synchronous media device processing unit, 622 asynchronous media device processing unit, 1900 computer

Claims

1. An information processing device comprising: a scene description generation unit that generates a scene description representing a scene in a 3D space in which a 3D object is placed; and stores in the scene description low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback; the low-latency interaction playback is a process in which interaction playback, which is triggered by the occurrence of an interaction in the 3D space and executes an action corresponding to the interaction, is executed with a delay lower than the maximum processing delay; and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

2. The information processing device according to claim 1, wherein the low-latency processing information further includes information indicating a processing policy for the low-latency interaction playback.

3. The information processing device according to claim 1, wherein the low-latency processing information further includes a processing priority indicating a priority of processing related to the action to be played back as the low-latency interaction.

4. The information processing device according to claim 1, wherein the low-delay processing information further includes a low-delay processing flag indicating whether the low-delay interaction playback is required.

5. The information processing device according to claim 1, wherein the low-latency processing information further includes an applied action indicating the action for which the low-latency interaction playback is performed.

6. The information processing device of claim 1, wherein the scene description generation unit further stores predictive trigger activation control information in the scene description, and the predictive trigger activation control information is information used by the processing unit that generates the action data to control the low-latency interaction playback in which the interaction is predicted.

7. The information processing device according to claim 6, wherein the predictive trigger activation control information includes a predictive trigger activation control threshold value indicating a threshold value used for determining a threshold value of the probability of occurrence of the interaction.

8. The information processing device of claim 1, wherein the scene description generation unit further stores predictive trigger activation offload control information in the scene description, and the predictive trigger activation offload control information is information used to control the low-latency interaction playback in which the output device predicts the interaction.

9. The information processing device according to claim 8, wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control threshold value indicating a threshold value used to determine the threshold value of the probability of occurrence of the interaction.

10. An information processing method that generates a scene description that represents a scene in a 3D space in which a 3D object is placed, and stores low-latency processing information in the scene description that includes a maximum processing delay and controls low-latency interaction playback, wherein the low-latency interaction playback is a process in which interaction playback that is triggered by the occurrence of an interaction in the 3D space and executes an action corresponding to the interaction is executed with a delay lower than the maximum processing delay, and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of the action on an output device.

11. An information processing device comprising: a scene processing unit that selects an action that corresponds to an interaction in a 3D space and that is capable of low-latency interaction playback based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, the low-latency interaction playback being stored in a scene description that represents a scene in a 3D space in which a 3D object is placed, and that performs the low-latency interaction playback for the selected action; wherein the low-latency interaction playback is a process in which interaction playback that executes the action in response to the occurrence of the interaction as a trigger is performed with a delay lower than the maximum processing delay; and the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

12. The information processing device according to claim 11, wherein the low-latency processing information further includes information indicating a processing policy for the low-latency interaction playback, and the scene processing unit performs the low-latency interaction playback in accordance with the processing policy.

13. The information processing device according to claim 11, wherein the low-latency processing information further includes a processing priority indicating a priority of processing related to the action to be played back in the low-latency interaction, and the scene processing unit processes the action with a higher priority preferentially.

14. The information processing device according to claim 11, wherein the low-latency processing information further includes a low-latency processing flag indicating whether the low-latency interaction playback is necessary, and the scene processing unit performs the low-latency interaction playback when the low-latency processing flag is true.

15. The information processing device according to claim 11, wherein the low-latency processing information further includes an applied action indicating the action for which the low-latency interaction playback is performed, and the scene processing unit selects from among the actions specified by the applied action.

16. The information processing device of claim 11, wherein the scene processing unit generates the action data, derives the probability of the interaction occurring, determines the trigger condition, and provides the action data to the output device when the trigger condition is met, based on predictive trigger activation control information stored in the scene description, as the low-latency interaction playback, and the predictive trigger activation control information is information used by the processing unit that generates the action data to control the low-latency interaction playback in which the interaction is predicted.

17. The information processing device described in claim 16, wherein the predictive trigger activation control information includes a predictive trigger activation control threshold indicating a threshold used for threshold determination of the occurrence probability, and the scene processing unit performs threshold determination of the occurrence probability using the predictive trigger activation control threshold to determine the trigger condition.

18. The information processing device of claim 11, wherein the scene processing unit generates the action data, derives the probability of the interaction occurring based on predictive trigger activation offload control information stored in the scene description, and provides the occurrence probability data, the action data, and the predictive trigger activation offload control information to the output device as the low-latency interaction playback, and the predictive trigger activation offload control information is information used by the output device to control the low-latency interaction playback in which the output device predicts the interaction.

19. The information processing device according to claim 18, wherein the predictive trigger activation offload control information includes a predictive trigger activation offload control threshold value indicating a threshold value used to determine the threshold value of the occurrence probability.

20. An information processing method comprising: selecting an action that corresponds to an interaction in a 3D space and that is capable of low-latency interaction playback based on low-latency processing information that includes a maximum processing delay and controls low-latency interaction playback, the low-latency interaction playback being stored in a scene description that represents a scene in a 3D space in which a 3D object is placed; and performing the low-latency interaction playback for the selected action; wherein the low-latency interaction playback is a process in which interaction playback that executes the action in response to the occurrence of the interaction as a trigger is executed with a delay lower than the maximum processing delay; and wherein the maximum processing delay indicates the maximum allowable time from the occurrence of the interaction to the output of action data of the action on an output device.

Citation Information

Patent Citations

  • Information processing device, information processing method, and information processing system

    WO2024009653A1

  • Information processing device and method

    WO2024024874A1

  • Data distribution system, data distribution method, data processing device, and data processing method

    WO2024034646A1